International network attack prediction method and system based on fusion of GNN and LLM

By constructing a multi-view dynamic graph neural network and a large language model to integrate multi-modal data, the problems of data integration and dynamic changes in network attack prediction among countries are solved, and more accurate and timely prediction results are achieved.

CN120389903APending Publication Date: 2025-07-29INST OF GEOGRAPHICAL SCI & NATURAL RESOURCE RES CAS

Patent Information

Application Number
CN202510716382.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

The prior art is difficult to effectively integrate multimodal data and consider dynamic changes in the prediction of inter-country cyberattacks, resulting in poor prediction accuracy and timeliness.

Method used

By constructing a multi-view dynamic graph neural network, combining large language models and LoRA methods, it integrates news events, cyber attack history, satellite remote sensing images and social media public opinion data to generate prediction results of cyber attacks among countries.

Benefits of technology

Achieve more accurate and timely prediction of inter-country cyberattacks, providing reasoning process-assisted decision-making support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120389903A_ABST
    Figure CN120389903A_ABST
Patent Text Reader

Abstract

The invention provides an inter-country network attack prediction method and system based on GNN and LLM fusion, and is applied to the field of network security. On the basis of a news event data set, a big language model is used for carrying out content analysis on a news text, in combination with network attack historical data, military facility change features in a satellite image are converted into text semantic embedding through a cross-modal alignment module, and a network attack prediction-oriented multi-modal data set is generated; the multi-modal data set is processed, layering is carried out according to time granularity, each layer of graph comprises country nodes, event nodes and relation nodes, features of different levels are aggregated through a dynamic time window, and a target directed multi-view dynamic graph composed of network attacks, news events, image events and public opinion events is generated; and processing the target event set based on the big language model in combination with the semantic information of the fusion graph data and the text data, and generating inter-country network attack prediction result information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of network security, and in particular to a method and system for predicting inter-country network attacks based on the fusion of GNN and LLM. Background Art

[0002] In the field of cybersecurity, predicting interstate cyberattacks is crucial. The current global cybersecurity landscape is complex and volatile, with interstate cyber confrontations occurring frequently. Traditional technologies face numerous limitations in addressing these challenges. Technically, past cyberattack predictions have often relied on a single data source, failing to fully capture the complex relationships between states. For example, relying solely on historical cyberattack data can provide a basic understanding of the attacks, but it fails to incorporate factors such as international relations and diplomatic dynamics for comprehensive analysis. For example, when predicting the likelihood of a cyberattack from country A against country B, failing to factor in the diplomatic tensions between the two countries stemming from trade frictions can often lead to biased predictions.

[0003] Existing technologies also have shortcomings in data processing and model building. Traditional methods lack efficient fusion methods when processing multimodal data (cyber attack data and news event text data). For example, when generating data sets for analysis, different types of data cannot be effectively integrated, resulting in subsequent modeling difficulties in exploring deep connections between data. In model construction, the models used are difficult to take into account the dynamic changes and multi-view features of the data. For example, traditional graph models cannot effectively interact and fuse features of the two views of cyber attacks and news events in the time dimension like multi-view dynamic graph neural networks, resulting in poor accuracy and timeliness in the prediction of cyber attacks between countries, making it difficult to meet the actual needs of network security protection.

[0004] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of the present disclosure, and therefore includes information that does not constitute prior art known to ordinary technicians in this field. Summary of the invention

[0005] The purpose of this application is to provide a method and system for predicting inter-state cyberattacks based on the fusion of GNN and LLM, which at least to some extent overcomes the problems of the existing technology. By obtaining news event data from GDELT and cyber attack history data from DAM, a multimodal dataset is constructed after analyzing the news text with a large language model, and processed into a target directed multi-view dynamic graph. Node embedding information is generated through multi-view dynamic graph neural network modeling. This information is then input into the large language model and fine-tuned using the LoRA method to achieve semantic fusion of graph and text data. Combined with the target event set, after feature extraction and association analysis, the large language model predicts and verifies the attack prediction result, and outputs the reasoning process to assist in decision-making.

[0006] Other features and advantages of the present application will become apparent from the following detailed description, or will be learned in part through the practice of the present invention.

[0007] According to one aspect of the present application, a method for predicting cross - country cyberattacks based on the fusion of GNN and LLM is provided, including: obtaining a target event set and a training sample set, where the training sample set includes a news event data set, historical cyberattack data, satellite remote sensing image data set, and social media public opinion data set; based on the news event data set, using a large language model to perform content analysis on news texts, combining historical cyberattack data, converting the change features of military facilities in satellite images into text semantic embeddings through a cross - modal alignment module, quantifying the sentiment polarity of social media public opinion into numerical features, and generating a multi - modal data set for cyberattack prediction; processing the multi - modal data set, stratifying it by time granularity, where each layer of the graph includes country nodes, event nodes, and relationship nodes, aggregating different - level features through a dynamic time window, and generating a target directed multi - view dynamic graph composed of cyberattacks, news events, image events, and public opinion events; modeling the target directed multi - view dynamic graph based on a multi - view dynamic graph neural network to generate node embedding information; injecting the node embedding information into each layer of the large language model, performing fine - tuning using the LoRA method, introducing an adversarial fusion module, and at the same time generating a causal explanation chain through attention weight visualization to form semantic information of the fused graph data and text data; processing the target event set based on the large language model combined with the semantic information of the fused graph data and text data, and simulating the change of attack probability in different scenarios through a counterfactual reasoning module to generate cross - country cyberattack prediction result information.

[0008] Another aspect of the present application is a cross - national cyber - attack prediction device based on the fusion of GNN and LLM, which is characterized by including: an acquisition module for acquiring a target event set and a training sample set, where the training sample set includes a news event data set, historical cyber - attack data, a satellite remote - sensing image data set, and a social - media public opinion data set; a processing module for, based on the news event data set, using a large - language model to perform content analysis on news texts, combining the historical cyber - attack data, converting the military - facility change features in satellite images into text semantic embeddings through a cross - modal alignment module, quantifying the sentiment polarity of social - media public opinion into numerical features, and generating a multi - modal data set for cyber - attack prediction; processing the multi - modal data set, stratifying it by time granularity, where each layer of the graph includes country nodes, event nodes, and relationship nodes, aggregating different - level features through a dynamic time window to generate a target directed multi - view dynamic graph composed of cyber - attacks, news events, image events, and public - opinion events; modeling the target directed multi - view dynamic graph based on a multi - view dynamic graph neural network to generate node embedding information; injecting the node embedding information into each layer of the large - language model, using the LoRA method for fine - tuning, introducing an adversarial fusion module, and simultaneously generating a causal explanation chain through attention - weight visualization to form the semantic information of the fused graph data and text data; processing the target event set based on the large - language model combined with the semantic information of the fused graph data and text data, simulating the change of attack probability under different scenarios through a counterfactual reasoning module, and generating cross - national cyber - attack prediction result information.

[0009] According to another aspect of the present application, there is provided a computer - readable storage medium having a computer program stored thereon, and when the computer program is executed by a second processor, it implements the above - mentioned cross - national cyber - attack prediction method based on the fusion of GNN and LLM.

[0010] The cross - national cyber - attack prediction method and system provided by the present application are committed to solving the problem of cross - national cyber - attack prediction and integrating the advantages of GNN and LLM. First, news event data is obtained from GDELT and historical cyber - attack data is obtained from DAM. After analyzing news texts with a large - language model, a multi - modal data set is constructed and processed into a target directed multi - view dynamic graph. Then, through modeling with a multi - view dynamic graph neural network, node embedding information is generated. Then, it is injected into the large - language model and fine - tuned using the LoRA method to achieve semantic fusion of graph and text data. Finally, combined with the target event set, through feature extraction, correlation analysis, etc., the large - language model predicts and verifies to obtain the attack prediction result, and can also output the reasoning process to assist in decision - making.

[0011] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1 The flowchart showing a method for predicting cross - country cyberattacks based on the fusion of GNN and LLM provided by an embodiment of the present application;

[0013] Figure 2 The structural schematic diagram showing a device for predicting cross - country cyberattacks based on the fusion of GNN and LLM provided by an embodiment of the present application;

[0014] Figure 3 The relationship diagram of time interval and data volume of a method for predicting cross - country cyberattacks based on the fusion of GNN and LLM provided by an embodiment of the present application. Detailed implementation manners

[0015] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are only for the purpose of illustration and explanation of the present invention, and are not used to limit the present invention.

[0016] The following combines Figure 1 to describe a method for predicting cross - country cyberattacks based on the fusion of GNN and LLM according to an exemplary embodiment of the present application. In one embodiment, the present application also proposes a method and system for predicting cross - country cyberattacks based on the fusion of GNN and LLM. As Figure 1 shown, the method is applied to a server and includes:

[0017] S101, obtaining a target event set and a training sample set.

[0018] In one embodiment, the news event data set selects the Global Database of Events, Language, and Tone (GDELT). This database collects news events worldwide and can provide rich event information; the historical data of cyberattacks uses the Digital Attack Map (DAM) data set, which records the process of cross - country cyberattacks; the satellite remote sensing image data set is obtained through a commercial remote sensing platform (such as PlanetLabs) and can provide remote sensing images with a resolution of ≥0.5 meters for sensitive areas such as military bases and ports in the target country; the social media public opinion data set is crawled through the APIs of platforms such as Twitter and Telegram and contains public opinion data. These data sources are extensive and authoritative, and can provide multi - dimensional information for the prediction task.

[0019] Retrieve data from the GDELT database, which provides information such as the URLs of events. Using these URLs, develop a distributed crawler script to obtain the original news web pages. During the retrieval process, for broken links, adopt the strategy of obtaining backup web pages from The Internet Archive to ensure data integrity. After that, use the web parsing tool Trafilatura to obtain the news text and perform duplicate merging operations. Since the same URL may correspond to multiple pieces of data, partly because the web page mentions multiple countries or multiple event types, and partly because of simple duplication, it is necessary to remove duplicate data and merge multiple pieces of data corresponding to the same URL into a larger event data set for subsequent analysis.

[0020] For satellite remote sensing image data, regularly obtain images of the target area through the API of a commercial remote sensing platform. Use object detection models such as YOLOv8 to identify military facilities (such as missile launch vehicles, radar stations) in the images, and convert visual features into text semantic descriptions through the CLIP model. Social media sentiment data is crawled in real time through the API by keywords (such as "war", "sanctions"). Use sentiment analysis tools such as VADER to quantify the sentiment polarity of the text and generate an emotion index from -100 to +100.

[0021] The DAM network attack data set records the network attack process between two countries over a period of time. In the original data table, there are several data records with the same fields except for time for the same pair of countries within a certain time period, which are regarded as one network attack event. First, divide the original database according to the attacking country and the attacked country to obtain a series of data item sets with the same attacker and the same attacked. Then sort by the attack time and judge the time difference between adjacent two records. Here, the "inflection point principle" is adopted, and 120 minutes is selected as the time interval threshold. If it is greater than this threshold, it is considered that the two adjacent data belong to two different attacks, otherwise it is regarded as the same attack. Select one of the same attack data to keep and delete the redundant data.

[0022] After the above processing, integrate the obtained news event data, network attack history data, satellite remote sensing image semantic descriptions, and social media sentiment indexes into a training sample set. This training sample set contains rich information. The news event data covers the time, location, participants of the event, and semantic features obtained through large language model analysis, etc.; the network attack history data contains key information such as attack time and attack type; the satellite remote sensing image data provides physical evidence of changes in military facilities; the social media sentiment data reflects the public's emotional tendency, laying a foundation for the subsequent generation of a multi-modal data set for network attack prediction and being the data cornerstone of the entire prediction model.

[0023] S102, based on a news event dataset, uses a large language model to perform content analysis on news texts. Combined with historical cyberattack data, the cross-modal alignment module converts the change features of military facilities in satellite images into text semantic embeddings, quantifies the emotional polarity of social media public opinion into numerical features, and generates a multimodal dataset for cyberattack prediction.

[0024] In one embodiment, based on a news event dataset, a crawler script is used to obtain the original news web page, which is processed by a parsing tool and combined with a duplicate merge operation to generate the original news text information. The URL is extracted from the news event dataset (GDELT database), and the original news web page is obtained using a distributed crawler script. During the acquisition process, for invalid links, backup web pages are obtained with the help of The Internet Archive to ensure data integrity. For example, when processing the 2020 GDELT event database, it was found that some URL links were invalid. By obtaining backup web pages from The Internet Archive, the relevant web page content was successfully obtained. Afterwards, the web page parsing tool Trafilatura was used to parse the original web page and extract the news text. Since the same URL may correspond to multiple data, a duplicate merge operation is required. When processing a batch of data, it was found that 10 data corresponded to the same URL, 3 of which were repeated because the web page mentioned multiple countries, and 2 were simply repeated. These 10 data were merged into an event data set to finally generate the original news text information.

[0025] At the same time, remote sensing image data (resolution ≥ 0.5 meters) of the target country's military bases, ports, and communication hubs are obtained through the satellite image API, and the YOLOv8 model is used to detect military facilities in the image (such as missile launchers and radar stations). The CLIP model is then used to map the visual features into textual semantic descriptions (such as "Satellite images show that three new radar facilities have been added to the border base of Country A"); public opinion data from platforms such as Twitter and Telegram are crawled through the social media API, and the VADER sentiment analysis tool is used to quantify the sentiment polarity of the text to generate a sentiment index ranging from -100 to +100 (such as "Anti-government sentiment index +75").

[0026] For the original news text information, a large language model (such as Qwen-2.5) is used to judge invalid content, judge the relevance to international relations, identify national roles, and generate event summaries, generating news text analysis information. In the judgment of invalid content, the model will identify invalid content such as "page not found" and "inaccessible" and mark them; in terms of the judgment of the relevance to international relations, the model screens out texts related to bilateral relations and eliminates irrelevant news such as domestic sports events in a certain country; when identifying national roles, the model determines the roles of each country from the news text. For example, for the news that "Country A provides economic assistance to Country B", it identifies A as the aid provider (event initiator) and B as the recipient (event bearer); in addition, the model generates summaries for long texts to reduce redundant information.

[0027] Through the LLM-Image cross-modal converter, the change features of military facilities in satellite images (such as runway expansion and troop concentration) are converted into text semantic embeddings (such as "the intensity of military activities has increased"); a sentiment analysis model is used to quantify the polarity of social media public opinion, generating a numerical public sentiment index (such as "Social Media Anti-US Sentiment Index -80").

[0028] For the historical data of cyberattacks, it is divided by the attacking country and the attacked country, sorted according to the attack time, and events are aggregated according to a time difference threshold of 120 minutes. After screening the data, processed cyberattack data information is generated. The DAM dataset records the cyberattack process between two countries over a period of time. First, it is divided by the attacking country and the attacked country. For example, the attack data of the United States on Russia is grouped together; then the data within each data set is sorted according to the attack time, and the "inflection point principle" is used to determine 120 minutes as the time difference threshold for aggregating events. For example, in the "United States - Russia" data set, the attack record times of 10:00 and 12:30 (time difference of 150 minutes) are regarded as two attacks, and 10:00 and 11:30 (time difference of 90 minutes) are regarded as one attack; after aggregating the attack events, duplicate or invalid data such as missing attack times and unclear types are screened out to generate processed cyberattack data information. An attack and event time correlation matrix is constructed to establish a time sequence correlation between each attack and news events, satellite image changes, and public opinion fluctuations within 72 hours before the attack. For example, a change in military deployment shown in satellite images 24 hours before the attack + a sharp increase in anti-sentiment on social media 12 hours before the attack → marked as a "highly correlated attack".

[0029] Based on news text analysis information, processed cyber-attack data information, semantic descriptions of satellite images, and social media sentiment indices, they are fused through a cross-modal alignment module: Entity alignment: Uniformly map the country entities in the news, the geographical coordinates in the satellite images, and the IP ownership locations in the attack data; Semantic alignment: Through contrastive learning, map the news text of "military exercises", the troop concentration features in the satellite images, and the tension sentiment indices in social media to the same semantic space; Temporal alignment: Construct a time synchronization mechanism to ensure the association of multi-source data within the same time window (such as aligning the satellite image changes at 14:00 on May 10, 2023 with the news report at 14:30); Feature enhancement: Generate composite features, such as the attack risk index = 0.4 × degree of change in military deployment + 0.3 × intensity of diplomatic conflicts + 0.2 × public hostility + 0.1 × historical attack frequency. Finally, generate a multi-dimensional tensor representation containing text features, image features, numerical features, and temporal features as a multi-modal dataset for cyber-attack prediction.

[0030] Taking "Country X launched a DDoS attack on Country Y at 10:00 on October 10, 2024" as an example, integrate the following information: Attack data: Attack time, type; News text: There was a diplomatic friction between Country X and Country Y regarding trade issues, Country X was the initiator, and Country Y was the recipient; Satellite image: Three days before the attack, there were newly added missile launch vehicles at the military base on the border of Country X (semantic description); Public opinion data: One day before the attack, the anti-sentiment index of social media towards Country X was -90. Form a multi-dimensional data record, so that the data not only contains direct attack information but also incorporates multi-dimensional backgrounds such as military deployment and social sentiment, providing more comprehensive feature inputs for the prediction model.

[0031] S103. Process the multi-modal dataset, stratify it by time granularity. Each layer of the graph contains country nodes, event nodes, and relationship nodes. Aggregate features at different levels through a dynamic time window to generate a target directed multi-view dynamic graph composed of cyber-attacks, news events, image events, and public opinion events.

[0032] In one implementation, network attack, news event, image event, and public opinion event data are respectively extracted from the constructed multi-modal dataset to generate corresponding subsets. For example, the multi-modal dataset contains network attack data, news event data, satellite image data, and social media public opinion data among multiple countries from January 1 to December 31, 2020. Through data screening and extraction operations, the data related to network attacks is extracted to form a subset, including information such as attack time, initiating country, target country, and attack type; the news event data is extracted to form a subset, covering release time, involved countries, content summary, etc.; the satellite image data is extracted to form an image event subset, including semantic descriptions of military facility changes, geographical coordinates, etc.; the social media public opinion data is extracted to form a subset, including sentiment index, keywords, etc.

[0033] Taking countries as nodes, the event initiator as the source node, and the recipient as the target node, node information is generated by combining satellite image geographical coordinate mapping and public opinion entity recognition. In the network attack data subset, if there is a record of "On March 5, 2020, Country A launched a DDoS attack on Country B", Country A is the source node and Country B is the target node; if there is a report of "Trade negotiation between Country A and Country B" in the news event subset, it also involves Countries A and B. In addition, the country to which the military facility belongs is determined through satellite image geographical coordinate mapping, such as a military base located on the border of Country A; the countries mentioned in the social media are determined through public opinion entity recognition, such as a tweet mentioning "the government of Country C". Each country node is attached with attributes such as GDP and population (obtained from other data channels), the image event node contains semantic descriptions such as "military exercise", and the public opinion event node contains numerical attributes such as "anti-US sentiment index -80" to form node information.

[0034] Based on the data of each subset, edge information is generated according to the event type and time causal relationship. The network attack data has an average of 1003 edges per graph, the news event data has an average of 710 edges per graph, the image event data has an average of 420 edges per graph, and the public opinion event data has an average of 580 edges per graph. For the network attack data, Country A launched a DDoS attack on March 5, 2020, and a SQL injection attack on May 10, forming edges between Countries A and B respectively, and the edge attributes are attack time and type; for news events such as trade negotiations and diplomatic statements between Countries A and B, edges are formed and annotated with summary and time; in the image event, a new missile launch vehicle is added to the base on the border of Country A, forming an edge of "Country A - image event", and the attribute is "military activity intensity +0.6"; in the public opinion event, the anti-A country sentiment index of the people in Country B is -90, forming an edge of "Country B - public opinion event", and the attribute is "emotional polarity -90". At the same time, edges across event types are generated according to the time causal relationship, such as a temporal correlation edge with a weight of 0.8 is generated between the military deployment change in the satellite image three days before the attack and the attack event.

[0035] Construct a directed graph based on the generated node information and edge information, divide it into four views according to cyber attacks, news events, image events, and public opinion events, and generate the preliminary structure of a directed multi-view dynamic graph. This structure contains 214 country nodes and 4 types of event nodes (news / attack / image / public opinion). For example, in the cyber attack view, the edge from country A's node to country B's node is labeled "DDoS attack on March 5, 2020"; in the image event view, country A's node is connected to the "border base expansion" image event node, and the edge attribute is "the runway was extended by 500 meters on March 2, 2020". Optimize the cross-view edge association through the graph attention mechanism, such as calculating the association weight between the news event "trade sanctions" and the image event "military exercise", and enhancing the attention coefficient of the "sanctions-exercise" cross-view edge, and finally generate the target directed multi-view dynamic graph.

[0036] In this dynamic graph, the interaction relationships between countries are presented through four views: the cyber attack view shows the attack historical trajectory, the news event view reflects diplomatic dynamics, the image event view provides evidence of military deployments, and the public opinion event view reflects social sentiment. For example, from the multi-view association of the expansion of country A's military facilities in the image view, country A's trade sanctions against country B in the news view, and the high anti-A sentiment among the people of country B in the public opinion view, it can be intuitively inferred that the cyber attack risk between countries A and B has increased. This graph structure provides multi-dimensional features for subsequent multi-view dynamic graph neural network modeling, and through cross-view feature fusion and time-stratified analysis, improves the spatio-temporal correlation and causal interpretability of attack prediction.

[0037] S104, perform modeling processing on the target directed multi-view dynamic graph based on the multi-view dynamic graph neural network to generate node embedding information.

[0038] In one implementation, the target directed multi-view dynamic graph consists of cyber attacks and news events and contains information such as nodes and edges. Analyzing its structure is to extract the key elements in the graph. For example, when analyzing the relationship between country A and country B, obtain node information from the graph, including the relevant attributes of country A and country B (such as GDP, population, etc.), and edge information, such as the type and time of cyber attacks, the summary and occurrence time of news events, etc. This information is organized into graph structure information, providing a basis for subsequent processing, just like disassembling a complex network relationship graph into a list of parts that are easy to understand and process.

[0039] Perform parameter initialization and model loading processing on the multi-view dynamic graph neural network to generate model initialization information. Taking the multi-view dynamic graph neural network as an example, setting appropriate initial parameters can help the model converge faster and avoid falling into local optimal solutions during the training process. The parameters that need to be initialized include the learning rate and weight decay coefficient, etc. The learning rate controls the step size of parameter updates during the training process. If the learning rate is set too large, the model will skip the optimal solution during training, resulting in non-convergence; if set too small, the training speed of the model will be extremely slow, requiring more training time and computing resources. Set the initial learning rate to 0.001. This value is determined through multiple experiments and optimizations. At this learning rate, the model can adjust parameters within a reasonable time, gradually approaching the optimal solution, effectively balancing the training speed and convergence effect. For example, during the training process, the model calculates the gradient based on the current parameters and loss function. When the learning rate is 0.001, the model will update the parameters in the direction of the gradient with an appropriate step size, enabling the model to move in the direction of reducing the loss.

[0040] The weight decay coefficient is used to prevent model overfitting. During the training process, the model will over-learn the noise and details in the training data, resulting in a decline in generalization ability on new data. The weight decay coefficient constrains the model parameters, preventing the parameter values from being too large, thereby avoiding the model from being too complex. Set the weight decay coefficient to 0.0001, which helps to appropriately adjust the weights of the model during training, suppress the excessive growth of parameters, and improve the generalization ability of the model. For example, when dealing with data on cyberattacks between countries, the model may learn some accidental but unrepresentative attack patterns. The weight decay mechanism can reduce the impact of these patterns on the model, making the model pay more attention to features with universality and regularity.

[0041] Utilize the model initialization information to process the graph structure information through the graph self-attention mechanism. The graph self-attention mechanism enables the model to automatically learn the importance weights between nodes. At a certain moment, there is a record in the cyberattack view that country A launches a DDoS attack on country B, and there is news about trade negotiations between countries A and B in the news event view. The graph self-attention mechanism will calculate the attention weights of each node to other nodes based on the attributes of the nodes (country A, country B) and the information of the edges between them. For example, for the node of country A, when calculating its features, it will pay more attention to the edges related to the DDoS attack and the parts related to itself in the news event of trade negotiations. The calculation formula is: where q, k, and v are the query, key, and value respectively. In this example, q is the feature vector of the node of country A, K is the key vector matrix of all nodes (including country A and country B), v is the value vector matrix of all nodes, and d kis the dimension of the key vector. Through this calculation, the initial node feature information of each node under the current view and snapshot is generated, and these features contain the relevant information of the node itself and its adjacent nodes.

[0042] The initial node feature information is processed using temporal interaction attention to learn the evolution pattern over time. Meanwhile, the complementary information of different views is utilized to generate the temporally fused feature information. The temporal interaction attention mechanism is used to capture the changes in node features over time and the complementary information between different views. Over a period of time, the frequency and type of cyberattacks between Country A and Country B have changed, and at the same time, news events have been continuously updated, covering the development of bilateral relations. The temporal interaction attention mechanism analyzes the initial node feature information at different times. For example, it compares the cyberattack-related features and news event-related features of Country A at different time points. It can be found that when the trade negotiations between Country A and Country B reach an impasse, the frequency of cyberattacks increases in the subsequent period. In this way, the evolution pattern over time is learned, and the information from different views (cyberattack view and news event view) is fused. The calculation formula can be expressed as: where M is the fused feature, α t is the attention weight at time t, and h t is the initial node feature information at time t. Through calculation, the temporally fused feature information containing temporal dynamic information and view complementary information is generated, enabling the model to better understand the changes in relations between countries over time.

[0043] The multi-graph attention mechanism is a technique that can integrate features from multiple views. Its core lies in using attention weights to measure the importance of different view features for generating node embedding information. In the scenario of cross-country cyberattack prediction involved in the present invention, there are two views: cyberattack and news event. The cyberattack view contains information such as attack time, attack type, attacking country, and target country; the news event view covers news release time, involved countries, news content summary, etc. The multi-graph attention mechanism calculates the corresponding attention weights based on the information of nodes and edges in each view, thereby determining the importance that should be assigned to the features of each view when generating node embedding information.

[0044] After the features of different views are weighted and fused through the multi-graph attention mechanism, node embedding information containing rich information will be generated. These node embedding information are no longer limited to the features of a single view, but integrate the key information of multiple views. Taking the relationship analysis between Country A and Country B as an example, in the network attack view, Country A launched multiple different types of attacks on Country B within a period of time. At the same time, in the news event view, there are news events such as trade negotiations and diplomatic statements between the two countries. The multi-graph attention mechanism will comprehensively consider this information from different views. It will analyze the features of the nodes (Country A and Country B) in each view, as well as the attributes of the edges between the nodes. The generated node embedding information contains both the historical behavior features of Country A and Country B in terms of network attacks and the dynamic change features of the relationship reflected by the news events between the two countries. This comprehensive node embedding information can represent the features of the country nodes more comprehensively, provide more valuable input for the subsequent large language model, and help improve the accuracy of predicting network attacks between countries.

[0045] The multi-graph attention mechanism will implement feature fusion and generation of node embedding information through a series of calculation steps. There are V views. For each node i, its feature representation in the v-th view is The multi-graph attention mechanism will calculate the attention weight α v for each view, and then obtain the embedding information z i of node i by weighted summation. The formula is as follows: Among them, the calculation of the attention weight α v is usually based on node features and edge information, and is normalized through the softmax function to ensure that the sum of the weights is 1. Through the multi-graph attention mechanism, the features from multiple views can be effectively fused into the node embedding information, providing more representative and comprehensive data for subsequent model processing.

[0046] S105, Inject the node embedding information into each layer of the large language model, use the LoRA method for fine-tuning, introduce an adversarial fusion module, and at the same time generate a causal explanation chain through attention weight visualization to form the semantic information of the fused graph data and text data.

[0047] In one implementation, the node embedding information is processed for dimension adaptation and feature space alignment to generate the adapted node embedding information. After the multi-view dynamic graph neural network generates the node embedding information, since the dimensions and feature spaces of this information may not match the input requirements of the large language model, dimension adaptation and feature space alignment are required. Taking the scenarios of network attack and news event analysis in Country A and Country B as an example, the node embedding information is initially a high-dimensional vector with a dimension of

[512] , while the large language model expects an input vector dimension of

[1024] . At this time, a specific linear transformation matrix needs to be used to process the node embedding information. Let the linear transformation matrix be W transform , whose dimension is [512, 1024]. By multiplying the matrix W transform by the original node embedding vector, the dimension of the node embedding information is converted from

[512] to

[1024] . In terms of format, if the original node embedding information is in a tensor format while the large language model requires a list format, then corresponding format conversion operations are needed. In addition, through the domain classifier in the adversarial fusion module (such as a binary classifier for discriminating graph features and text feature domains), the feature domain difference loss is calculated, and the feature space alignment parameters are optimized in the reverse direction, forcing the graph features and text features to be distributed consistently in the semantic space, and finally generating the adapted node embedding information so that it can be smoothly input into the large language model.

[0048] Select the large language model (LLaMA) for loading and freeze its pre-trained weights after loading. This is because the pre-trained weights have been learned on a large amount of data and contain rich general knowledge. Freezing these weights can avoid unnecessary modifications during subsequent fine-tuning, while reducing the computational cost of training and avoiding overfitting. The loaded LLaMA model has L layers, and the weight matrices of each layer are W layer1 , W layer2 ,..., W layerL . Freezing these weights means that they will no longer be updated during the training process. At the same time, initialize the trainable rank decomposition matrices of LoRA. In the LoRA method, for the weight matrix W layer of each layer, it is approximated by two low-rank matrices A and B, that is, W layer= W0 + δW = W0 + B × A. Here, W0 is the original weight matrix, and δW is the adjusted part approximated by the low-rank matrix product. During initialization, matrices A and B are usually initialized with small random values, and cross-modal alignment constraints are injected: calculate the attention weights of the pre-trained LLM for typical cross-modal samples (such as "military exercise news - satellite image features"), and use them as prior knowledge for initializing matrix A. For example, align the initial value of matrix A with the activation pattern of text features related to "military". For example, for the A matrix of a certain layer with dimensions [r, d] and the B matrix with dimensions [d, r] (where r is the rank of the low-rank decomposition and d is a parameter related to the dimension of the original weight matrix), the values of these two matrices can be randomly initialized using a uniform distribution or a normal distribution, such as A ~ U(-∈, ∈), B ~ U(-∈, ∈), where ∈ is a small positive number, such as 0.01. After these operations, the initial state information of the generation model is generated, preparing for injecting node embedding information into the model and fine-tuning later.

[0049] According to the layer structure of the large language model, the adapted node embedding information is injected layer by layer. Taking the LLaMA model with L layers as an example, starting from the first layer, the adapted node embedding information is fused with the original input (such as the word embedding representation of text data) as an additional input. The input of the i-th layer is X input (including text-related feature representations), and the adapted node embedding information is a node . During injection, the two are combined through concatenation or other fusion methods. For example, using the concatenation method, the input of the i-th layer is updated to [X input , a node . At the same time, an attention consistency module is inserted between Transformer layers. This module calculates the cross-attention weights between graph features and text features, forcing their representations in the intermediate layer to maintain semantic consistency. For example, by minimizing the cosine distance loss between graph features and text features, cross-modal feature interaction is ensured. In this way, after layer-by-layer injection operations, each layer contains node embedding information and text-related information, generating model input information with node embeddings, providing a basis for the model to learn the features of graph data and text data simultaneously in subsequent training.

[0050] In the cross-country cyber attack prediction model based on the integration of GNN and LLM, the LoRA method is used to fine-tune the model input information with node embeddings. LoRA freezes the weights of the pre-trained large language model (such as LLaMA) and injects trainable rank decomposition matrices into each layer of the model, thus achieving efficient fine-tuning of the model with a minimal number of parameters. The weight matrix of a certain layer of the model is W. According to the LoRA method, it is decomposed into W = W0 + δW, where W0 is the pre-trained weight and remains unchanged, and δW is approximated by the product of two low-rank matrices A and B, that is, δW = BA. During training, only matrices A and B are updated. To illustrate with a simple example, if the original dimension of the weight matrix of this layer is d×d (d = 1024), and the low rank r = 16 is set in LoRA, then the dimension of A is d×r (i.e., 1024×16), and the dimension of B is r×d (i.e., 16×1024).

[0051] During the training process, according to the backpropagation algorithm, the gradients of the loss function with respect to A and B are calculated using the model input information with node embeddings, and the values of A and B are updated in the way of gradient descent. At this time, the loss function includes the generation loss (such as the cross-entropy between the attack probability prediction and the true value) and the adversarial loss (the discrimination error of the domain classifier for graph-text features). By synchronously optimizing these two types of losses, the fusion effect of cross-modal features is enhanced. For example, using the Adam optimizer, its update formula is: where, A t and B t are the matrix values of the current training step, α is the learning rate (set to 0.0001), and are the gradients of the loss function with respect to A and B at the current step. Through continuous iterative training, A and B are gradually updated, enabling the model to better adapt to the cross-country cyber attack prediction task. As the training progresses, A and B are continuously updated, and finally the fine-tuned model parameter information is obtained. These parameter information contain the optimization results after fusing the features of graph data and text data. After multiple rounds of training (the number of training rounds is T = 5), the values of A and B will converge to a suitable state. At this time, the updated A and B are combined with the pre-trained weight W0 to obtain the fine-tuned weight matrix W final = W0 + BA. Each layer of the model is processed in the same way, so as to obtain the fine-tuned parameter information of the entire model. These parameter information record the optimization results of the model in the process of learning the data features related to cross-country cyber attacks, providing a more accurate model parameter basis for subsequent joint semantic modeling.

[0052] Taking the large language model with the Transformer architecture as an example, it contains multiple attention layers and fully connected layers inside. During joint semantic modeling, the node embedding information of graph data and the word embedding representation of text data interact with each other in these layers. In the attention layer, the calculation of the query, key, and value matrices comprehensively considers the characteristics of both types of data. The node embedding information is h node , and the word embedding representation of text data is h text . For a certain attention head, when calculating the query matrix Q, key matrix K, and value matrix V, h node and h text are combined in a specific way. For example, Q = W Q [h node ; h text , K = W K [h node ; h text , V = W V [h node ; h text , where W Q , W K , W V are learnable weight matrices, and [;] represents the concatenation operation. In this way, the model can capture the correlation between graph data and text data. For example, it can analyze the potential patterns of cyberattacks between countries (graph data) in the context of specific news events (text data).

[0053] In the multi-layer structure of the model, through multiple attention calculations and fully connected layer transformations, the semantics of graph data and text data are continuously fused and more representative features are extracted. At each layer, the model adjusts the feature representation according to the input data and the current parameters, so that the fused features can better reflect the internal connection between cyberattacks between countries and related news events. For example, in a certain layer, the fused features are processed through a non-linear activation function (such as the ReLU function: f(x) = max(0, x)) to enhance the expression ability of the model. The fused feature vector x, after being processed by the ReLU function, gets y = f(x). This operation can highlight important features and suppress irrelevant information, so that the model can focus more on the key semantic information related to cyberattack prediction.

[0054] After multiple layers of joint semantic modeling, the model finally outputs semantic information that fuses graph data and text data. These information are a comprehensive understanding of the network attack-related situations between countries, including various potential patterns and relationships mined from the data. For example, the model output includes semantic vector representations in aspects such as the relationship situation between countries, the trend of attack possibility, and the impact degree of relevant events. Taking Country A and Country B as an example, the fused semantic information may indicate that under the current trade friction (reflected by text data), the probability of Country A launching a specific type of cyber attack on Country B (reflected by graph data) has an upward trend in the next period of time, and this trend is closely related to the recent diplomatic dynamics between the two countries. At the same time, through the attention weight visualization tool, the attention head weights related to "military exercise image features - border conflict news" in the Transformer layer are extracted to generate a causal explanation chain, such as "satellite images show the missile deployment of Country A → news reports on the military exercise of Country A → GNN extracts the attack chain features → LLM predicts that the attack probability increases by 23%", forming an interpretable reasoning path. These fused semantic information provide rich and valuable inputs for the subsequent prediction of cyber attacks between countries, helping to improve the accuracy and reliability of the prediction.

[0055] S106, process the target event set based on the large language model combined with the semantic information of the fused graph data and text data, and simulate the change of attack probability in different scenarios through the counterfactual reasoning module to generate the prediction result information of cyber attacks between countries.

[0056] In one implementation, based on the large language model, feature extraction processing is performed on the semantic information of the fused graph data and text data to generate fused data feature information. Taking Country A and Country B as an example, the semantic information of the fused graph data and text data includes the historical behaviors of the two countries in terms of cyberattacks (such as graph data features like attack types and frequencies), satellite image semantic embeddings (such as text descriptions of military base expansions), public opinion sentiment indices (such as the anti-US sentiment index on social media -80), and the semantics contained in relevant news events (such as text data features like diplomatic relations and trade dynamics). The large language model deeply analyzes this information through its internal multi-layer neural network structure. The model will perform semantic understanding on the words and sentences in the text, and at the same time, combine the features of nodes and edges in the graph data, satellite image semantic embeddings, and public opinion sentiment indices to extract representative fused data features. For example, through the processing of the fused information "Recently, the trade negotiation between Country A and Country B broke down. Satellite images of a military base on the border of Country A show new missile launch vehicles, and the anti-A sentiment index on social media reaches -90", the model extracts features such as "Degree of trade relationship tension +0.8", "Military deployment threat level +0.7", "Hostility of social sentiment +0.9", "Upward trend of cyberattack risk". In terms of technical implementation, the model uses the attention mechanism to assign different weights to different parts of the information to highlight key information. The fused data is represented as X fusion , and the feature weight calculated by the attention mechanism is α. After weighted processing, fused data feature information F fusion is generated. The calculation formula can be expressed as: where represents different parts of the fused data, and α i is the corresponding weight. The large language model can extract valuable features for subsequent predictions from complex fused information.

[0057] The target event set includes a large number of historical cyberattack sequences, historical news event text sequences, image event time series, and public opinion event time series. For historical cyberattack sequences, key information such as attack time, attack type, and attack intensity needs to be parsed. For example, in a historical cyberattack sequence, the record states, "Country A launched a DDoS attack against Country B on January 5, 2024. The attack lasted for two hours and impacted some critical network services in Country B." From this record, features such as the attack time (January 5, 2024), attack type (DDoS attack), and attack intensity (affecting some critical network services) can be extracted. For historical news event text sequences, a large language model is used for parsing to extract information such as the countries involved, event themes, and event sentiment. For example, from the news text "'Countries A and B issued strong statements on a territorial dispute, and relations between the two sides are tense,'" features such as the countries involved (Countries A and B), the event theme (territorial dispute), and the event sentiment (tension) can be extracted. For time series of image events, the system analyzes the time points and semantic descriptions of changes in military facilities (e.g., "On January 3, 2024, a new radar facility was added to a border base of Country A"). For time series of public opinion events, the system analyzes the temporal fluctuations of public sentiment indexes (e.g., "On January 4, 2024, the social media anti-Country A sentiment index dropped to -90"). By extracting and organizing this information, it generates characteristic information for the target event, providing basic data for subsequent correlation analysis.

[0058] The fused data feature information and the target event feature information are correlated and analyzed, and the causal chain of multi-source events within a dynamic time window (such as 72 hours before the attack) is combined to explore the potential connection between the two. Taking the previously extracted features as an example, the "degree of tension in trade relations" and "military deployment threat level" in the fused data features and the "territorial disputes", "emotional tendencies of tense relations between the two sides", "changes in military deployment in satellite images 3 days before the attack", and "surges in negative public opinion 1 day before the attack" in the target event features may be causally related. Through analysis, it can be found that when trade relations are tense, there are territorial disputes, and the escalation of military deployment and the increase in hostility in public opinion occur continuously within 72 hours, the possibility of a cyber attack will increase significantly. When performing correlation analysis, some statistical methods or machine learning algorithms can be used. Use cosine similarity to measure the degree of correlation between features. For the fused data feature vector F fusion and the target event feature vector F fusion , the cosine similarity calculation formula is: By calculating the similarity between different feature vectors, it is possible to determine which fused data features are closely related to the target event features. If the calculated cosine similarity between "Degree of trade relation tension" and "Territorial disputes" and "Emotional tendency of strained bilateral relations" is relatively high, it indicates a strong association between them. After organizing this association information, association analysis result information is generated, providing strong support for subsequent predictions.

[0059] Input the association analysis result information into the prediction module of the large language model. Use the large language model combined with the counterfactual reasoning module to simulate the change of attack probability in different scenarios, and generate preliminary cross-country cyber attack prediction information. Taking Country A and Country B as an example, the association analysis results show that in historical cases, when the trade relations between Country A and Country B are tense, there are territorial disputes, the military deployment is upgraded, and the public opinion hostility is high, the probability of cyber attacks is relatively high. The large language model generates target scenarios (such as "Suppose Country A cancels trade sanctions" and "Suppose Country B withdraws military deployment") through the counterfactual reasoning module, and uses the graph neural network to generate the graph structure features after scenario perturbation (such as after canceling trade sanctions, updating the edge weight of trade relations; after withdrawing military deployment, updating the node features of military activities), and inputs them into the large language model to simulate the attack probability distribution in different scenarios. For example, after receiving the original association information, the model first predicts that the probability of Country A launching a cyber attack on Country B within the next week in the basic scenario is 70%, and the attack types may be DDoS attacks or malware attacks targeting critical infrastructure; then simulate the "cancel trade sanctions" scenario. After the graph neural network updates the edge weight of trade relations, the LLM predicts that the attack probability drops to 45%, and the attack type is more inclined to low-intensity network reconnaissance. This preliminary prediction result includes the probability distribution of different scenarios, but still needs to be verified and adjusted.

[0060] Conduct result verification and adjustment processing on the preliminary cross-country cyber attack prediction information, and generate cross-country cyber attack prediction result information with scenario simulation by comparing with the historical counterfactual case library. The historical counterfactual case library stores historical prediction and actual attack data in similar scenarios (such as "After Country C cancelled sanctions against Country D in 2023, the actual attack probability dropped from 60% to 30%"). During verification, compare the prediction result of the current counterfactual scenario with the historical cases with high similarity in the case library. If the prediction error of the basic scenario's predicted attack probability of 70% exceeds 15% compared with a similar case in 2023, adjust the current prediction according to the actual result of the historical case. Combine domain expert knowledge. For example, if experts point out that the current domestic political elections in Country A may cause it to avoid high-intensity attacks, then reduce the high-intensity attack probability by 20% in the "basic scenario" prediction.

[0061] After calibration and adjustment, the final prediction results with scenario simulation are generated. For example: Base scenario: The probability of country A launching a cyber attack against country B is 55%, and the attack types are DDoS attack (probability 40%) or network penetration (probability 35%); Canceling trade sanctions scenario: The attack probability drops to 30%, mainly for low-intensity reconnaissance; Military deployment upgrade scenario: The attack probability rises to 80%, possibly targeting critical infrastructure. This result provides multi-dimensional prediction support for decision-makers through scenario probability distribution and causal explanation chains (such as the conduction path of "trade sanctions → military deployment → public opinion sentiment"), facilitating the formulation of preventive strategies in advance.

[0062] This application aims to solve the problem of predicting cyber attacks between countries, integrating the advantages of graph neural networks (GNN) and large language models (LLM) to improve prediction accuracy and interpretability. First, news event data is obtained from the Global Database of Events, Language, and Tone (GDELT), and historical cyber attack data is obtained from the Digital Attack Map (DAM). The large language model is used to deeply analyze news texts, and a multi-modal data set is constructed by combining cyber attack data, and then processed to generate a target directed multi-view dynamic graph. Then, a multi-view dynamic graph neural network is used to model the graph. Through structure parsing, parameter initialization, graph self-attention encoding, time interaction attention processing, and multi-graph attention fusion, node embedding information is generated. Then, it is injected into each layer of the large language model and fine-tuned using the LoRA method to achieve deep semantic fusion of graph data and text data. Finally, based on the fused semantic information, combined with the target event set, through steps such as feature extraction and correlation analysis, the large language model predicts and calibrates and adjusts to obtain the prediction results of cyber attacks between countries, and at the same time outputs the reasoning process to enhance the decision support ability.

[0063] In one implementation, as Figure 2 shown, this application also provides a device for predicting cyber attacks between countries based on the fusion of GNN and LLM, including:

[0064] An acquisition module 201, configured to acquire a target event set and a training sample set, where the training sample set includes a news event data set, historical cyber attack data, satellite remote sensing image data set, and social media public opinion data set;

[0065] A processing module 202 is configured to perform content analysis on news texts using a large language model based on a news event dataset, combine historical network attack data, convert the change features of military facilities in satellite images into text semantic embeddings through a cross-modal alignment module, quantify the sentiment polarity of social media public opinion into numerical features, and generate a multi-modal dataset for network attack prediction; process the multi-modal dataset, stratify it by time granularity, with each layer graph containing country nodes, event nodes, and relationship nodes, aggregate different-level features through a dynamic time window, and generate a target directed multi-view dynamic graph composed of network attacks, news events, image events, and public opinion events; model and process the target directed multi-view dynamic graph based on a multi-view dynamic graph neural network to generate node embedding information; inject the node embedding information into each layer of the large language model, perform fine-tuning using the LoRA method, introduce an adversarial fusion module, and simultaneously generate a causal explanation chain through attention weight visualization to form semantic information integrating graph data and text data; process the target event set based on the large language model combined with the integrated graph data and text data semantic information, simulate the change of attack probability in different scenarios through a counterfactual reasoning module, and generate national network attack prediction result information.

[0066] Each embodiment in this application is described in a related manner. For the same and similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the embodiments of the method, electronic device, and readable storage medium for predicting international cyber attacks based on the fusion of GNN and LLM, since they are basically similar to the embodiments of the method for predicting international cyber attacks based on the fusion of GNN and LLM described above, the description is relatively simple, and the relevant parts can be referred to the partial description of the embodiments of the method for predicting international cyber attacks based on the fusion of GNN and LLM described above.

Claims

1. A method for predicting cross - country cyber attacks based on the fusion of GNN and LLM, characterized in that, Including: Obtain a target event set and a training sample set, where the training sample set includes a news event data set, network attack historical data, satellite remote sensing image data set, and social media public opinion data set; Based on the news event data set, use a large language model to perform content analysis on news texts, combine network attack historical data, and through a cross-modal alignment module, convert the military facility change features in satellite images into text semantic embeddings, and quantify the sentiment polarity of social media public opinion into numerical features to generate a multi-modal data set for network attack prediction; Process the multi-modal data set, stratify it by time granularity, each layer of the graph contains country nodes, event nodes, and relationship nodes, and aggregate different hierarchical features through a dynamic time window to generate a target directed multi-view dynamic graph composed of network attacks, news events, image events, and public opinion events; Perform modeling processing on the target directed multi-view dynamic graph based on the multi-view dynamic graph neural network to generate node embedding information; Inject the node embedding information into each layer of the large language model, use the LoRA method for fine-tuning, introduce an adversarial fusion module, and at the same time generate a causal explanation chain through attention weight visualization to form the semantic information of the fused graph data and text data; Based on the large language model combined with the semantic information of the fused graph data and text data, process the target event set, and through a counterfactual reasoning module, simulate the change of attack probability in different scenarios to generate the network attack prediction result information between countries.

2. The method according to claim 1, wherein Based on the news event data set, use a large language model to perform content analysis on news texts, combine network attack historical data, and through a cross-modal alignment module, convert the military facility change features in satellite images into text semantic embeddings, and quantify the sentiment polarity of social media public opinion into numerical features to generate a multi-modal data set for network attack prediction, including: Based on the news event data set, use a crawler script to obtain the original news web pages, process them through a parsing tool, and combine duplicate item merging operations to generate the original news text information; Obtain the remote sensing image data of the military bases, ports, and communication hubs of the target country through the satellite image API, and use the social media API to crawl the public opinion data of the relevant country; For the original news text information, use a large language model to perform invalid content judgment, international relations relevance judgment, country role recognition, and event summary generation to generate news text analysis information; Through the LLM-Image cross-modal converter, convert the military facility change features in satellite images into text semantic descriptions; Use a sentiment analysis model to quantify the polarity of social media public opinion and generate a numerical public sentiment index; For the network attack historical data, divide it by the attacking country and the attacked country, sort it according to the attack time, aggregate events according to a time difference threshold of 120 minutes, and after screening the data, generate the processed network attack data information; Construct an attack and event time correlation matrix to establish a temporal correlation between each attack and the news events, satellite image changes, and public opinion fluctuations within 72 hours before the attack. Based on news text analysis information, processed cyber-attack data information, satellite image semantic descriptions, and social media sentiment indices, they are fused through a cross-modal alignment module to generate a multi-dimensional tensor representation containing text features, image features, numerical features, and temporal features, serving as a multi-modal dataset for cyber-attack prediction.

3. The method according to claim 1, characterized in that Process the multi-modal dataset, stratify it by time granularity. Each layer of the graph contains country nodes, event nodes, and relationship nodes. Aggregate features at different levels through a dynamic time window to generate a target directed multi-view dynamic graph composed of cyber-attacks, news events, image events, and public opinion events, including: Process the multi-modal dataset, stratify it by time granularity. Each layer of the graph contains country nodes, event nodes, and relationship nodes. Aggregate features at different levels through a dynamic time window, and extract data on cyber-attacks, news events, image events, and public opinion events to generate corresponding subsets; Taking countries as nodes, the event initiator as the source node, and the recipient as the target node, combine satellite image geographical coordinate mapping and public opinion entity recognition to generate node information, and generate edge information according to the event type and temporal causal relationship; Construct a directed graph containing 214 country nodes and 4 types of event nodes, divided into 4 views according to cyber-attacks, news events, image events, and public opinion events. Optimize cross-view edge associations through a graph attention mechanism to generate a target directed multi-view dynamic graph.

4. The method according to claim 1, wherein Based on a multi-view dynamic graph neural network, model and process the target directed multi-view dynamic graph to generate node embedding information, including: Conduct structural analysis processing on the target directed multi-view dynamic graph to generate graph structure information; Conduct parameter initialization and model loading processing on the multi-view dynamic graph neural network to generate model initialization information; Based on the model initialization information, process the graph structure information, and use graph self-attention to encode each snapshot of each view to generate initial node feature information; Use temporal interaction attention to process the initial node feature information, learn the evolutionary pattern in time, and at the same time utilize the complementary information of different views to generate temporal fusion feature information; Use multi-graph attention to process the temporal fusion feature information, fuse features from multiple views to generate node embedding information.

5. The method according to claim 4, characterized in that Inject the node embedding information into each layer of the large language model, fine-tune it using the LoRA method, introduce an adversarial fusion module, and at the same time generate a causal explanation chain through attention weight visualization to form the semantic information of the fused graph data and text data, including: Perform dimension adaptation and feature space alignment on the node embedding information, and force the graph features and text features to have the same semantics through the adversarial fusion module; Load the large language model and freeze the pre-trained weights, and inject cross-modal alignment constraints when initializing the LoRA matrix; Inject the adapted node embedding information layer by layer, and insert an attention consistency module between Transformer layers to ensure cross-modal feature interaction; When using LoRA for fine-tuning, synchronously optimize the adversarial loss and the generation loss, and generate a causal explanation chain through an attention weight visualization tool to form the semantic information of the fused graph data and text data.

6. The method according to claim 1, wherein Processing the target event set based on a large language model combined with semantic information of fused graph data and text data, simulating the change of attack probability in different scenarios through a counterfactual reasoning module, and generating information on the prediction results of cross-border cyber attacks, including: Processing the target event set based on a large language model combined with semantic information of fused graph data and text data, generating a target scenario through a counterfactual reasoning module, using a graph neural network to generate graph structure features after scenario perturbation, and inputting them into the large language model to simulate the attack probability distribution in different scenarios; When extracting cross-modal features from the fused data feature information, introducing satellite image semantic embedding and public opinion sentiment index; When analyzing the feature information of the target event, associating the time series of image events and public opinion events; Combining the causal chains of multi-source events within a dynamic time window during the correlation analysis; After the prediction module outputs preliminary prediction information including the scenario probability distribution, it is verified and adjusted by comparing with the historical counterfactual case library to generate information on the prediction results of cross-border cyber attacks with scenario simulation.

7. An inter-country network attack prediction device based on the fusion of GNN and LLM, characterized in that, The device includes: An acquisition module, configured to acquire a target event set and a training sample set, where the training sample set includes a news event data set, historical cyber attack data, a satellite remote sensing image data set, and a social media public opinion data set; A processing module, configured to, based on the news event data set, use a large language model to perform content analysis on news texts, combine historical cyber attack data, convert the change features of military facilities in satellite images into text semantic embeddings through a cross-modal alignment module, quantify the sentiment polarity of social media public opinion into numerical features, and generate a multi-modal data set for cyber attack prediction; process the multi-modal data set, layer it by time granularity, with each layer of the graph including country nodes, event nodes, and relationship nodes, aggregate different-level features through a dynamic time window to generate a target directed multi-view dynamic graph composed of cyber attacks, news events, image events, and public opinion events; perform modeling processing on the target directed multi-view dynamic graph based on a multi-view dynamic graph neural network to generate node embedding information; inject the node embedding information into each layer of the large language model, perform fine-tuning using the LoRA method, introduce an adversarial fusion module, and simultaneously generate a causal explanation chain through attention weight visualization to form the semantic information of the fused graph data and text data; process the target event set based on the large language model combined with the semantic information of the fused graph data and text data, simulate the change of attack probability in different scenarios through a counterfactual reasoning module, and generate information on the prediction results of cross-border cyber attacks.

8. An electronic device, characterized in that, Including: A first processor; And a memory for storing executable instructions of the first processor; Wherein, the first processor is configured to execute the cross-border cyber attack prediction method based on the fusion of GNN and LLM according to any one of claims 1 to 6 by executing the executable instructions.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a second processor, it implements the cross-border cyber attack prediction method based on the fusion of GNN and LLM according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Network attack identification method based on improved Swinin-Transform model

    CN117375924A

  • Identifying a distributed denial of service (DDoS) attack within a network and defending against such an attack

    US20060010389A1

  • Automatic event graph construction method and device for multi-source vulnerability information

    US20230035121A1

Cited By

  • A multi-stage vulnerability mining and threat tracing method and system

    CN122640245A