Network security situation awareness prediction system and method based on large model
Through the cleaning and feature processing of network security logs and traffic data, combined with cross-modal semantic alignment technology, multimodal large models are used to predict attack behavior, the problem of difficulty in capturing multi-stage attack behavior patterns in the existing technology is solved, and high-accurate situational awareness and forward-looking prediction are achieved.
Patent Information
- Application Number
- CN202510688584.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-08-15
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing cybersecurity situation awareness and prediction technologies are difficult to effectively capture the coordinated model of multi-stage attack behavior. Especially when facing massive heterogeneous data, it is impossible to accurately predict the specific stage of the attack, and lacks forward-looking defense capabilities.
By cleaning and analyzing the network security log and network traffic data and extracting key fields, a structured event stream is generated, and modal feature processing is performed. The two are mapped to a unified semantic space using cross-modal semantic alignment technology, fine-grained feature transfer interactive reasoning is performed, and the multimodal big model is used to predict the specific stage of attack behavior.
It realizes accurate capture of the trajectory of attack behavior, improves the accuracy and forward-looking nature of situational awareness, and can more accurately predict the development trend of attack behavior, and meets the needs of refined defense.
Smart Images

Figure CN120498800A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of network security technology, and more specifically, to a network security situation awareness and prediction system and method based on a large model. Background Art
[0002] With the rapid development of information technology and the increasing prevalence of network applications, cyberspace has become a vital support for critical infrastructure and social operations, yet it also faces increasingly severe security threats. In particular, with the construction of new power systems, the digitalization and networking of power systems continues to deepen, leading to profound changes across the energy supply, user, and grid sides. New low-voltage control services, such as new power load management, distributed photovoltaics, and connected vehicle charging stations, are key components of these systems. These widely deployed edge computing devices and peripheral terminals (such as smart energy units and their downstream devices) are characterized by numerous access points, small scale, and uncontrollable physical security. Furthermore, these devices are interconnected with open networks and directly contact the controlled objects, significantly increasing the attack surface and opportunities for the power grid. Attackers can exploit hidden and sophisticated means to launch advanced persistent network attacks (APTs) against a large number of edge devices, thereby threatening the security of core networks and critical business systems. Therefore, developing an effective cybersecurity situational awareness and prediction solution is crucial to comprehensively understand the cybersecurity status of this critical power infrastructure, promptly identify targeted potential threats, and predict future trends in attack behavior.
[0003] Specifically, cybersecurity situational awareness aims to provide a macro, dynamic security perspective, while situational prediction focuses on anticipating attackers' intentions and next actions, thereby providing critical support for security decision-making and proactive defense in power systems. These approaches are a core component in enhancing cybersecurity defense-in-depth capabilities. Currently, several technical attempts at cybersecurity situational awareness and prediction, such as those based on rule-based association analysis and network traffic statistical learning, rely primarily on rule bases and statistical models predefined by security experts, making them incapable of addressing the dynamic and evolving attack behavior characteristics of multi-stage attacks. In particular, when faced with massive amounts of heterogeneous data, existing situational awareness methods often analyze a single data source in isolation, resulting in a disconnect between multimodal information and an inability to effectively capture the coordinated behavioral patterns left by attackers during stages such as infiltration, lateral movement, and data theft. Furthermore, most existing methods focus on post-event alerting or simple risk assessments, and are limited in their ability to accurately predict the specific attack stages (such as reconnaissance, weaponization, delivery, exploitation, installation, command and control, and target achievement), making them incapable of meeting the demands of refined, proactive defense.
[0004] Therefore, we look forward to an optimized network security situation awareness prediction system and method based on a large model. Summary of the Invention
[0005] In order to solve the above technical problems, the present application is proposed. The embodiment of the present application provides a network security situation awareness prediction system and method based on a large model, which first cleans and parses the network security log and extracts key fields to generate a structured event stream, and then performs modal characterization processing to extract its deep semantic feature representation. At the same time, the network traffic data is subjected to time series modeling to learn the time series feature pattern of the network traffic. Then, cross-modal semantic alignment technology is adopted to map the semantic features of the network security log and the time series features of the network traffic to a unified semantic space to achieve feature dimension alignment, and through fine-grained feature transfer interactive reasoning on the two, the potential correlation and collaborative behavior pattern between network traffic and log events are deeply mined, so as to predict the specific stage of the attack behavior based on cross-modal complementary information and a multi-modal large model that has been fine-tuned by the domain. This method can accurately capture the trajectory of attack behavior and improve the accuracy and foresight of situation awareness by making full use of the complementary information between heterogeneous data.
[0006] According to one aspect of the present application, a network security situation awareness prediction method based on a large model is provided, which includes:
[0007] Obtain security log data and network traffic data;
[0008] Performing data cleaning, format parsing, and key field extraction on the security log data to obtain a structured log event stream;
[0009] Performing modal characterization on the structured log event stream to obtain a log event modal semantic feature encoding vector;
[0010] Extracting network traffic time series features from the network traffic data to obtain a network traffic modal time series pattern feature encoding vector;
[0011] The network traffic modal time series pattern feature encoding vector and the log event modal semantic feature encoding vector are input into a multimodal large model fine-tuned for the security field to obtain a security situation awareness result, which is a predicted attack stage label.
[0012] According to another aspect of the present application, a network security situation awareness prediction system based on a large model is provided, which includes:
[0013] Data acquisition module, used to obtain security log data and network traffic data;
[0014] A security log data preprocessing module is used to perform data cleaning, format parsing and key field extraction on the security log data to obtain a structured log event stream;
[0015] A modality characterization module, configured to perform modality characterization on the structured log event stream to obtain a log event modality semantic feature encoding vector;
[0016] A time series feature extraction module, configured to extract network traffic time series features from the network traffic data to obtain a network traffic modal time series pattern feature encoding vector;
[0017] The security situation awareness module is used to input the network traffic modal time series pattern feature encoding vector and the log event modal semantic feature encoding vector into a multimodal large model fine-tuned for the security field to obtain a security situation awareness result, wherein the security situation awareness result is a predicted attack stage label.
[0018] Compared with the existing technology, the network security situation awareness prediction system and method based on a large model provided by this application first cleans and parses the network security log and extracts key fields to generate a structured event stream, and then extracts its deep semantic feature representation through modal characterization processing. At the same time, the network traffic data is time-series modeled to learn the time-series feature pattern of the network traffic. Then, cross-modal semantic alignment technology is used to map the semantic features of the network security log and the time-series features of the network traffic to a unified semantic space to achieve feature dimension alignment, and through fine-grained feature transfer interactive reasoning on the two, the potential correlation and collaborative behavior pattern between network traffic and log events are deeply mined, so as to predict the specific stage of the attack behavior based on cross-modal complementary information and a multi-modal large model that has been fine-tuned by the domain. This method can accurately capture the trajectory of attack behavior and improve the accuracy and foresight of situation awareness by making full use of the complementary information between heterogeneous data. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The above and other purposes, features, and advantages of the present application will become more apparent through a more detailed description of the embodiments of the present application in conjunction with the accompanying drawings. The accompanying drawings are intended to provide a further understanding of the embodiments of the present application and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the present application and do not constitute a limitation of the present application. In the drawings, the same reference numerals generally represent the same components or steps.
[0020] Figure 1 This is a flowchart of a network security situation awareness prediction method based on a large model according to an embodiment of the present application.
[0021] Figure 2 This is a data flow diagram of a large-model-based network security situation awareness prediction method according to an embodiment of the present application.
[0022] Figure 3 This is a flowchart of sub-step S5 of the network security situation awareness prediction method based on a large model according to an embodiment of the present application.
[0023] Figure 4 This is a flowchart of sub-step S51 of the network security situation awareness prediction method based on a large model according to an embodiment of the present application.
[0024] Figure 5 This is a flowchart of sub-step S52 of the network security situation awareness prediction method based on a large model according to an embodiment of the present application.
[0025] Figure 6 This is a flowchart of sub-step S522 of the network security situation awareness prediction method based on a large model according to an embodiment of the present application.
[0026] Figure 7 This is a block diagram of a large model-based network security situation awareness and prediction system according to an embodiment of the present application. DETAILED DESCRIPTION
[0027] As used in this application and the claims, unless the context clearly indicates otherwise, the words "a," "an," "an," and / or "the" are not intended to refer to the singular but may include the plural. Generally speaking, the terms "comprises" and "include" only indicate the inclusion of the steps and elements specifically identified, and these steps and elements do not constitute an exclusive list. A method or apparatus may also include other steps or elements.
[0028] Although the present application makes various references to certain modules in the system according to embodiments of the present application, any number of different modules can be used and run on the user terminal and / or server. The modules are illustrative only, and different aspects of the system and method can use different modules.
[0029] Flowcharts are used in this application to illustrate the operations performed by the systems according to the embodiments of the present application. It should be understood that the preceding or following operations are not necessarily performed in exact order. Instead, the various steps may be processed in reverse order or simultaneously, as needed. Furthermore, other operations may be added to these processes, or one or more operations may be removed from these processes.
[0030] Below, the exemplary embodiments according to the present application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application, and it should be understood that the present application is not limited to the exemplary embodiments described herein.
[0031] It is worth noting that in this application, all actions to obtain data are carried out in compliance with the relevant data protection laws and policies of the country where they are located and with the authorization given by the owner of the corresponding device.
[0032] In response to the technical problems described in the above background technology, this application proposes a network security situation awareness prediction method based on a large model, which first cleans and parses the network security log and extracts key fields to generate a structured event stream, and then performs modal feature processing to extract its deep semantic feature representation. At the same time, the network traffic data is time-series modeled to learn the time-series feature pattern of the network traffic. Then, cross-modal semantic alignment technology is used to map the semantic features of the network security log and the time-series features of the network traffic to a unified semantic space to achieve feature dimension alignment, and through fine-grained feature transfer interactive reasoning on the two, the potential correlation and collaborative behavior pattern between network traffic and log events are deeply mined, so as to predict the specific stage of the attack behavior based on cross-modal complementary information and a multi-modal large model that has been fine-tuned by the domain. This method can accurately capture the trajectory of attack behavior and improve the accuracy and foresight of situation awareness by making full use of the complementary information between heterogeneous data.
[0033] Figure 1 This is a flowchart of a network security situation awareness prediction method based on a large model according to an embodiment of the present application. Figure 2 Figure 1 is a data flow diagram of a network security situation awareness prediction method based on a large model according to an embodiment of the present application. Figure 1 and Figure 2 As shown, the network security situation awareness prediction method based on the large model includes the following steps: S1, obtaining security log data and network traffic data; S2, performing data cleaning, format parsing and key field extraction on the security log data to obtain a structured log event stream; S3, performing modal characterization on the structured log event stream to obtain a log event modal semantic feature encoding vector; S4, extracting network traffic time series features from the network traffic data to obtain a network traffic modal time series pattern feature encoding vector; S5, inputting the network traffic modal time series pattern feature encoding vector and the log event modal semantic feature encoding vector into a multimodal large model fine-tuned for the security field to obtain a security situation awareness result, which is a predicted attack stage label.
[0034] In the above-mentioned large-scale model-based network security situational awareness prediction method, step S1 obtains security log data and network traffic data. It should be understood that since network security situational awareness prediction requires the fusion of multi-source heterogeneous data to fully capture attack behavior characteristics, the attacker's activity traces are usually scattered in security logs (such as system logs, firewall logs) and network traffic (such as traffic packet timing changes, protocol characteristics). Therefore, in order to build a multimodal data input covering the entire life cycle of the attack, this application uses a standardized data acquisition interface to obtain raw security logs and network traffic data from network devices, servers, and security protection systems in real time. Among them, the security log records specific security events and system activities, and the network traffic reflects the behavior pattern of network communication. The combination of the two can provide a more complete view. Specifically, the security log data includes text fields such as event type, timestamp, source / destination IP, and operation instructions, while the network traffic data includes time series data such as traffic size, protocol type, and session duration. In this way, the comprehensiveness and real-time nature of the model input data are ensured, laying a data foundation for subsequent multimodal feature extraction and association reasoning.
[0035] Specifically, security log data typically contains a wealth of information, such as text fields like event type, timestamp, source / destination IP addresses, and operation instructions. Each entry is like a piece of a puzzle that, when combined, can depict the security status of the entire system. By accessing firewall logs, intrusion detection system (IDS) logs, and other relevant log files, clues to abnormal behavior can be captured. It is worth noting that different types of logs may use different formats, so during the collection process, compatibility and parsing accuracy must be considered to avoid information loss or misunderstanding.
[0036] At the same time, network flow data provides another perspective on the behavioral patterns of network communications. This data, consisting of time-series data such as traffic volume, protocol type, and session duration, can reveal the volume of information flowing within the network and its dynamic trends. Monitoring network traffic not only reveals unusual traffic patterns but also identifies potential threat signals. For example, certain types of attacks may cause traffic surges or unusual protocol usage. Because network flow data is highly real-time, its collection must ensure prompt response and documentation of any suspicious changes. This requires the deployment of efficient collection tools and techniques that can accomplish this task without impacting network performance.
[0037] Combining these two types of data can provide a more complete picture of cybersecurity situational awareness and prediction. During implementation, all necessary data sources should be identified and the characteristics and structure of each data source should be thoroughly analyzed. It is also crucial to fully consider privacy protection and compliance requirements when designing data collection strategies. In particular, when processing security logs involving sensitive personal information, relevant laws and regulations must be followed to ensure the legal use of the data. This means implementing appropriate technical measures during the collection process, such as data encryption and anonymization, to prevent unauthorized access and personal information leakage.
[0038] In the above-mentioned large-model-based network security situation awareness prediction method, the step S2 performs data cleaning, format parsing and key field extraction on the security log data to obtain a structured log event stream. It should be understood that, considering that there are usually noise data (such as duplicate records, invalid fields), heterogeneous formats (such as Syslog, JSON, CSV) and redundant information in the original security log, directly inputting the information into the large model for information processing will lead to feature learning bias and low computational efficiency. Therefore, in order to construct a standardized, high-information-density log event data input, this application is based on regular expression matching and pattern recognition algorithms to remove duplicates, filter outliers and unify the format of the security log data, and extract key fields that are strongly related to the attack stage (such as high-risk operation instructions, abnormal login behavior) to form a structured log event stream. Specifically, first, by constructing a security event keyword library (such as "privilege escalation", "abnormal file access", etc.), a regular expression matching algorithm is used to quickly identify and extract key information fields that are closely related to security events. Subsequently, based on pattern recognition algorithms, key fields with core value are extracted, such as timestamp, event source IP / MAC address, destination IP / MAC address, port number, protocol type, event ID, operation type, user account, and power system-specific control instructions or status information. Ultimately, this information is organized into a unified structured data stream (for example, a JSON object stream or data table). In this way, the quality and availability of log data can be significantly improved, the obstacles caused by format differences can be eliminated, and subsequent data processing can be more focused on effective information, thereby improving the accuracy and efficiency of situational awareness prediction.
[0039] In the above-mentioned large-model-based network security situation awareness prediction method, the step S3 performs modal characterization on the structured log event stream to obtain a log event modal semantic feature encoding vector. In a specific example of the present application, the step S3 includes: performing modal semantic encoding based on the BERT model on the structured log event stream to obtain the log event modal semantic feature encoding vector. It should be understood that, considering that the traditional bag-of-words model or TF-IDF method cannot capture the context-dependent semantic relationship in the log event (for example, the combination of "file deletion" and "system backup failure" at a specific time sequence may imply a data destruction attack). Therefore, in order to deeply extract the contextual semantic representation of the structured log event stream, the present application is based on the pre-trained BERT model, and fine-tunes it to adapt to the semantic understanding of security field text to perform modal characterization on the structured log event stream. Specifically, the BERT model (Bidirectional Encoder Representations from Transformers) is a deep bidirectional pre-trained language model that can capture complex dependencies between words and contextual semantic information in the text. In this application, the BERT model is applied to security log event data. By taking the structured log event stream as the input text sequence and utilizing BERT's multi-layer self-attention mechanism to deeply encode each word unit, it is possible to effectively capture the semantic associations between entities, actions, and objects in the event description, and generate a fixed-dimensional log event modal semantic feature encoding vector through a pooling layer, thereby converting discrete log events into continuous semantic vector representations that contain attack intent, providing highly discriminative feature representations for subsequent network situational awareness.
[0040] In the above-mentioned network security situation awareness prediction method based on a large model, the step S4 extracts network traffic time series features from the network traffic data to obtain a network traffic modal time series pattern feature encoding vector. In a specific example of the present application, the step S4 includes: extracting network traffic time series features based on a Bi-LSTM model on the network traffic data to obtain the network traffic modal time series pattern feature encoding vector. It should be understood that since network traffic has dynamic time series dependencies (such as the traffic surge of a DDoS attack is usually accompanied by periodic fluctuations of a specific protocol port), traditional statistical features (such as mean and variance) are difficult to characterize hidden anomalies in long-range time series patterns. Therefore, in order to effectively capture the time series dynamics and potential attack patterns of traffic data, the present application uses a bidirectional long short-term memory network (Bi-LSTM) to perform time series analysis on the network traffic data, and simultaneously learns the forward and backward dependencies of the traffic data through bidirectional time series modeling to fully understand the time series evolution laws in the traffic data. Specifically, the Bi-LSTM model introduces two independent LSTM layers to process the input sequences of forward and reverse time series respectively, so that the model can capture the historical information before the current time step while also taking into account the potential impact of future time steps. In this application, network traffic data (such as traffic size, protocol type, session duration and other time series data) is input into the Bi-LSTM model as a multidimensional time series signal. Each LSTM unit in the model is responsible for processing the multidimensional joint features of a time step in the input sequence, and through mechanisms such as forget gates, input gates and output gates, the information of each time step in the sequence is selectively memorized and forgotten, thereby capturing the time series evolution laws of multidimensional features such as traffic size, protocol distribution, and session frequency, and identifying attack signals such as periodic anomalies and sudden fluctuations in traffic, such as the continuous growth characteristics of encrypted traffic in the data transmission stage. Finally, the extracted network traffic data time series features are mapped into a fixed-dimensional network traffic modal time series pattern feature encoding vector through a fully connected layer, which is used to characterize the time series evolution laws and potential attack patterns in the traffic data, providing strong traffic information support for subsequent situational awareness.
[0041] In the above-mentioned large-scale model-based network security situational awareness prediction method, in step S5, the network traffic modal temporal pattern feature encoding vector and the log event modal semantic feature encoding vector are input into a multimodal large-scale model fine-tuned for the security field to obtain a security situational awareness result, which is a predicted attack stage label. Figure 3 FIG. 5 is a flowchart of sub-step S5 of the network security situation awareness prediction method based on a large model according to an embodiment of the present application. Figure 3As shown, the step S5 includes the steps of: S51, performing semantic alignment processing on the network traffic modal time series pattern feature coding vector and the log event modal semantic feature coding vector to obtain an aligned network traffic modal time series pattern feature coding vector and an aligned log event modal semantic feature coding vector; S52, performing cross-modal feature semantic granularity transfer reasoning interaction on the aligned network traffic modal time series pattern feature coding vector and the aligned log event modal semantic feature coding vector to obtain a cross-modal interaction response reasoning coding vector; S53, inputting the cross-modal interaction response reasoning coding vector into the multimodal large model fine-tuned for the security field to obtain the security situation awareness result.
[0042] Specifically, the step S51 performs semantic alignment processing on the network traffic modal time series pattern feature coding vector and the log event modal semantic feature coding vector to obtain an aligned network traffic modal time series pattern feature coding vector and an aligned log event modal semantic feature coding vector. Specifically, considering that security logs and network traffic data are two heterogeneous modal data, there are differences between the two in data structure, feature space and semantic expression, and direct fusion may lead to information loss or misleading. Therefore, the present application further performs semantic alignment processing on the network traffic modal time series pattern feature coding vector and the log event modal semantic feature coding vector to construct a basis for cross-modal feature fusion. Among them, Figure 4 FIG is a flowchart of sub-step S51 of the network security situation awareness prediction method based on a large model according to an embodiment of the present application. Figure 4 As shown, the step S51 includes the steps of: S511, calculating the eigenvalue granularity global correlation matrix between the network traffic modal time series pattern feature coding vector and the log event modal semantic feature coding vector to obtain a cross-modal eigenvalue granularity global correlation matrix; S512, performing convolution coding and nonlinear activation on the cross-modal eigenvalue granularity global correlation matrix to obtain a cross-modal co-space semantic alignment coding matrix; S513, mapping the network traffic modal time series pattern feature coding vector and the log event modal semantic feature coding vector to the cross-modal co-space semantic alignment coding matrix respectively to obtain the aligned network traffic modal time series pattern feature coding vector and the aligned log event modal semantic feature coding vector.
[0043] More specifically, the step S511 calculates the eigenvalue granularity global correlation matrix between the network traffic modal temporal pattern feature coding vector and the log event modal semantic feature coding vector to obtain a cross-modal eigenvalue granularity global correlation matrix. Specifically, since the discrete events (semantic information) recorded by the security log and the continuous behavior (temporal pattern) reflected by the network traffic are two different aspects describing the same network security situation, there must be a potential correlation between the two (for example, an abnormal login log may correspond to an abnormal traffic connection). Based on this, in order to utilize the feature correlation between the network traffic modal temporal pattern feature coding vector and the log event modal semantic feature coding vector to align the feature space of the two, and then realize the effective fusion of cross-modal features, the present application further calculates the cross-modal eigenvalue granularity global correlation matrix between the network traffic modal temporal pattern feature coding vector and the log event modal semantic feature coding vector through matrix multiplication operations to reveal the potential correlation pattern and mutual dependence strength between log events and network traffic data, thereby providing key information guidance for subsequent cross-modal features.
[0044] More specifically, in step S512, the cross-modal feature value granularity global correlation matrix is convolutionally encoded and nonlinearly activated to obtain a cross-modal co-space semantic alignment coding matrix. It should be understood that the present application takes into account that the cross-modal feature value granularity global correlation matrix may contain some redundant associations (such as the random correlation between normal traffic fluctuations and irrelevant log events), and directly using it for alignment will lead to feature mapping deviation. Therefore, in order to extract significant patterns in cross-modal associations and establish a unified semantic space, the present application uses a convolutional neural network (CNN) to extract local patterns from the cross-modal feature value granularity global correlation matrix, for example, using multi-scale convolution kernels (such as 3×3, 5×5) to capture cross-modal association patterns of different granularities (such as local matching of short-term traffic anomalies and single high-risk logs, global association of long-term traffic trends and log event sequences), and enhances the nonlinear expression ability through the ReLU activation function, and finally outputs the cross-modal co-space semantic alignment coding matrix. In this way, the model can filter out noise interference and retain cross-modal correlation features driven by attack behavior, such as the strong correlation between the persistent high level of encrypted traffic and frequent file access logs during the data theft phase.
[0045] More specifically, in step S513, the network traffic modal temporal pattern feature coding vector and the log event modal semantic feature coding vector are respectively mapped to the cross-modal co-space semantic alignment coding matrix to obtain the aligned network traffic modal temporal pattern feature coding vector and the aligned log event modal semantic feature coding vector. That is, based on the cross-modal co-space semantic alignment coding matrix, the network traffic modal time series pattern feature coding vector and the log event modal semantic feature coding vector are linearly transformed respectively, that is, the network traffic modal time series pattern feature coding vector and the cross-modal co-space semantic alignment coding matrix are projected through matrix multiplication operation to generate an aligned network traffic modal time series pattern feature coding vector; similarly, the same operation is performed on the log event modal semantic feature coding vector to obtain an aligned log event modal semantic feature coding vector, so that the network traffic modal time series pattern feature coding vector and the log event modal semantic feature coding vector are mapped to a unified semantic space, so that the "periodic anomaly" in the network traffic time series pattern and the "authority escalation operation" in the log event semantics form an explainable correspondence on the same dimension, providing a feature alignment basis for subsequent cross-modal data interaction and situational awareness, and promoting the cross-modal feature interaction effect.
[0046] Specifically, step S52 performs cross-modal feature semantic granularity transfer reasoning interaction on the aligned network traffic modal temporal pattern feature encoding vector and the aligned log event modal semantic feature encoding vector to obtain a cross-modal interaction response reasoning encoding vector. It should be understood that aligning network traffic temporal pattern features and log event semantic features into the same semantic space only solves the basics of information "communication" but does not fully explore the deep-level, causal interaction relationship between the two. Complex network attacks (such as APTs) often manifest as the co-evolution between log events and traffic patterns. Simple feature splicing or weighted summation cannot model the complex nonlinear interactions between cross-modal features. Therefore, in order to deeply explore the collaborative behavior pattern of multimodal features, this application further performs fine-grained feature interaction transfer reasoning and global context transfer aggregation on the aligned network traffic modal time series pattern feature encoding vector and the aligned log event modal semantic feature encoding vector to mine the multimodal collaborative signals in the evolution of the attack stage, such as the spatiotemporal coupling of low-frequency heartbeat traffic and scheduled task logs in the command and control (C2) stage, thereby generating a cross-modal interactive response reasoning encoding vector as the core input of the subsequent attack behavior prediction and situational awareness model. Among them, Figure 5 FIG is a flowchart of sub-step S52 of the network security situation awareness prediction method based on a large model according to an embodiment of the present application. Figure 5As shown, the step S52 includes the steps of: S521, performing feature transfer response reasoning based on local granularity on the aligned network traffic modal time series pattern feature coding vector and the aligned log event modal semantic feature coding vector to obtain a sequence of network traffic-log event cross-modal feature local transfer response coding matrices; S522, performing transfer response reasoning sequence transfer coding on the sequence of network traffic-log event cross-modal feature local transfer response coding matrices to obtain the cross-modal interaction response reasoning coding vector.
[0047] More specifically, in a specific example of the present application, step S521 includes: first, performing an ordered arrangement based on the eigenvalue size on the aligned network traffic modal temporal pattern feature coding vector and the aligned log event modal semantic feature coding vector to obtain an aligned network traffic modal temporal pattern feature ordered arrangement coding vector and an aligned log event modal semantic feature ordered arrangement coding vector, which is expressed as follows:
[0048] V1'=sort(V1)
[0049] V2'=sort(V2)
[0050] Among them, V1 represents the feature encoding vector of the aligned network traffic modal temporal pattern, V2 represents the feature encoding vector of the aligned log event modal semantic feature, sort(·) represents the sorting operation on the vector elements, V1 ‘ Represents the ordered arrangement encoding vector of the aligned network traffic modal temporal pattern features, V2 ‘ Represents the ordered arrangement encoding vector of the aligned log event modal semantic features.
[0051] That is, by sequentially arranging the feature encoding vectors of the aligned network traffic modal temporal patterns and the semantic feature encoding vectors of the aligned log event modalities, a standardized representation is constructed that is independent of the order of feature arrangement. This allows the model to strip away the interference factors caused by the order of input feature positions during processing and instead focus on the inherent distribution patterns and relative strength relationships of the feature values. This provides a unified and stable feature foundation for subsequent cross-modal deep interactive reasoning, adapting to the processing requirements of large multimodal models for structured inputs and enhancing the comparability between features. Through this sequential arrangement processing, the model is endowed with the ability to perceive the invariance of the input feature arrangement, avoiding analytical bias caused by differences in feature positions.
[0052] Then, the aligned network traffic modal temporal pattern feature ordered arrangement coding vector and the aligned log event modal semantic feature ordered arrangement coding vector are subjected to equal granularity feature segmentation to obtain a sequence of ordered coding vectors of network traffic modal local temporal feature and a sequence of ordered coding vectors of log event modal local semantic feature, which can be expressed as follows:
[0053] Split(V1')={x1,x2,...,x i ,...,x n}
[0054] Split(V2')={y1,y2,...,y i ,...,y n}
[0055] Among them, x1, x2, x i and x n They represent the first, second, i-th and n-th ordered coding vectors of local temporal sequence characteristics of network traffic modalities in the sequence of ordered coding vectors of local temporal sequence characteristics of network traffic modalities, respectively. n is the number of ordered coding vectors of local temporal sequence characteristics of network traffic modalities, y1, y2, y i and y n where represents the first, second, i-th, and n-th ordered encoding vectors of the modal local semantic features of the log event, respectively. Split(·) represents the feature segmentation function.
[0056] That is, by synchronously dividing the ordered arrangement coding vectors of the aligned network traffic modal temporal pattern features and the ordered arrangement coding vectors of the aligned log event modal semantic features into continuous sub-vector segments of equal granularity along the dimensions, the network traffic temporal features and the log semantic features form an alignable fine-grained feature group within the same dimensional interval, generating a sequence of ordered coding vectors of the local temporal features of the network traffic modal and a sequence of ordered coding vectors of the local semantic features of the log event modal, providing a structured input basis for the subsequent precise interaction and collaborative reasoning of cross-modal local features.
[0057] Finally, each corresponding group of ordered coding vectors of local temporal features of network traffic modality and ordered coding vectors of local semantic features of log event modality in the sequence of ordered coding vectors of local temporal features of network traffic modality and the sequence of ordered coding vectors of local semantic features of log event modality are input into the transfer response inference unit to obtain the sequence of local transfer response coding matrices of network traffic-log event cross-modal features, which can be expressed as follows:
[0058]
[0059] Among them, W i represents the linear transformation matrix, ReLU(·) represents the ReLU activation function, Represents matrix multiplication, M i Represents x i with y iThe network traffic-log event cross-modal feature local transfer response encoding matrix.
[0060] Specifically, by leveraging the powerful fitting capabilities of deep learning, we deeply model the complex, nonlinear interactions between the local temporal features of network traffic modalities and the local semantic features of log event modalities within corresponding local vectors. This quantifies and encodes interaction patterns such as local similarities, differences, alignments, conditional dependencies, and co-activation or inhibition between the two types of features within a specific feature value range. This encapsulates rich and structured interaction information within the corresponding local regions, preserving more interaction details. The resulting sequence of local transfer response encoding matrices for network traffic-log event cross-modal features more accurately captures the strength of the element-by-element correlation between network traffic and log events, rather than simply using scalar scores. This allows the model to deeply explore potential connections between multimodal data.
[0061] Figure 6 Flowchart of sub-step S522 of the network security situation awareness prediction method based on a large model according to an embodiment of the present application. Figure 5 As shown, the step S522 includes the steps of: S5221, based on the migration connection probability density between each group of corresponding network traffic modal local temporal feature ordered coding vectors and log event modal local semantic feature ordered coding vectors, regularizing the individual network traffic-log event cross-modal feature local transfer response coding matrices in the sequence of the network traffic-log event cross-modal feature local transfer response coding matrices to obtain a sequence of optimized network traffic-log event cross-modal feature local transfer response coding matrices; S5222, performing transfer coding based on the attention mechanism on the sequence of the optimized network traffic-log event cross-modal feature local transfer response coding matrices to obtain the cross-modal interaction response inference coding vector.
[0062] In a preferred example of the present application, the step S5221, based on the migration connection probability density between each group of corresponding network traffic modal local temporal feature ordered coding vectors and log event modal local semantic feature ordered coding vectors, performs regularization constraints on each network traffic-log event cross-modal feature local transfer response coding matrix in the sequence of network traffic-log event cross-modal feature local transfer response coding matrices to obtain a sequence of optimized network traffic-log event cross-modal feature local transfer response coding matrices.
[0063] Here, considering that the segmentation granularity of the feature segmentation operation affects the domain size of the network traffic-log event cross-modal feature local transfer response encoding matrix, it will also directly affect the expression of the interaction pattern between the corresponding aligned network traffic modal temporal pattern feature ordered arrangement encoding vector and the aligned log event modal semantic feature ordered arrangement encoding vector. Based on this, the present application further performs regularized constraint optimization on each network traffic-log event cross-modal feature local transfer response encoding matrix.
[0064] Specifically, since the network traffic-log event cross-modal feature local transfer response encoding matrix is a transfer response between two local domains, it actually uses its row vector as a benchmark to express the spatial measure of the transfer response space, and the row vector length is also the representation of the above-mentioned domain scale size. If the domain scale size, that is, the row vector length L, is introduced as the domain scale strength constraint, then the dimensionality reduction space representation (that is, the F norm ‖M i ‖ F ), should follow a Poisson-like relationship:
[0065]
[0066] where (·)! represents factorial, ‖·‖ F represents the Frobenius norm, L represents M i The row vector length, e (·) represents an exponential function with a natural constant as the base, and λ represents the low-rank space measurement expression parameter.
[0067] That is, L acts as a domain scale strength constraint on the dimensionality reduction representation frequency of the local transfer response encoding matrix of the network traffic-log event cross-modal feature L times.
[0068] In this case, when the migration coupling is generated by a Poisson-like process with a coupling strength coefficient L and a mean expectation λ, i ‖2 and ‖y i ‖2 The boundary connection representation described by the spatial metric can further determine the migration connection probability density between the two local regions as:
[0069]
[0070] Where ρ represents x i and y i The probability density of migration connections between them, |·| represents the calculation of the modulus, and ‖·‖2 represents the calculation of the second norm.
[0071] Then, the local area size L is gradually optimized using the migration connection probability density:
[0072] L ′ =ρL
[0073] Among them, L ′ represents L after gradual optimization.
[0074] That is, under the condition of strictly ensuring the expected correlation degree is L, the regularization constraint of the probability density mean fluctuation of each row is used to determine the regularization of the global space metric, so that the structured interaction information in the local space metric can avoid the risk of local overfitting. Finally, with the optimized L ' Reconstrain the network traffic-log event cross-modal feature local transfer response encoding matrix M i Sparsity:
[0075]
[0076] Among them, M' i Indicates M i The corresponding optimized network traffic-log event cross-modal feature local transfer response encoding matrix.
[0077] That is, by enforcing the local transfer of network traffic-log event cross-modal features, the low-rank expression of the response encoding matrix and the local size L ' Matching,balances feature strength and sparsity to improve the overall expression effect of the sequence,of network traffic-log event cross-modal feature local transfer response encoding,matrix.
[0078] In a specific example of the present application, step S5222 includes: first, flattening each optimized network traffic-log event cross-modal feature local transfer response coding matrix in the sequence of the optimized network traffic-log event cross-modal feature local transfer response coding matrix to obtain a sequence of network traffic-log event cross-modal feature local transfer response coding vectors, which is expressed as follows:
[0079] vec(M′ i )=v i
[0080] Among them, vec(·) represents the matrix flattening operation, v i Represents M′ i Expand the obtained network traffic-log event cross-modal feature local transfer response encoding vector.
[0081] Specifically, the structured interaction information of the optimized network traffic-log event cross-modal feature local transfer response encoding matrix is converted into a linear sequence form, providing a unified input format for the subsequent integration of interaction information from different local regions and capturing their sequential dependencies, facilitating the model's learning of the evolution of interaction patterns across feature intervals. The generated sequence of network traffic-log event cross-modal feature local transfer response encoding vectors retains the element-level interaction details in the optimized network traffic-log event cross-modal feature local transfer response encoding matrix, providing a data foundation for analyzing the sequential associations and global contextual dependencies of interaction patterns across different eigenvalue intervals, and enhancing the model's ability to capture the global characteristics of multimodal feature interactions.
[0082] Then, using the Frobenius norm of each optimized network traffic-log event cross-modal feature local transfer response coding matrix in the sequence of the optimized network traffic-log event cross-modal feature local transfer response coding matrix as a constraint factor, the sequence of the network traffic-log event cross-modal feature local transfer response coding vectors is feature modulated to obtain a sequence of modulated network traffic-log event cross-modal feature local transfer response coding vectors, which is expressed as follows:
[0083]
[0084] Among them, exp(·) represents the exponential function operation with e as the base, is the feature modulation function based on the attention mechanism.
[0085] That is, the Frobenius norm of the optimized network traffic-log event cross-modal feature local transfer response encoding matrix is used as a constraint factor to adjust the weights of the sequence of flattened network traffic-log event cross-modal feature local transfer response encoding vectors. By quantifying the overall strength of the interactive information in each local area, the contribution of the corresponding feature vector is dynamically adjusted to highlight the feature influence of high interaction intensity areas and suppress the interference of low-value information, providing feature input that is more focused on effective interactive information for the subsequent accurate prediction of attacks.
[0086] Finally, the sequence of the modulated network traffic-log event cross-modal feature local transfer response encoding vector is transferred and encoded based on the LSTM model to obtain the cross-modal interaction response inference encoding vector, which is expressed as follows:
[0087]
[0088] Among them, M1 represents the network traffic-log event cross-modal feature local transfer response encoding matrix between x1 and y1, M n Represents x n with y nThe network traffic-log event cross-modal feature local transfer response encoding matrix between them, LSTM(·) represents the LSTM model, v f Represents the cross-modal interaction response reasoning encoding vector.
[0089] That is, the long short-term memory network (LSTM) is used to deeply process the sequence of modulated network traffic-log event cross-modal feature local transfer response encoding vectors to capture the long-term dependencies and time series characteristics in the sequence, thereby comprehensively and compactly representing the overall and deep interactive response characteristics between the aligned network traffic modal temporal pattern feature encoding vector and the aligned log event modal semantic feature encoding vector, providing more discriminative input for subsequent security situation awareness and attack stage prediction.
[0090] Specifically, in step S53, the cross-modal interaction response inference encoding vector is input into the multimodal macro model fine-tuned for the security domain to obtain the security situation awareness result. It should be understood that since the general multimodal macro model lacks targeted modeling of network security domain knowledge, direct application will lead to attack phase prediction bias. Therefore, in order to improve the model's ability to discriminate the characteristics of the APT attack chain stage, this application is based on the pre-trained multimodal macro model and performs domain-adaptive fine-tuning using security domain data (such as ATT&CK tactical labels and attack chain samples). Specifically, during the fine-tuning phase, attack phase labels are used as supervisory signals, and a contrastive learning loss function is employed to narrow the distance between multimodal features within the same attack phase while simultaneously widening the distance between features from different phases. During the inference phase, the cross-modal interaction response inference encoding vector is input into the fine-tuned large model. The large model's terminal layer comprises one or more fully connected layers and a classification layer (such as a Softmax layer), which is responsible for mapping the input cross-modal interaction response features to a probability distribution within a predefined set of attack phase labels (e.g., phase labels consistent with the Kill Chain or MITRE ATT&CK framework: reconnaissance, weaponization, delivery, exploitation, installation, command and control (C2), and target achievement). The label with the highest probability is then output as the prediction result. This approach fully leverages the advantages of the large model in processing complex inputs and performing high-level reasoning, achieving high-precision perception of the network security situation and accurate judgment of the specific attack phase. This provides timely and accurate early warning information for critical infrastructure such as power systems, thereby supporting more effective security decision-making and proactive defense strategies.
[0091] In summary, a large-scale model-based network security situational awareness prediction method based on the embodiment of the present application is illustrated, which first cleans and parses the network security log and extracts key fields to generate a structured event stream, and then performs modal characterization processing to extract its deep semantic feature representation. At the same time, the network traffic data is subjected to time series modeling to learn the time series feature pattern of the network traffic. Then, cross-modal semantic alignment technology is used to map the semantic features of the network security log and the time series features of the network traffic to a unified semantic space to achieve feature dimension alignment, and through fine-grained feature transfer interactive reasoning on the two, the potential correlation and collaborative behavior pattern between network traffic and log events are deeply mined, so as to predict the specific stage of the attack behavior based on cross-modal complementary information and a multi-modal large model that has been fine-tuned by the domain. This method can accurately capture the trajectory of attack behavior and improve the accuracy and foresight of situational awareness by making full use of the complementary information between heterogeneous data.
[0092] Furthermore, a network security situation awareness and prediction system based on a large model is also provided.
[0093] Figure 7 FIG is a block diagram of a network security situation awareness prediction system based on a large model according to an embodiment of the present application. Figure 7 As shown, according to an embodiment of the present application, a network security situation awareness prediction system 100 based on a large model includes: a data acquisition module 110, used to acquire security log data and network traffic data; a security log data preprocessing module 120, used to perform data cleaning, format parsing and key field extraction on the security log data to obtain a structured log event stream; a modal characterization module 130, used to perform modal characterization on the structured log event stream to obtain a log event modal semantic feature encoding vector; a time series feature extraction module 140, used to extract network traffic time series features from the network traffic data to obtain a network traffic modal time series pattern feature encoding vector; a security situation awareness module 150, used to input the network traffic modal time series pattern feature encoding vector and the log event modal semantic feature encoding vector into a multimodal large model fine-tuned for the security field to obtain a security situation awareness result, wherein the security situation awareness result is a predicted attack stage label.
[0094] Here, those skilled in the art will understand that the specific operations of each module in the above-mentioned large model-based network security situation awareness prediction system have been referred to above. Figures 1 to 6 The description of the large model-based cybersecurity situation awareness prediction method has been introduced in detail, and therefore, its repeated description will be omitted.
[0095] The basic principles of the present invention have been described above in conjunction with specific embodiments. However, it should be noted that the advantages, strengths, and effects mentioned in the present invention are merely illustrative and non-limiting, and should not be construed as necessarily possessed by each embodiment of the present invention. Furthermore, the specific details of the above embodiments are provided for illustrative purposes and to facilitate understanding, and are not intended to be limiting. These details do not necessarily limit the present invention to being implemented using these specific details.
[0096] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described or recorded in detail in a certain embodiment, please refer to the relevant description of other embodiments. In the several embodiments provided by the present invention, it should be understood that the disclosed system and method can be implemented in other ways. For example, the system embodiment described above is only schematic. For example, the unit division is only a logical function division, and there may be other division methods in actual implementation. The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the scheme of this embodiment.
[0097] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above and that the invention can be embodied in other specific forms without departing from the spirit or essential characteristics of the invention. Therefore, the embodiments should be considered in all respects as illustrative and non-restrictive, and the scope of the invention is defined by the appended claims, not the foregoing description, and all variations within the meaning and range of equivalents of the claims are intended to be encompassed therein. Any reference to a figure in a claim should not be construed as limiting the claim to which it relates.
[0098] In addition, it is obvious that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units stated in the system claims can also be implemented by one unit through software or hardware.
[0099] Finally, it should be noted that the above description has been provided for the purpose of illustration and description. In addition, the above embodiments are intended only to illustrate the technical solutions of the present invention and are not intended to be limiting. Although the technical solutions may be modified or replaced with equivalents with reference to the preferred embodiments, they do not depart from the spirit and scope of the technical solutions of the present invention.
Claims
1. A network security situation awareness prediction method based on a large model, characterized by: include: Obtain security log data and network traffic data; Performing data cleaning, format parsing, and key field extraction on the security log data to obtain a structured log event stream; Performing modal characterization on the structured log event stream to obtain a log event modal semantic feature encoding vector; Extracting network traffic time series features from the network traffic data to obtain a network traffic modal time series pattern feature encoding vector; The network traffic modal time series pattern feature encoding vector and the log event modal semantic feature encoding vector are input into a multimodal large model fine-tuned for the security field to obtain a security situation awareness result, which is a predicted attack stage label.
2. The network security situation awareness prediction system and method based on a large model according to claim 1 is characterized in that: Performing modal characterization on the structured log event stream to obtain a log event modal semantic feature encoding vector includes: The structured log event stream is subjected to modal semantic encoding based on the BERT model to obtain the log event modal semantic feature encoding vector.
3. The network security situation awareness prediction method based on a large model according to claim 1 is characterized in that: Extracting network traffic time series features from the network traffic data to obtain a network traffic modal time series pattern feature encoding vector includes: The network traffic data is subjected to network traffic time series feature extraction based on the Bi-LSTM model to obtain the network traffic modal time series pattern feature encoding vector.
4. The network security situation awareness prediction method based on a large model according to claim 1 is characterized in that: Inputting the network traffic modal temporal pattern feature encoding vector and the log event modal semantic feature encoding vector into a multimodal large model fine-tuned for the security field to obtain a security situation awareness result, including: Performing semantic alignment processing on the network traffic modal time series pattern feature coding vector and the log event modal semantic feature coding vector to obtain an aligned network traffic modal time series pattern feature coding vector and an aligned log event modal semantic feature coding vector; Performing cross-modal feature semantic granularity transfer reasoning interaction on the aligned network traffic modal time series pattern feature encoding vector and the aligned log event modal semantic feature encoding vector to obtain a cross-modal interaction response reasoning encoding vector; The cross-modal interactive response reasoning encoding vector is input into the multimodal large model fine-tuned for the security domain to obtain the security situation awareness result.
5. The network security situation awareness prediction method based on a large model according to claim 4 is characterized in that: Performing semantic alignment processing on the network traffic modal time series pattern feature coding vector and the log event modal semantic feature coding vector to obtain an aligned network traffic modal time series pattern feature coding vector and an aligned log event modal semantic feature coding vector, including: Calculating a eigenvalue granularity global correlation matrix between the network traffic modal time series pattern feature encoding vector and the log event modal semantic feature encoding vector to obtain a cross-modal eigenvalue granularity global correlation matrix; Performing convolutional coding and nonlinear activation on the cross-modal feature value granularity global correlation matrix to obtain a cross-modal co-spatial semantic alignment coding matrix; The network traffic modal temporal pattern feature coding vector and the log event modal semantic feature coding vector are respectively mapped to the cross-modal co-space semantic alignment coding matrix to obtain the aligned network traffic modal temporal pattern feature coding vector and the aligned log event modal semantic feature coding vector.
6. The network security situation awareness prediction method based on a large model according to claim 5 is characterized in that: Performing cross-modal feature semantic granularity transfer reasoning interaction on the aligned network traffic modal time series pattern feature encoding vector and the aligned log event modal semantic feature encoding vector to obtain a cross-modal interaction response reasoning encoding vector, including: Performing local granularity-based feature transfer response reasoning on the aligned network traffic modal temporal pattern feature encoding vector and the aligned log event modal semantic feature encoding vector to obtain a sequence of network traffic-log event cross-modal feature local transfer response encoding matrices; The sequence of the network traffic-log event cross-modal feature local transfer response coding matrix is subjected to transfer response reasoning sequence transfer coding to obtain the cross-modal interaction response reasoning coding vector.
7. The network security situation awareness prediction method based on a large model according to claim 6 is characterized in that: Performing local granularity-based feature transfer response reasoning on the aligned network traffic modal temporal pattern feature encoding vector and the aligned log event modal semantic feature encoding vector to obtain a sequence of network traffic-log event cross-modal feature local transfer response encoding matrices, including: The aligned network traffic modal time series pattern feature coding vector and the aligned log event modal semantic feature coding vector are sequentially arranged based on the eigenvalue size to obtain an aligned network traffic modal time series pattern feature sequentially arranged coding vector and an aligned log event modal semantic feature sequentially arranged coding vector; Performing equal-granularity feature segmentation on the aligned network traffic modal temporal pattern feature ordered arrangement coding vector and the aligned log event modal semantic feature ordered arrangement coding vector to obtain a sequence of network traffic modal local temporal feature ordered coding vectors and a sequence of log event modal local semantic feature ordered coding vectors; Each group of corresponding ordered coding vectors of local temporal features of the network traffic modalities and ordered coding vectors of local semantic features of the log event modalities in the sequence of ordered coding vectors of local temporal features of the network traffic modalities and the sequence of ordered coding vectors of local semantic features of the log event modalities are input into the transfer response inference unit to obtain a sequence of local transfer response coding matrices of the network traffic-log event cross-modal features.
8. The network security situation awareness prediction method based on a large model according to claim 7 is characterized in that: Performing transfer response inference sequence transfer encoding on the sequence of the network traffic-log event cross-modal feature local transfer response encoding matrix to obtain the cross-modal interaction response inference encoding vector, including: Based on the migration connection probability density between each group of corresponding ordered coding vectors of local temporal features of network traffic modality and ordered coding vectors of local semantic features of log event modality, regularization constraints are performed on each network traffic-log event cross-modal feature local transfer response coding matrix in the sequence of network traffic-log event cross-modal feature local transfer response coding matrices to obtain a sequence of optimized network traffic-log event cross-modal feature local transfer response coding matrices; The sequence of the optimized network traffic-log event cross-modal feature local transfer response encoding matrix is subjected to transfer encoding based on the attention mechanism to obtain the cross-modal interaction response inference encoding vector.
9. The network security situation awareness prediction method based on a large model according to claim 8 is characterized in that: Performing transfer encoding based on an attention mechanism on the sequence of the optimized network traffic-log event cross-modal feature local transfer response encoding matrix to obtain the cross-modal interaction response inference encoding vector, including: Flattening each optimized network traffic-log event cross-modal feature local transfer response coding matrix in the sequence of optimized network traffic-log event cross-modal feature local transfer response coding matrices to obtain a sequence of network traffic-log event cross-modal feature local transfer response coding vectors; Using the Frobenius norm of each optimized network traffic-log event cross-modal feature local transfer response coding matrix in the sequence of the optimized network traffic-log event cross-modal feature local transfer response coding matrix as a constraint factor, feature modulating the sequence of the network traffic-log event cross-modal feature local transfer response coding vectors to obtain a sequence of modulated network traffic-log event cross-modal feature local transfer response coding vectors; The sequence of the modulated network traffic-log event cross-modal feature local transfer response encoding vector is subjected to transfer encoding based on the LSTM model to obtain the cross-modal interaction response inference encoding vector.
10. A network security situation awareness and prediction system based on a large model, characterized by: include: Data acquisition module, used to obtain security log data and network traffic data; A security log data preprocessing module is used to perform data cleaning, format parsing and key field extraction on the security log data to obtain a structured log event stream; A modality characterization module, configured to perform modality characterization on the structured log event stream to obtain a log event modality semantic feature encoding vector; A time series feature extraction module, configured to extract network traffic time series features from the network traffic data to obtain a network traffic modal time series pattern feature encoding vector; The security situation awareness module is used to input the network traffic modal time series pattern feature encoding vector and the log event modal semantic feature encoding vector into a multimodal large model fine-tuned for the security field to obtain a security situation awareness result, wherein the security situation awareness result is a predicted attack stage label.
Citation Information
Cited By
Cantonese pronunciation training system based on speech recognition feedback
CN120319226A
Wireless network dynamic security defense method and device based on risk evolution reasoning
CN122054149A
Method and device for dynamic security defense of wireless network based on risk evolution reasoning
CN122054149B