Network security situation awareness method and system based on multi-mode large model
Through the network security situation awareness method based on multimodal large model, the problem of insufficient ability to analyze complex network attacks and multi-source data in the prior art is solved, and more efficient and accurate network security situation awareness and automated response are achieved.
Patent Information
- Application Number
- CN202411889010.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-20
- Publication Date
- 2025-05-16
AI Technical Summary
Existing cybersecurity situation awareness technologies rely on a single data source, static rules and manual participation, making it difficult to cope with complex cyber attacks and multi-source data analysis, resulting in slow response, high false positive rates and limited analysis capabilities.
A multimodal deep situational awareness model is adopted based on multimodal large model, and a multimodal depth situational awareness and prediction model is constructed through orthogonal sequence fusion, teacher-student model, fast gradient symbol method, deep balanced multimodal fusion technology and bidirectional encoder structure to realize comprehensive analysis and automated response of multi-source data.
It improves the accuracy and efficiency of network security situation awareness, reduces the false alarm and missed alarm rates, enhances the ability to identify complex attack patterns, and improves the intelligence level of network security protection.
Smart Images

Figure CN120017301A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of network security and artificial intelligence, and in particular to a network security situation awareness method and system based on a multimodal large model. Background Art
[0002] With the rapid development of information technology, the network environment is becoming increasingly complex, and network security issues have become the focus of global attention. Traditional network security technologies often rely on a single data source and a static rule system, which makes it difficult to cope with the growing number of network attacks and hidden threat behaviors. In order to improve the level of intelligence in network security protection, it is necessary to adopt advanced technical means to conduct real-time monitoring and analysis of network security situations.
[0003] Network security situation awareness is one of the key technologies for maintaining cyberspace security. Its purpose is to promptly detect and respond to security threats by real-time monitoring and analysis of network activities.
[0004] Traditional network security situational awareness methods mainly rely on manually set security rules and thresholds, as well as analysis of single data sources. These methods often have slow responses, high false alarm rates, and limited analytical capabilities when faced with new attack methods and complex network environments. With the continuous evolution and upgrading of network attack methods, traditional security protection measures are increasingly unable to meet the growing security needs.
[0005] Existing network security situation awareness technologies still have some limitations and challenges:
[0006] Most existing network security situation awareness systems rely on a single data source, such as network traffic or system logs, and lack the ability to comprehensively analyze multi-source data, which limits the identification and response to complex attack patterns.
[0007] Although some cybersecurity situational awareness systems can monitor network activities in real time, they are insufficient in terms of automated response. In many cases, the identification and response of security incidents still require human participation, resulting in delayed responses.
[0008] The system generates a high proportion of false positives and false negatives when processing large amounts of data, which not only increases the workload of security analysts but also ignores real security threats;
[0009] In addition, there is a lack of effective data fusion and analysis mechanisms in processing and fusing multimodal data.
[0010] In response to the above problems, this application proposes a network security situation awareness method and system based on a multimodal large model. Summary of the invention
[0011] The present invention proposes the following technical solutions to address one or more technical deficiencies in the above-mentioned prior art.
[0012] Based on the first aspect of the present application, a network security situation awareness method based on a multimodal large model is proposed, comprising:
[0013] S1: Define a unique orthogonal sequence S for the data source i of multimodal data i , the orthogonal sequence S i Logging with each data source i (t) Combine to generate a new coded data sequence E i (t), and the coded data sequence E i (t) is added to obtain a composite signal C(t), and the orthogonal sequence S i Performing a matching operation with the composite signal C(t) to obtain information of a single data source;
[0014] S2: extracting features and performing knowledge distillation on the multimodal data in the single data source through a teacher-student model;
[0015] S3: Generate adversarial samples of the original multimodal data of the single data source by using a fast gradient sign method and an iterative fast gradient sign method, perform cross-modal association analysis on the multimodal data, and perform fusion processing and cross-modal alignment processing on the extracted features;
[0016] S4: establishing an equalization system through deep equalization multimodal fusion technology, and calculating the optimal transmission path between the processed features of different modes and integrating the features in the equalization system by minimizing the cost function, the lower semi-continuous function, the joint probability distribution and the Sinkhorn algorithm;
[0017] S5: Based on the bidirectional encoder structure and integrated features, a multimodal deep situational awareness and prediction model is constructed, and the adversarial sample is used for training, and the residual period prediction technology is integrated into the multimodal deep situational awareness and prediction model, and the multimodal deep situational awareness and prediction model is used for network security situational awareness.
[0018] Furthermore, before S1, the step also includes acquiring the multimodal data in the network and performing data enhancement processing, including rotating, scaling and color transforming image data in the multimodal data; and,
[0019] Synonym replacement and sentence reorganization are performed on the text data in the multimodal data.
[0020] Furthermore, the cross-modal association analysis of the multimodal data specifically includes: extracting deep features from data of different modalities, capturing the interactive behavior of data between different modalities, constructing fine-grained association relationships between data of different modalities, and constructing the association relationship between hash codes and labels through a cross-modal retrieval and matching method based on hash learning and a latent semantic matrix of learning labels.
[0021] Constructing the association between hash codes and tags can improve the retrieval efficiency of cross-modal data;
[0022] Cross-modal correlation analysis of multimodal data can accurately capture the essential characteristics of network security threats and improve the accuracy and efficiency of threat detection.
[0023] Furthermore, the equalization system includes a recursive fusion layer, a modal feature injection layer and an equalization state layer. The recursive fusion layer calculates the output of this layer based on the result of the previous iteration, the modal feature injection layer combines the injected modal features to obtain residual fusion features, and the equalization state layer searches for the equalization state of the internal and cross-modal features of the equalization system through a black box solver.
[0024] The equalization system can thoroughly encode rich information inside and outside different modalities from low to high levels, realize effective downstream multimodal learning, and can be easily inserted into various multimodal frameworks.
[0025] Furthermore, the multimodal deep situation awareness and prediction model based on the bidirectional encoder structure and integrated features specifically includes:
[0026] S501: Initialize the parameters of the bidirectional encoder;
[0027] S502: Segmenting the text data in the multimodal data and converting it into a token sequence, and converting the non-text data into a text-compatible form through vector quantization or embedding technology;
[0028] S503: Input the processed multimodal data into the bidirectional encoder, capture the context information of each element in the token sequence through forward propagation and backward propagation, and fuse the features in the context information to obtain the context-related features of each token sequence;
[0029] S504: Add a multi-head attention mechanism to the bidirectional encoder to capture the dependencies between the context-related features, and assign attention weights according to the importance of the context features;
[0030] S505: interacting the features of data of different modalities through the equalization system, learning the correlation of cross-modal data, and learning the representation of shared features through the multi-layer structure of the bidirectional encoder;
[0031] S506: Add a periodic prediction module to the output of the bidirectional encoder to integrate the residual periodic prediction technology to obtain the multimodal depth perception and prediction model.
[0032] Building a multimodal deep situational awareness and prediction model can ensure the uniqueness and integrity of each modality during the feature fusion process.
[0033] Furthermore, the parameters of the bidirectional encoder are initialized including word embedding matrix, position encoding and other network parameters.
[0034] Furthermore, in response to the results of network security situation perception and prediction, preset security response measures are triggered, and the intelligent decision tree adjusts the security response measure strategy according to the real-time feedback of the network security situation.
[0035] Based on the second aspect of the present application, a network security situation awareness system based on a multimodal large model is proposed, comprising:
[0036] Orthogonal sequence fusion module: defines a unique orthogonal sequence S for the data source i of multimodal data i , the orthogonal sequence S i Logging with each data source i (t) Combine to generate a new coded data sequence E i (t), and the coded data sequence E i (t) is added to obtain a composite signal C(t), and the orthogonal sequence S i Performing a matching operation with the composite signal C(t) to obtain information of a single data source;
[0037] Initial processing module: extracting features and performing knowledge distillation on the multimodal data in the single data source through a teacher-student model;
[0038] Reprocessing module: generating adversarial samples of the original multimodal data of the single data source by fast gradient sign method and iterative fast gradient sign method, performing cross-modal association analysis on the multimodal data, and performing fusion processing and cross-modal alignment processing on the extracted features;
[0039] Integration module: Establish an equalization system through deep equalization multimodal fusion technology, and calculate the optimal transmission path between the processed features of different modes and perform feature integration in the equalization system by minimizing the cost function, the lower semi-continuous function, the joint probability distribution and the Sinkhorn algorithm;
[0040] Prediction module: Based on the bidirectional encoder structure and integrated features, a multimodal deep situational awareness and prediction model is constructed, and the adversarial sample is used for training, and the residual cycle prediction technology is integrated into the multimodal deep situational awareness and prediction model, and the multimodal deep situational awareness and prediction model is used for network security situational awareness.
[0041] Based on the third aspect of the present application, a computer program product is proposed, which has one or more computer programs thereon, and when the one or more computer programs are executed by a computer processor, any of the methods described above is implemented.
[0042] The technical effect of the present application is that the present application constructs a deep learning model capable of real-time network security situation analysis and future trend prediction by generating adversarial networks and optimal transmission Transformer fusion technology, and combining self-attention mechanism and time series analysis technology; and dynamically adjusts the preprocessing strategy according to data characteristics and changes in the network environment, optimizes the feature learning process by adversarial learning, integrates features of different modalities by cross-modal correlation analysis, improves the model's ability to identify and predict network security threats, can comprehensively capture and analyze network security threats, improve the accuracy of security threat identification, reduce the false alarm rate and missed alarm rate of network security threats, and enable security analysts to focus more effectively on dealing with real threats, thereby improving the intelligence level of network security protection, reducing security risks caused by human factors, enhancing the ability to identify complex attack patterns, and improving the adaptability and flexibility of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Other features, objects and advantages of the present application will become more apparent from the detailed description of non-limiting embodiments made with reference to the following drawings.
[0044] Figure 1 It is a flowchart of a network security situation awareness method based on a multimodal large model provided according to an embodiment of the present invention.
[0045] Figure 2 It is a flowchart of building a multimodal deep situational awareness and prediction model provided according to an embodiment of the present invention.
[0046] Figure 3 It is a framework diagram of a network security situation awareness system based on a multimodal large model provided according to an embodiment of the present invention.
[0047] Figure 4 A schematic diagram of the structure of a computer system suitable for implementing an electronic device of an embodiment of the present application is shown. DETAILED DESCRIPTION
[0048] The present application will be further described in detail below in conjunction with the accompanying drawings and embodiments. It is to be understood that the specific embodiments described herein are only used to explain the relevant invention, rather than to limit the invention. It should also be noted that, for ease of description, only the parts related to the relevant invention are shown in the accompanying drawings.
[0049] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0050] Figure 1 The present invention provides a flowchart of a network security situation awareness method based on a multimodal large model, including:
[0051] S1: Define a unique orthogonal sequence S for the data source i of multimodal data i , the orthogonal sequence S i Logging with each data source i (t) Combine to generate a new coded data sequence E i (t), and the coded data sequence E i (t) is added to obtain a composite signal C(t), and the orthogonal sequence S i Performing a matching operation with the composite signal C(t) to obtain information of a single data source;
[0052] S2: extracting features and performing knowledge distillation on the multimodal data in the single data source through a teacher-student model;
[0053] S3: Generate adversarial samples of the original multimodal data of the single data source by using a fast gradient sign method and an iterative fast gradient sign method, perform cross-modal association analysis on the multimodal data, and perform fusion processing and cross-modal alignment processing on the extracted features;
[0054] S4: establishing an equalization system through deep equalization multimodal fusion technology, and calculating the optimal transmission path between the processed features of different modes and integrating the features in the equalization system by minimizing the cost function, the lower semi-continuous function, the joint probability distribution and the Sinkhorn algorithm;
[0055] S5: Based on the bidirectional encoder structure and integrated features, a multimodal deep situational awareness and prediction model is constructed, and the adversarial sample is used for training, and the residual period prediction technology is integrated into the multimodal deep situational awareness and prediction model, and the multimodal deep situational awareness and prediction model is used for network security situational awareness.
[0056] It should be noted that, before S1, the step also includes obtaining the multimodal data in the network and performing data enhancement processing, including rotating, scaling and color transformation of image data in the multimodal data; and,
[0057] Synonym replacement and sentence reorganization are performed on the text data in the multimodal data.
[0058] It should be noted that the cross-modal association analysis of the multimodal data specifically includes: extracting deep features from data of different modalities, capturing the interactive behavior of data between different modalities, constructing fine-grained association relationships between data of different modalities, and constructing the association relationship between hash codes and labels through a cross-modal retrieval and matching method based on hash learning and a latent semantic matrix of learning labels.
[0059] It should be noted that the equalization system includes a recursive fusion layer, a modal feature injection layer and a balanced state layer. The recursive fusion layer calculates the output of this layer based on the result of the previous iteration. The modal feature injection layer obtains the residual fusion feature by combining the injected modal features. The balanced state layer searches for the balanced state of the internal and cross-modal features of the equalization system through a black box solver.
[0060] It should be noted that if Figure 2 As shown, the multimodal deep situation awareness and prediction model based on the bidirectional encoder structure and integrated features specifically includes:
[0061] S501: Initialize the parameters of the bidirectional encoder;
[0062] S502: Segmenting the text data in the multimodal data and converting it into a token sequence, and converting the non-text data into a text-compatible form through vector quantization or embedding technology;
[0063] S503: Input the processed multimodal data into the bidirectional encoder, capture the context information of each element in the token sequence through forward propagation and backward propagation, and fuse the features in the context information to obtain the context-related features of each token sequence;
[0064] S504: Add a multi-head attention mechanism to the bidirectional encoder to capture the dependencies between the context-related features, and assign attention weights according to the importance of the context features;
[0065] S505: interacting the features of data of different modalities through the equalization system, learning the correlation of cross-modal data, and learning the representation of shared features through the multi-layer structure of the bidirectional encoder;
[0066] S506: Add a periodic prediction module to the output of the bidirectional encoder to integrate the residual periodic prediction technology to obtain the multimodal depth perception and prediction model.
[0067] It should be noted that the parameters for initializing the bidirectional encoder include word embedding matrix, position encoding and other network parameters.
[0068] It should be noted that in response to the results of network security situation perception and prediction, the preset security response measures are triggered, and the intelligent decision tree adjusts the security response measure strategy according to the real-time feedback of the network security situation.
[0069] It should be noted that by aggregating multimodal data in the network environment through EDR\NDR and other devices, in order to cope with the concealment and complexity of network attack characteristics, this application introduces data enhancement technology to improve data diversity and the generalization ability of the model.
[0070] It should be noted that orthogonal sequence fusion (OSF) technology uses orthogonality to distinguish and process data streams from different sources. In this application, the log data generated by different security monitoring systems have different log formats. Orthogonal sequence fusion technology can efficiently perform cross-system analysis without losing important information, better manage and analyze multimodal data, and reduce mutual interference between data to a certain extent.
[0071] It should be noted that the knowledge distillation process refers to using the teacher model as a pre-trained large model, allowing the student model to imitate the soft labels of the teacher model, and then using the cross-entropy loss to compare the outputs of the two models, so as to transfer the knowledge of the complex teacher model to the smaller student model, so that the student model can maintain high accuracy and have lower computing resource requirements.
[0072] It should be noted that deep learning technology is combined with knowledge distillation methods to extract key features from multimodal data. For text data, natural language processing technology is used to extract semantic features. For image and video data, computer vision technology is used to extract visual features. For network traffic data, traffic patterns and behavioral characteristics are analyzed. Finally, through the teacher-student model structure, the knowledge of the expert system is distilled into the student model to enhance the model's ability to identify complex network threats.
[0073] It should be noted that the present application automatically discovers and utilizes the intrinsic connections between data of different modalities through cross-modal correlation analysis, and generates richer and more comprehensive feature representations, which can more accurately capture the essential characteristics of network security threats, and improve the expressiveness of features and the generalization of models through adversarial enhancement strategies, thereby improving the accuracy and efficiency of network security threat detection.
[0074] It should be noted that the fast gradient sign method (FGSM) and the iterative fast gradient sign method (I-FGSM) are effective mechanisms for attacking neural network models. They use gradient information to generate malicious samples to test the security of the model by calculating the gradient of the model loss function relative to the input data and using the sign of the gradient to determine the direction of the perturbation.
[0075] It should be noted that convolutional neural networks (CNN) and long short-term memory networks (LSTM) are used to extract deep features from data of different modalities, obtain feature representations with more semantic information, build a model based on deep learning, use the attention mechanism to capture subtle interactions between different modalities, construct fine-grained associations between modalities, and a cross-modal retrieval and matching method based on hash learning. By learning the latent semantic matrix of the labels, the association between hash codes and labels is constructed, which can improve retrieval efficiency.
[0076] It should be noted that the cross-modal alignment process includes attention alignment and semantic alignment, and the cross-modal alignment process is implicitly achieved by using the internal mechanisms of the Transformer structure and the BERT-based structure.
[0077] It should be noted that the use of deep equalization multimodal fusion technology to establish an equalization system to process the features of different modes can ensure that the features of each mode remain unique and complete during the fusion process.
[0078] It should be noted that the balanced system finds fixed points in the dynamic multimodal fusion process and models the correlation between features in an adaptive and recursive manner. It can thoroughly encode rich information inside and outside different modalities from low to high levels, realize effective downstream multimodal learning, and can be easily inserted into various multimodal frameworks.
[0079] It should be noted that the optimal transmission path is determined by minimizing the cost function, which is used to measure the cost of transmitting one probability measure to another probability measure. The present application finds the optimal transmission path by optimizing the joint probability distribution γ so that the marginal probability distribution of γ is equal to the probability measure of the modal characteristics respectively; the stability and effectiveness of the transmission path are ensured by the lower semi-continuous function; numerical methods, such as the Sinkhorn algorithm, are used to approximately solve the optimal transmission problem, improve computational efficiency and reduce computational cost.
[0080] It should be noted that the residual period prediction technology enhances the modeling ability of the model for long-term dependencies by learning the cycle period of time series data and predicting the periodic residual component. The cycle period and the backbone network train the model to reveal the internal cycle in the data;
[0081] The model adopts a bidirectional encoder structure to improve the understanding of time series data, while taking into account the past and future information of time series data, thus enhancing the model's predictive ability.
[0082] It should be noted that adding a multi-head attention mechanism to the bidirectional encoder is conducive to identifying key information of data of different modalities, capturing dependencies in different dimensions, and strengthening the learning of key features.
[0083] In a specific embodiment, multimodal data features that have undergone feature fusion are input into a multimodal deep situational awareness and prediction model, the self-attention layer of the model captures the dependencies between features, and position encoding is added to the model so that it can understand the position information of words in a sequence. A multi-head attention mechanism is used to process features of different subspaces in parallel to enhance the expressive power of the model. The normalization layer of the model stabilizes the training process, and a feedforward network is applied in the transformer layer of the model to further extract features and perform normalization to output the final prediction result.
[0084] Figure 3 The framework of a network security situation awareness system based on a multimodal large model is shown, including an orthogonal sequence fusion module a, a primary processing module b, a reprocessing module c, an integration module d and a prediction module e.
[0085] In a specific embodiment, the orthogonal sequence fusion module a is configured to: define a unique orthogonal sequence S for a data source i of the multimodal data: i , the orthogonal sequence S i Logging with each data source i (t) Combine to generate a new coded data sequence E i (t), and the coded data sequence E i (t) is added to obtain a composite signal C(t), and the orthogonal sequence S i A matching operation is performed with the composite signal C(t) to obtain information of a single data source.
[0086] In a specific embodiment, the initial processing module b is configured to: perform feature extraction and knowledge distillation on the multimodal data in the single data source through a teacher-student model.
[0087] In a specific embodiment, the reprocessing module c is configured to: generate adversarial samples of the original multimodal data of the single data source through the fast gradient sign method and the iterative fast gradient sign method, perform cross-modal correlation analysis on the multimodal data, and perform fusion processing and cross-modal alignment processing on the extracted features.
[0088] In a specific embodiment, the integration module d is configured to: establish an equalization system through deep equalization multimodal fusion technology, and in the equalization system, calculate the optimal transmission path between the processed features of different modes by minimizing the cost function, the lower semi-continuous function, the joint probability distribution and the Sinkhorn algorithm and perform feature integration.
[0089] In a specific embodiment, the prediction module e is configured to: construct a multimodal deep situational awareness and prediction model based on the bidirectional encoder structure and integrated features, and use the adversarial sample for training, integrate the residual cycle prediction technology into the multimodal deep situational awareness and prediction model, and use the multimodal deep situational awareness and prediction model to perform network security situation awareness.
[0090] Reference below Figure 4 , which shows a schematic diagram of the structure of a computer system suitable for implementing an electronic device of an embodiment of the present application. Figure 4 The electronic device shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.
[0091] like Figure 4 As shown, the computer system includes a central processing unit (CPU) 401, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 402 or a program loaded from a storage part 408 into a random access memory (RAM) 403. Various programs and data required for system operation are also stored in the RAM 403. The CPU 401, the ROM 402, and the RAM 403 are connected to each other via a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.
[0092] The following components are connected to the I / O interface 405: an input section 406 including a keyboard, a mouse, etc.; an output section 407 including a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 408 including a hard disk, etc.; and a communication section 409 including a network interface card such as a LAN card, a modem, etc. The communication section 409 performs communication processing via a network such as the Internet. A drive 410 is also connected to the I / O interface 405 as needed. A removable medium 411, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 410 as needed, so that a computer program read therefrom is installed into the storage section 408 as needed.
[0093] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable storage medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 409, and / or installed from the removable medium 411. When the computer program is executed by the central processing unit (CPU) 401, the above functions defined in the method of the present application are executed. It should be noted that the computer-readable storage medium of the present application can be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium can be, for example, - but not limited to - a system, device or device of electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to, an electrical connection with one or more conductors, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device, or device. In the present application, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which a computer-readable program code is carried. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable storage medium other than a computer-readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, device, or device. The program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to: wireless, wireline, optical cable, RF, etc., or any suitable combination of the foregoing.
[0094] Computer program code for performing the operations of the present application may be written in one or more programming languages or a combination thereof, including object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0095] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present application. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0096] The modules involved in the embodiments of the present application may be implemented by software or by hardware.
[0097] As another aspect, the present application further provides a computer-readable storage medium, which may be included in the electronic device described in the above embodiment; or may exist independently without being assembled into the electronic device. The above computer-readable storage medium carries one or more programs, and when the above one or more programs are executed by the electronic device, the electronic device: defines a unique orthogonal sequence S for the data source i of the multimodal data i , the orthogonal sequence S i Logging with each data source i (t) Combine to generate a new coded data sequence E i (t), and the coded data sequence Ei (t) is added to obtain a composite signal C(t), and the orthogonal sequence S i A matching operation is performed with the composite signal C(t) to obtain information of a single data source; feature extraction and knowledge distillation are performed on the multimodal data in the single data source through a teacher-student model; adversarial samples of the original multimodal data of the single data source are generated through a fast gradient sign method and an iterative fast gradient sign method, cross-modal correlation analysis is performed on the multimodal data, and the extracted features are fused and aligned across modalities; an equalization system is established through deep equalization multimodal fusion technology, and the optimal transmission path between the processed features of different modalities is calculated in the equalization system by minimizing the cost function, the lower semi-continuous function, the joint probability distribution and the Sinkhorn algorithm, and feature integration is performed; a multimodal deep situational awareness and prediction model is constructed based on a bidirectional encoder structure and integrated features, and the adversarial samples are used for training, the residual period prediction technology is integrated into the multimodal deep situational awareness and prediction model, and the multimodal deep situational awareness and prediction model is used for network security situational awareness.
[0098] Finally, it should be noted that the above embodiments are only intended to illustrate rather than limit the technical solutions of the present invention. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that the present invention can still be modified or replaced by equivalents. Any modification or partial replacement that does not depart from the spirit and scope of the present invention should be included in the scope of the claims of the present invention.
Claims
1. A network security situation awareness method based on a multimodal large model, characterized in that: include: S1: Define a unique orthogonal sequence S for the data source i of multimodal data i , the orthogonal sequence S i Logging with each data source i (t) Combine to generate a new coded data sequence E i (t), and the coded data sequence E i (t) is added to obtain a composite signal C(t), and the orthogonal sequence S i Performing a matching operation with the composite signal C(t) to obtain information of a single data source; S2: extracting features and performing knowledge distillation on the multimodal data in the single data source through a teacher-student model; S3: Generate adversarial samples of the original multimodal data of the single data source by using a fast gradient sign method and an iterative fast gradient sign method, perform cross-modal association analysis on the multimodal data, and perform fusion processing and cross-modal alignment processing on the extracted features; S4: establishing an equalization system through deep equalization multimodal fusion technology, in which the optimal transmission path between the processed features of different modes is calculated and the features are integrated by minimizing the cost function, the lower semi-continuous function, the joint probability distribution and the Sinkhorn algorithm; S5: Based on the bidirectional encoder structure and integrated features, a multimodal deep situational awareness and prediction model is constructed, and the adversarial sample is used for training, and the residual period prediction technology is integrated into the multimodal deep situational awareness and prediction model, and the multimodal deep situational awareness and prediction model is used for network security situational awareness.
2. The method according to claim 1, characterized in that The step S1 also includes obtaining the multimodal data in the network and performing data enhancement processing, including rotating, scaling and color transforming the image data in the multimodal data; and, Synonym replacement and sentence reorganization are performed on the text data in the multimodal data.
3. The method according to claim 1, characterized in that The cross-modal association analysis of the multimodal data specifically includes: extracting deep features from data of different modalities, capturing the interactive behavior of data between different modalities, constructing fine-grained association relationships between data of different modalities, and constructing the association relationship between hash codes and labels through a cross-modal retrieval and matching method based on hash learning and a latent semantic matrix of learning labels.
4. The method according to claim 1, characterized in that: The equalization system includes a recursive fusion layer, a modal feature injection layer and an equalization state layer. The recursive fusion layer calculates the output of this layer based on the result of the previous iteration. The modal feature injection layer obtains the residual fusion feature by combining the injected modal features. The equalization state layer searches for the equalization state of the internal and cross-modal features of the equalization system through a black box solver.
5. The method according to claim 1, characterized in that The multimodal deep situation awareness and prediction model based on the bidirectional encoder structure and integrated features specifically includes: S501: Initialize the parameters of the bidirectional encoder; S502: Segmenting the text data in the multimodal data and converting it into a token sequence, and converting the non-text data into a text-compatible form through vector quantization or embedding technology; S503: Input the processed multimodal data into the bidirectional encoder, capture the context information of each element in the token sequence through forward propagation and backward propagation, and fuse the features in the context information to obtain the context-related features of each token sequence; S504: Add a multi-head attention mechanism to the bidirectional encoder to capture the dependencies between the context-related features, and assign attention weights according to the importance of the context features; S505: interacting the features of data of different modalities through the equalization system, learning the correlation of cross-modal data, and learning the representation of shared features through the multi-layer structure of the bidirectional encoder; S506: Add a periodic prediction module to the output of the bidirectional encoder to integrate the residual periodic prediction technology to obtain the multimodal depth perception and prediction model.
6. The method according to claim 5, characterized in that The parameters of the bidirectional encoder are initialized including word embedding matrix, position encoding and other network parameters.
7. The method according to claim 1, characterized in that In response to the results of network security situation perception and prediction, the preset security response measures are triggered, and the intelligent decision tree adjusts the security response strategy based on real-time feedback of the network security situation.
8. A network security situation awareness system based on a multimodal large model, characterized in that: include: Orthogonal sequence fusion module: defines a unique orthogonal sequence S for the data source i of multimodal data i , the orthogonal sequence S i Logging with each data source i (t) Combine to generate a new coded data sequence E i (t), and the coded data sequence E i (t) is added to obtain a composite signal C(t), and the orthogonal sequence S i Performing a matching operation with the composite signal C(t) to obtain information of a single data source; Initial processing module: extracting features and performing knowledge distillation on the multimodal data in the single data source through a teacher-student model; Reprocessing module: generating adversarial samples of the original multimodal data of the single data source by fast gradient sign method and iterative fast gradient sign method, performing cross-modal association analysis on the multimodal data, and performing fusion processing and cross-modal alignment processing on the extracted features; Integration module: Establish an equalization system through deep equalization multimodal fusion technology, and calculate the optimal transmission path between the processed features of different modes and perform feature integration in the equalization system by minimizing the cost function, the lower semi-continuous function, the joint probability distribution and the Sinkhorn algorithm; Prediction module: Based on the bidirectional encoder structure and integrated features, a multimodal deep situational awareness and prediction model is constructed, and the adversarial sample is used for training, and the residual cycle prediction technology is integrated into the multimodal deep situational awareness and prediction model, and the multimodal deep situational awareness and prediction model is used for network security situational awareness.
9. A computer program product having one or more computer programs thereon, characterized in that: When the one or more computer programs are executed by a computer processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Network security threat detection method and system based on multi-modal artificial intelligence
CN120658466A
A multi-modal artificial intelligence based cyber security threat detection method and system
CN120658466B
Network security threat identification method and system based on network security knowledge structure perception
CN121333806A