A Network Anomaly Detection Method Based on Multimodal Federated Active Learning

Through the multimodal federated active learning network anomaly detection method, the problems of data dispersion and privacy security in a distributed network environment are solved, efficient and accurate network anomaly detection is achieved, adapting to node heterogeneity and reducing communication overhead.

CN119210899BActive Publication Date: 2025-06-13南京智能计算科技发展有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411698904.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-26
Publication Date
2025-06-13
Estimated Expiration
2044-11-26

AI Technical Summary

Technical Problem

Existing network anomaly detection methods are difficult to process decentralized data in distributed network environments, and traditional centralized model training methods are prone to data leakage and privacy violations, and single modal data cannot fully capture complex cyber attack behaviors.

Method used

A network anomaly detection method based on multimodal federated active learning is adopted, and aggregation of multimodal features and fuses text, vision, voice and other data is introduced, allowing distributed nodes to independently select learning strategies, and optimize communication efficiency through adaptive learning optimization strategies and progressive federated synchronization mechanisms.

Benefits of technology

On the premise of protecting data privacy, the accuracy and efficiency of network anomaly detection are improved, the comprehensiveness and accuracy of detection are enhanced, the cost of data labeling and communication overhead are reduced, and the node heterogeneity is adapted to the distributed environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119210899B_ABST
    Figure CN119210899B_ABST
Patent Text Reader

Abstract

The present invention discloses a network anomaly detection method based on multi-modal federated active learning, including the steps of: fusing data of several modalities through multi-modal feature aggregation to generate a unified feature representation; introducing an active learning mechanism in the federated learning framework to allow distributed nodes to autonomously select learning strategies according to the characteristics and distribution of their own data; coping with node heterogeneity in the distributed environment through an adaptive learning optimization strategy; optimizing communication efficiency and reducing the frequency of global synchronization through a progressive federated synchronization mechanism; this method combines multi-modal data fusion and federated active learning to improve the accuracy and efficiency of network anomaly detection while protecting data privacy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of network anomaly detection, and particularly to a network anomaly detection method based on multi-modal federated active learning. Background Art

[0002] With the booming development of technologies such as the Internet, cloud computing, and the Internet of Things, the global network traffic and the number of connected devices have shown an explosive growth, and the complexity and diversity of network attacks have also been continuously upgraded. Network anomaly detection, as a key technology to ensure network security, aims to discover abnormal behaviors existing in network systems, such as malicious attacks, data leakage, and abnormal operations. However, the distributed characteristics of modern network environments and data privacy requirements have brought new challenges to network anomaly detection. First, in cloud computing and Internet of Things applications, data is usually distributed among different physical devices or nodes and cannot be centrally collected and processed. Such a distributed environment requires anomaly detection methods to be able to adapt to the decentralized network architecture and perform distributed data analysis and model training. Second, many organizations and individual users have a high concern for data privacy and security. Traditional centralized model training methods usually need to upload a large amount of data to a central server for processing, which is likely to cause data leakage and privacy infringement. Therefore, how to perform efficient network anomaly detection without centralizing data has become an urgent need. Finally, in actual network security scenarios, various data modalities such as network traffic, log information, system calls, and sensor data can be used to detect network abnormal behaviors. Different modalities of data contain different information dimensions. Relying solely on single-modal data may miss important abnormal features and cannot comprehensively reflect complex network attack behaviors. Therefore, how to fuse multi-modal data to enhance the accuracy and comprehensiveness of network anomaly detection has become a key issue.

[0003] Currently, the main network anomaly detection methods are as follows: Rule-based anomaly detection methods rely on predefined rules or thresholds to judge abnormal behaviors. They usually use fixed attack signature libraries, feature rule sets, etc. to identify known threats. Such methods have a high detection rate for known attacks, but have poor detection capabilities for unknown threats and are difficult to adapt to complex and dynamically changing network environments. To address the limitations of rule-based methods, in recent years, machine learning technologies have gradually been applied to network anomaly detection. By learning the patterns of normal and abnormal behaviors from historical data, machine learning models can detect unknown threats that are not clearly defined. However, most machine learning methods rely on centralized data collection and model training, which is not feasible in modern distributed network environments, especially when the data involves privacy-sensitive content, and centralized processing will cause privacy leakage problems. In addition, most existing machine learning methods only use single-modal data for detection, such as only using network traffic data or log information, resulting in limited detection coverage. Summary of the Invention

[0004] To this end, a network anomaly detection method based on multi-modal federated active learning is required. By combining multi-modal data fusion and federated active learning, the accuracy and efficiency of network anomaly detection are improved while protecting data privacy.

[0005] To achieve the above object, the inventor provides a network anomaly detection method based on multi-modal federated active learning, including the steps of:

[0006] S1, fusing data of several modalities through multi-modal feature aggregation to generate a unified feature representation;

[0007] S2, introducing an active learning mechanism in the federated learning framework to allow distributed nodes to independently select learning strategies according to the characteristics and distribution of their own data;

[0008] S3, coping with node heterogeneity in the distributed environment through an adaptive learning optimization strategy;

[0009] S4, optimizing communication efficiency and reducing the frequency of global synchronization through a progressive federated synchronization mechanism.

[0010] As a preferred embodiment of the present invention, several modalities include: text modality, visual modality, and speech modality. The text modality includes log files, the visual modality includes data stream data, and the speech modality includes voice communication data.

[0011] As a preferred embodiment of the present invention, the step S1 further includes:

[0012] S101, where several modalities include the text modality. Extract text features from the data source of the text modality. For the text features, perform context-level partitioning on them, and each level of context is represented by a bidirectional encoder representation model to form a hierarchical context embedding representation of the text modality , and use an attention mechanism to adaptively select informative text features between different levels , the expression is:

[0013] ;

[0014] Among them, represents a time-related adaptive weight, which is dynamically adjusted as the time step t changes, represents the self-attention mechanism, represents a dynamic weighting coefficient related to the attention weight;

[0015] S102, several modalities include the visual modality. Extract data stream data from the data source of the visual modality, and fuse the data stream data through a multi-view feature learning model. The expression is:

[0016] ;

[0017] Among them, represents the visual feature representation after multi-view fusion, represents the adaptive weight coefficient of the i-th view, represents the geometric transformation matrix. N represents the total number of views, which is used to align the features of the i-th camera view to the same geometric space as the j-th camera view, represents the feature representation extracted from the i-th camera view, represents the feature representation extracted from the j-th camera view, is used to balance the geometric consistency loss and the overall loss in the feature fusion process, represents the geometric consistency constraint loss function;

[0018] S103, several modules include the speech modality. Extract speech communication data from the data source of the speech modality. For the speech communication data, fuse it through a differential signal decomposition and neural modality alignment mechanism. The expression is:

[0019] ;

[0020] Among them, represents the speech feature representation, represents the differential signal decomposition, which decomposes the input speech into several semantic components , represents the weights of different semantic components, represents the frequency alignment operation, represents the alignment weight. DTW represents the time domain alignment operation, and T represents the time alignment parameter, represents the contribution degree of each semantic component, and M represents the number of semantic components;

[0021] S104, an adaptive fusion mechanism based on modal differential collaborative modeling, actively utilizes modal differences, and adjusts the fusion of each modal feature through a modal collaborative conflict resolution mechanism. The expression is:

[0022] ;

[0023] Among them, and represent the features of any modality, represents different modal indices, and It represents the collaborative learning enhancement weight, which controls the influence degree of the remaining modules on the current k-modal enhancement. It represents feature collaborative learning. It represents the weight for controlling conflicting features. It represents conflicting features. It represents the fused multi-modal features.

[0024] As a preferred embodiment of the present invention, the text features include system log features, and the system log features are hierarchically divided by context according to time, location, and operation type.

[0025] As a preferred embodiment of the present invention, step S2 further includes:

[0026] S201, In the distributed federated learning model, each node autonomously selects a learning strategy to perform multi-objective optimization of the node. The multi-objective optimization function of the node The expression is:

[0027] ;

[0028] Wherein, represents the learning strategy of node b, represents the influence weight of the th factor, represents the fused multi-modal features, represents the th factor related to the learning strategy

[0029] S202, Each node generates pseudo-labels through self-supervised signals and continuously learns and updates through unlabeled data. The expression is:

[0030] ;

[0031] Wherein, represents the parameters of node b after rounds of training, represents the current training parameters, represents the learning rate, which is used to control the update of parameters, represents the gradient of the loss function, represents the unlabeled dataset, represents the loss function, represents the error on the unlabeled data x, represents the pseudo-labels generated through self-supervision;

[0032] S203, Dynamically adjust the communication frequency through the change amount dynamic evaluation mechanism and the adaptive communication control strategy. The expression is:

[0033] ;

[0034] Among them, represents the communication trigger probability of node in the t-th round, which is used to determine whether to synchronize with the global model. represents the parameter change amount. is the set change threshold. is the adjustment parameter. represents the exponential function with the natural number e as the base.

[0035] As a preferred embodiment of the present invention, step S3 further includes:

[0036] S301. Based on the adaptive learning optimization strategy, dynamically adjust the task allocation strategy through the Q-learning algorithm. The expression is:

[0037] ;

[0038] Among them, represents the task allocation strategy of node . represents the computing resources, storage resources, and network bandwidth of node . represents the current node task executing the action task allocation of the long-term cumulative reward function. represents the discount factor, which is used to balance the immediate reward and the long-term reward. represents the mathematical expectation;

[0039] S302. According to the resource status of each node, dynamically prune the non-critical layers in the model for lightweighting. The expression is:

[0040] ;

[0041] Among them, represents the initial complexity of the distributed federated learning model. represents the computing resources, storage resources, and network bandwidth of node . represents the weight parameter that determines the influence of each resource on pruning. represents the complexity of the distributed federated learning model after pruning.

[0042] As a preferred embodiment of the present invention, step S4 further includes: S401. After each round of training, each node dynamically selects important parameters for compression and synchronization according to the differences of the distributed federated learning model. Meanwhile, a gradual fusion strategy is adopted. When performing global updates, the updates from each node are gradually fused in the form of weight adjustment, and the expression is:

[0043] ;

[0044] where, represents the parameters of the global model in rounds, represents the learning rate, represents the weight of node at round t, represents the compressed differential parameters, represents the parameters of node .

[0045] Different from the prior art, the beneficial effects achieved by the above technical solutions are as follows:

[0046] (1) Most existing network anomaly detection methods rely on single-modal data sources. The detection accuracy and comprehensiveness of this method are often limited by the singularity of the data source and cannot capture complex abnormal behaviors. This method can capture complex abnormal behaviors in the network environment from different dimensions by introducing multi-modal data sources, such as text, vision, voice, etc. The multi-modal data fusion method enhances the comprehensiveness and accuracy of network anomaly detection by combining the advantages of each modality, and solves the deficiency that single-modal data cannot cover complex attack scenarios;

[0047] (2) The federated autonomous learning method of this method allows each distributed node to independently determine the optimal learning strategy according to its own local data distribution and characteristics, without relying on the unified guidance of the global model. This autonomous learning mechanism can ensure that each node performs the most effective model update for its unique data structure during the training process, thereby reducing the data annotation cost and improving the learning efficiency and performance of the local model; at the same time, each node can reduce frequent parameter exchanges with other nodes while making full use of its own data resources, reducing communication overhead;

[0048] (3) In a distributed network environment, there may be differences in hardware configuration, computing power, data type, etc. among different nodes. The federated autonomous learning method proposed by this method can perform adaptive adjustment according to the heterogeneity of each node. Nodes with low computing power can choose lightweight learning strategies, while high-performance nodes can perform more local training tasks to ensure the efficient cooperation of the entire federated system. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1It is the method framework diagram described in the specific implementation manner. Specific implementation manner

[0050] To describe in detail the technical content, structural features, achieved objectives and effects of the technical solution, the following will be described in detail with specific embodiments in conjunction with the accompanying drawings.

[0051] As Figure 1 shown, this embodiment provides a network anomaly detection method based on multi-modal federated active learning, which is characterized by including the steps:

[0052] S1, fuse the data of several modalities through multi-modal feature aggregation to generate a unified feature representation; specifically: fuse the data of different modalities such as text, vision, and speech. The data source of each modality represents different dimensions of information in the network. For example, the system log file is in text modality, the video stream data is in vision modality, and the voice communication data is in speech modality. After these modality data are subjected to feature extraction and preprocessing, a multi-modal fusion strategy is adopted to combine the information of different modalities to generate a unified feature representation.

[0053] S2, introduce an active learning mechanism in the federated learning framework to allow distributed nodes to independently select learning strategies according to the characteristics and distribution of their own data; in this way, the global synchronization and the dependence on labeled data can be significantly reduced. Each node analyzes the characteristics of the local data, such as specific network traffic patterns, user behavior logs, etc., and selects the most suitable distributed federated learning model structure for local data for independent training. During the autonomous learning process, each node can independently evaluate the performance of the local model and dynamically adjust the training strategy according to data changes without relying on the unified guidance of the global model. For some nodes, this mechanism allows them to continue to improve the performance of the local model through unsupervised learning strategies in the case of unlabeled data. In addition, when the updates of the nodes are small, they will not synchronize with the global model frequently, thus effectively reducing the communication frequency and overhead.

[0054] S3, adopt an adaptive learning optimization strategy to cope with node heterogeneity in the distributed environment; specifically, each node dynamically selects an appropriate task load and model complexity according to its own hardware configuration, computing resources, storage capacity, and network bandwidth. The system can allocate tasks according to the resource status of the nodes: for nodes with low computing power, lightweight model training tasks are allocated; for high-performance nodes, more complex training tasks are allocated. This adaptive learning optimization strategy ensures the efficient cooperation of all nodes in the heterogeneous environment and makes full use of the hardware capabilities of each node.

[0055] S4. Optimize communication efficiency and reduce the frequency of global synchronization through a progressive federated synchronization mechanism. Specifically, this method is different from the synchronization method in traditional federated learning. This method proposes an innovative progressive federated synchronization mechanism aimed at optimizing communication efficiency and reducing the frequency of global synchronization. This mechanism not only relies on model difference evaluation but also introduces model compression technology and a gradual fusion strategy, which can further improve communication efficiency and synchronization effects.

[0056] In the specific implementation process of step S1 in the above embodiment, step S1 includes the following steps:

[0057] S101. First, extract features from data sources of different modalities. These features represent information in different dimensions of the network, such as text features in system logs, visual features in surveillance videos, and audio features in voice communications. For each modality of data, use a specific feature extraction method, and represent and normalize these features so that they can be uniformly processed. For text features, perform context-level partitioning on them and extract multi-level features. Events in system logs can be divided into different levels according to time, location, and operation type. The context of each level is represented by a bidirectional encoder representation model to form a hierarchical context embedding representation of the text modality , and use the attention mechanism to enable the system to adaptively select the most informative text features between different levels , the expression is:

[0058] ;

[0059] Among them, represents a time-related adaptive weight that dynamically adjusts with the change of the time step t, represents the self-attention mechanism, represents a dynamic weighting coefficient related to the attention weight, which is responsible for adjusting the influence degree of the attention mechanism during multi-layer context aggregation. This coefficient is updated through gradient descent in the model training stage and can be adaptively adjusted according to task-specific requirements;

[0060] S102. For video stream data, although traditional convolutional neural networks can extract visual features, it is difficult to handle dynamic scene changes in surveillance videos or images. This embodiment proposes a multi-view feature learning model based on geometric consistency constraints, which can process complex visual scenes in real time, especially multi-views and scene changes in surveillance videos. The expression is:

[0061] ;

[0062] Among them, It represents the visual feature representation after multi-view fusion, which is a global representation formed by integrating the visual features extracted from different camera views. It represents the adaptive weight coefficient of the i-th camera view. It represents the geometric transformation matrix. N represents the total number of views, which is used to align the features of the i-th camera view to the same geometric space as the j-th camera view. It represents the feature representation extracted from the i-th camera view. It represents the feature representation extracted from the j-th camera view. It is used to balance the geometric consistency loss and the overall loss in the feature fusion process. It represents the geometric consistency constraint loss function, which is used to minimize the geometric error between the features of multiple views, ensure that the features extracted from different views can be aligned, and guarantee that the model obtains consistent feature representations from different views.

[0063] S103. For speech data, traditional speech modeling methods mainly rely on fixed convolutional or recurrent network structures, but they perform limitedly when dealing with speech signals with strong background noise, emotional changes, and multiple speaker patterns. To more accurately capture the dynamic features in speech and enhance the adaptability to complex scenarios, this embodiment proposes a speech modeling method based on differential signal decomposition and neural modality alignment mechanism, and the expression is:

[0064] ;

[0065] Where, It represents the speech feature representation. It represents the differential signal decomposition, which decomposes the input speech into multiple semantic components. , It represents the weights of different semantic components. It represents the frequency alignment operation. It represents the alignment weight. DTW represents the time-domain alignment operation, and T represents the time alignment parameter. It represents the contribution degree of each semantic component, and M represents the number of semantic components.

[0066] S104. Based on the adaptive fusion mechanism of modal differential collaborative modeling, it actively utilizes the modal differences. Through the "modal collaborative conflict resolution" mechanism, the differences between modalities are made into advantages in the fusion process, rather than simply pursuing alignment. This embodiment not only improves at the feature alignment level, but also deeply utilizes the differences between modalities to dynamically adjust the fusion of each modal feature to improve the model's ability to understand complex data, and the expression is:

[0067] ;

[0068] Among them, and represent features of any modality, represents different modality indices, and represent the co-learning enhancement weights, which control the influence degree of other modules on the enhancement of the current k modality, represents feature co-learning, represents controlling the weights of conflicting features, represents conflicting features, is the fused multi-modal feature.

[0069] In the specific implementation process of step S2 in the above embodiment, step S2 includes the following steps:

[0070] S201, in the distributed federated learning framework, each node autonomously selects a learning strategy, which is not only based on a single goal, such as model accuracy, but combines various factors such as the computing resources, network conditions, and local data characteristics of the node, such as data distribution, quality, and noise level, to determine the optimal learning strategy. Each node regards the learning strategy selection problem as a multi-objective optimization problem, assigns weights to different goals through an adaptive algorithm, and finally autonomously selects the most suitable learning strategy for the current data and system conditions. The multi-objective optimization function expression of the node is:

[0071] ;

[0072] Among them, represents the learning strategy of node b, represents the influence weight of the th factor, is the fused multi-modal feature, represents the metric between the th factor related to the strategy and, such as the accuracy, resource consumption, and synchronization cost of the distributed federated learning model. This function contains multiple sub-goals, such as minimizing computing resource consumption, minimizing communication frequency, and minimizing the impact of data noise. Each node dynamically adjusts the weight of each factor , for example: when the computing resources of the node are insufficient, the weight of the resource consumption goal can be increased, so as to select a learning strategy with less computing overhead;

[0073] S202. In traditional federated learning, nodes usually rely on the update guidance of the global model. In this embodiment, each node can dynamically evaluate the performance of the local model through a self-supervised learning mechanism and adjust the model parameters in real time according to the changes in the local data distribution, reducing the dependence on labeled data. Through self-supervised signals, the model of each node can generate pseudo-labels and continuously learn through unlabeled data. The expression for the model update method is:

[0074] ;

[0075] Among them, represents the parameters of node b after rounds of training, represents the current training parameters, represents the learning rate, which is used to control the update of the parameters, represents the loss function, represents the error on the unlabeled data x, represents the pseudo-labels generated through self-supervision, represents the gradient related to the loss function, represents the unlabeled dataset;

[0076] S203. In federated learning, the synchronization frequency between nodes and the global model directly affects the communication overhead and system efficiency. In the traditional federated learning framework, each node needs to frequently synchronize parameters with the global model, which consumes a large amount of network bandwidth and computing resources. To optimize the communication frequency, in this embodiment, through a model change amount dynamic evaluation mechanism and an adaptive communication control strategy, the communication frequency is dynamically adjusted according to the magnitude of the local model change amount, thereby reducing unnecessary global synchronization, reducing the communication overhead, and improving the system efficiency. The adaptive communication optimization expression based on the model change amount is:

[0077] ;

[0078] Among them, represents the communication trigger probability of node at the t-th round, which is used to determine whether to synchronize with the global model, represents the parameter change amount, represents the set change threshold, is the adjustment parameter that controls the steepness of the communication probability function, so that when the model changes greatly, the communication probability increases rapidly, represents the exponential function with the natural number e as the base.

[0079] In the specific implementation process of step S3 in the above embodiment, step S3 includes the following steps:

[0080] S301. Based on the adaptive learning optimization strategy, the present invention dynamically adjusts the task allocation strategy through the Q-learning algorithm, rather than relying solely on predefined rules or static parameters to allocate tasks. Instead, it dynamically adjusts task allocation through the algorithm. Through node feedback, such as computing speed, resource consumption, etc., the system can learn the resource status of each node and find the optimal task allocation strategy through trial and error and learning, ensuring the maximization of the resource utilization rate of heterogeneous nodes. The expression is:

[0081] ;

[0082] Among them, represents the task allocation strategy of node , represents the computing resources, storage resources, and network bandwidth of node , represents the long-term cumulative reward function for the current node task to execute the action task allocation , is the discount factor, used to balance immediate rewards and long-term rewards, represents the mathematical expectation. Through the Q-learning algorithm, this embodiment can dynamically adjust the task load to ensure the optimal utilization of resources for each node. The task allocation of each node not only depends on the current resource status but also optimizes the strategy based on the feedback of past task executions.

[0083] S302. Based on the adaptive model complexity adjustment of hierarchical model pruning, in this embodiment, according to the resource status of each node, an adaptive model complexity adjustment mechanism is provided. This method not only considers the hardware resources of the nodes but also dynamically prunes non-critical layers in the model according to the resource status of each node, such as neuron connections with smaller weights, to generate a lightweight model version. Nodes can prune some layers or channels of the neural network according to their own computing capabilities to reduce the model complexity, enabling low-resource nodes to also participate in complex tasks. The expression is:

[0084] ;

[0085] Among them, represents the initial complexity of the distributed federated learning model, represents the computing resources, storage resources, and network bandwidth of node , represents that the weight parameter determines the impact of each resource on pruning, represents the complexity of the distributed federated learning model after pruning. If the computing resources are scarce and the storage and bandwidth are also limited, It will approach 1, which means that the pruning ratio of the distributed federated learning model is relatively large, generating a lightweight distributed federated learning model. On the contrary, if the computing resources are sufficient and the storage and bandwidth are abundant, it will approach 0, meaning that the model pruning ratio is small and the complete model structure is retained.

[0086] In the specific implementation process of step S4 in the above embodiment, step S4 includes the following steps:

[0087] S401. In this embodiment, through the model compression and gradual fusion strategy, the communication efficiency in federated learning is greatly optimized, and the frequency of global synchronization is reduced. Specifically, after each round of training, the nodes do not immediately synchronize all model parameters, but dynamically select important parameters for compression and synchronization according to the model differences. At the same time, the system adopts a gradual fusion strategy. When the global model is updated, the model updates from each node are gradually fused in a more flexible weight adjustment manner, thereby reducing unnecessary synchronization overhead and ensuring the synchronization effect and model convergence speed. The expression is:

[0088] ;

[0089] Among them, represents the parameters of the global model in rounds, represents the learning rate, represents the weight of node at round t, represents the compressed differential parameters, represents the parameters of node .

[0090] To prove the effectiveness of this method, a public dataset was used for verification. Specifically as follows: In the Internet of Things (IoT) in smart cities, the network dataset IOT-23 is generated by various edge devices, including smart sensors, cameras, vehicles, air quality detectors, and network traffic monitors, etc. Such datasets are characterized by multimodality and are suitable for anomaly detection tasks in the field of network security, and can help detect distributed denial of service (DDoS) attacks, data leaks, network congestion, and other network threats. The performance comparison experimental results of different network anomaly detection methods are shown in Table 1:

[0091] Table 1: Performance comparison experimental results of different network anomaly detection methods

[0092]

[0093] The experimental results show that this method performs significantly better than other methods in network anomaly detection performance. As can be seen from the table, the accuracy, recall rate, and F1 value of this method reach 97.1%, 95.1%, and 96.1% respectively, far higher than other models. This indicates that by introducing multi-modal federated active learning, this method can effectively capture complex abnormal behaviors and improve the comprehensiveness and accuracy of detection. In contrast, traditional methods such as BiGRU, Siamese-BiGRU, and BERT+WMD perform less well in these metrics, with the accuracy and recall rate generally between 70% and 83%, and the F1 value between 66.8% and 78.4%. This shows that the model of this method can not only process multi-modal data in a distributed environment, but also greatly improve the overall detection efficiency while reducing false negatives and false positives.

[0094] It should be noted that although the above embodiments have been described in this article, the patent protection scope of the present invention is not limited thereby. Therefore, based on the innovative concept of the present invention, any changes and modifications made to the embodiments described in this article, or equivalent structural or equivalent process transformations made using the content of the specification and drawings of the present invention, and the direct or indirect application of the above technical solutions to other related technical fields, are all included in the patent protection scope of the present invention.

Claims

1. A network anomaly detection method based on multimodal federated active learning, characterized in that: Includes steps: S1, through multimodal feature aggregation, the data of several modalities are fused to generate a unified feature representation; S2, introduces active learning mechanism into the federated learning framework, allowing distributed nodes to autonomously select learning strategies based on the characteristics and distribution of their own data; S3, copes with node heterogeneity in distributed environments through adaptive learning optimization strategies; S4, optimizes communication efficiency and reduces the frequency of global synchronization through a progressive federated synchronization mechanism; Step S1 includes: S101, several modalities including text modality, extract text features from the data source of the text modality, divide the text features into context levels, and represent the context of each level through a bidirectional encoder representation model to form a hierarchical context embedding representation of the text modality , and use the attention mechanism to adaptively select informative text feature representations between different layers , the expression is: ; in, represents a time-dependent adaptive weight that is dynamically adjusted as the time step t changes. represents the self-attention mechanism, Represents the dynamic weighting coefficient related to the attention weight; S102, several modalities including visual modality, extracting data stream data from the data source of the visual modality, and fusing the data stream data through a multi-view feature learning model, the expression is: ; in, represents the visual feature representation after multi-view fusion, represents the adaptive weight coefficient of the i-th view, Represents the geometric transformation matrix, N represents the total number of viewing angles, and is used to align the features of the i-th camera viewing angle to the geometric space consistent with the j-th camera viewing angle. represents the feature representation extracted from the i-th camera view, represents the feature representation extracted from the j-th camera view, Used to balance the geometric consistency loss and the overall loss in the feature fusion process. represents the geometric consistency constraint loss function; S103, several modules include speech modality, and speech communication data is extracted from the data source of the speech modality. The speech communication data is fused through differential signal decomposition and neural modality alignment mechanism, and the expression is: ; in, represents the speech feature representation, Represents differential signal decomposition, which decomposes the input speech into several semantic components , represents the weights of different semantic components, represents the frequency alignment operation, represents the alignment weight, DTW represents the time domain alignment operation, T represents the time alignment parameter, represents the contribution of each semantic component, and M represents the number of semantic components; S104, an adaptive fusion mechanism based on modal differentiation collaborative modeling, actively utilizes modal differences and adjusts the fusion of modal features through modal collaborative conflict resolution mechanism. The expression is: ; in, and represents the characteristics of any mode, Indicates different modal indexes, and Represents the collaborative learning enhancement weight, which controls the influence of other modules on the current k-modal enhancement. Represents feature collaborative learning, represents the weight of controlling conflicting features, Indicates conflict characteristics, Represents the fused multimodal features; Step S2 includes: S201, in the distributed federated learning model, each node independently selects a learning strategy to perform multi-objective optimization of the node. The multi-objective optimization function of the node The expression is: ; in, represents the learning strategy of node b, Indicates The influence weight of each factor is represents the fused multimodal features, Representation and Learning Strategies Related The measurement between factors; In S202, each node generates a pseudo label through self-supervision signals and continuously learns and updates through unlabeled data. The expression is: ; in, Indicates that node b is The parameters after round training, represents the current training parameters, Represents the learning rate, which is used to control the update of parameters. represents the gradient of the loss function, represents an unlabeled dataset, represents the loss function, represents the error on the unlabeled data x, represents the pseudo-label generated by self-supervision; S203, dynamically adjust the communication frequency through the dynamic evaluation mechanism of the variation and the adaptive communication control strategy, the expression is: ; in, Representation Node The communication trigger probability in round t is used to decide whether to synchronize with the global model. Indicates the parameter change, is the set change threshold, To adjust the parameters, represents an exponential function with natural number e as base; Step S3 includes: S301, based on the adaptive learning optimization strategy, dynamically adjust the task allocation strategy through the Q-learning learning algorithm, the expression is: ; in, Representation Node The task allocation strategy Representation Node computing resources, storage resources, and network bandwidth, Indicates the current node task Execute action task allocation The long-term cumulative reward function is represents the discount factor, which is used to weigh the immediate reward and long-term reward, represents the mathematical expectation; S302, according to the resource status of each node, dynamically prune the non-critical layers in the model to make it lightweight, the expression is: ; in, represents the initial complexity of the distributed federated learning model, Representation Node computing resources, storage resources, and network bandwidth, Indicates the weight parameter that determines the impact of each resource on pruning. Represents the complexity of the distributed federated learning model after pruning.

2. The network anomaly detection method based on multimodal federated active learning according to claim 1 is characterized in that: The several modalities include: a textual modality including log files, a visual modality including data stream data, and a voice modality including voice communication data.

3. The network anomaly detection method based on multimodal federated active learning according to claim 1 is characterized in that: The text features include system log features, and the system log features are divided into context levels according to time, location and operation type.

4. The network anomaly detection method based on multimodal federated active learning according to claim 1 is characterized in that: S401, after each round of training, each node dynamically selects important parameters for compression and synchronization based on the differences in the distributed federated learning model, and adopts a gradual fusion strategy. During global updates, the updates from each node are gradually integrated in a weight adjustment manner. The expression is: ; in, Represents the global model in Wheel parameters, represents the learning rate, Representation Node The weight in round t, Represents the difference parameter after compression, Representation Node Parameters.

Citation Information

Patent Citations

  • Federal learning-based edge heterogeneous network intrusion detection method and system

    CN118573442A

  • Power grid health assessment and analysis method based on multiple modes

    CN118657404A