A train communication network intrusion detection method based on multi-scale residual networks
An intrusion detection method combining multi-scale residual networks and global-local attention mechanisms solves the problems of class imbalance and ambiguous class boundaries in train communication networks, thereby improving the accuracy and recognition capability of intrusion detection.
Patent Information
- Application Number
- CN202411591538.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-08
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2044-11-08
AI Technical Summary
Existing technologies in train communication networks suffer from class imbalance and ambiguous class boundaries, resulting in low intrusion detection accuracy and difficulty in effectively identifying covert intrusion traffic.
An intrusion detection method based on multi-scale residual networks and global and local attention mechanisms is adopted. By constructing a train communication network topology model, combining multi-scale residual networks and global and local attention mechanisms, and using an improved focus loss function to optimize model training, multi-scale feature information is extracted and attention is strengthened to key features, reducing redundant feature interference and improving detection accuracy.
It improves the accuracy, recall, and F1 score of intrusion detection in train communication networks, enhances the ability to identify covert intrusion traffic, and alleviates the problems of class imbalance and boundary ambiguity.
Smart Images

Figure CN119652559B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of train communication network technology, and in particular relates to a train communication network intrusion detection method based on multi-scale residual networks. Background Technology
[0002] Train communication security is closely related to passenger safety and information security, and is also a crucial guarantee for the safe operation of high-speed trains. With the expansion of the railway network and the increase in intelligent onboard equipment, the openness of train communication networks is constantly increasing, and the amount of data transmission is also rising sharply. This poses a severe challenge to the security of train communication networks. Faced with an increasingly complex network environment, intrusion detection technology can provide a proactive defense strategy for the security of train communication networks. By monitoring network traffic in real time, it can effectively identify network intrusions. Intrusion detection technology is crucial for ensuring the security of train communication networks.
[0003] With the increasing volume of communication data, data-driven intrusion detection technology has seen new developments. Machine learning-based intrusion detection relies on domain knowledge and experience to manually extract relevant features from historical network traffic data. This allows the intrusion detection model to learn the differences between normal and intrusive traffic, thereby detecting potential intrusion behaviors in real-time traffic. However, as intrusion methods continue to evolve, the boundary between some covert intrusion traffic and normal traffic is becoming increasingly blurred. Because traditional machine learning-based intrusion detection methods rely excessively on manually extracted features, their identification and detection performance is often unsatisfactory when faced with these difficult-to-classify intrusion traffic types.
[0004] In real-world train communication network environments, intrusion events are typically low-probability events, and the class distribution of intrusion detection datasets exhibits an objective imbalance. In imbalanced tasks, data-driven algorithms tend to favor majority class samples, neglecting the detection accuracy of minority class samples. Therefore, the development of a reliable and accurate intrusion detection method is urgently needed in train communication network scenarios characterized by imbalanced traffic data and ambiguous class boundaries. Summary of the Invention
[0005] The purpose of this invention is to provide an intrusion detection method for train communication networks based on multi-scale residual networks, thereby addressing the problem of low intrusion detection accuracy in existing technologies under scenarios with class imbalance and ambiguous class boundaries. To achieve the above objective, this invention adopts the following technical solution:
[0006] This invention provides a train communication network intrusion detection method based on multi-scale residual networks and global and local attention mechanisms, comprising the following steps:
[0007] S01: Establish a train communication network topology model and define the performance parameters of the switching nodes, terminal equipment, and links in the train communication network topology model;
[0008] S02: Under the train communication network topology model, establish a traffic model for the mixed transmission of normal traffic and intrusion traffic;
[0009] S03: Collect traffic data packets in the link as the original data samples in the intrusion detection dataset. After preprocessing the original data samples by parsing, normalizing and time windowing, a data sample containing statistical features is obtained.
[0010] S04: Merge all data samples in chronological order to form an intrusion detection dataset, and randomly shuffle the dataset to divide it into a training set and a test set in a 6:4 ratio; the training set is used for training the MSRNet-GLAM model for intrusion detection of train communication networks, and the test set is used for evaluating the inspection performance of the MSRNet-GLAM model for intrusion detection of train communication networks.
[0011] S05: Construct the MSRNet-GLAM model for train communication network intrusion detection. Input the training set data samples into the MSRNet-GLAM model for training. In each round of training, the MSRNet-GLAM model outputs the detection results of the data samples. Input the test set data samples into the MSRNet-GLAM model for intrusion detection. Check and evaluate the performance of the MSRNet-GLAM model for train communication network intrusion detection. If it shows good intrusion detection performance, then intrusion detection of the train communication network is achieved.
[0012] By combining multi-scale residual networks and global and local attention mechanisms, the MSRNet-GLAM model can learn key feature information of data samples at multiple scales. When data samples are input into the MSRNet-GLAM model for training, an improved focus loss function is constructed to measure the difference between the detection results and the actual labels of these data samples. With minimizing the output of the loss function as the optimization objective, the difference between the MSRNet-GLAM model output and the actual labels is gradually reduced through multiple rounds of training, thus optimizing the detection performance of the MSRNet-GLAM model. The final result is a well-trained MSRNet-GLAM model that has fully learned the feature patterns of the intrusion detection dataset, and the difference between the detection results and the actual labels in each round of training. Minimizing this difference is used as the optimization objective to optimize the detection performance of the MSRNet-GLAM model on the training set data samples in each round of training. Finally, a well-trained MSRNet-GLAM model, fully learning the feature patterns of the intrusion detection dataset, is obtained. The test set data samples to be detected are input into the trained MSRNet-GLAM model. The MSRNet-GLAM model detects each sample in the test set based on the feature patterns learned in the training set. The performance of the MSRNet-GLAM intrusion detection model is evaluated using accuracy, recall, precision, and F1 score. If the MSRNet-GLAM intrusion detection model exhibits good accuracy, recall, precision, and F1 score on the test set, it can be used for intrusion detection in train communication networks.
[0013] In a further preferred embodiment of the above scheme, in step S01, the construction of the train communication network topology consists of an interconnected train backbone network and a vehicle formation network. Each vehicle formation network's switching node is connected to its subordinate terminal equipment (ED) for data transmission. The definition of the switching node mainly includes the number of switching nodes, the number of switching node ports, and IP address configuration. The definition of the terminal equipment mainly includes the number of terminal nodes and the IP address configuration. The definition of the link parameters mainly includes the number of links, link bandwidth, and link latency. In this invention, establishing a train communication network topology model and defining the performance parameters of the switching nodes, terminal equipment, and links in the train communication network topology model specifically includes: first, constructing a train communication network topology, which consists of an interconnected train backbone network and a vehicle formation network, with each vehicle formation network's switching node connected to its subordinate terminal equipment (ED) for data transmission; and then defining the performance parameters of the switching nodes, terminal equipment, and links in the train communication network topology model.
[0014] In a further preferred embodiment of the above scheme, in step S02, the normal traffic includes monitoring data, process data, message data, stream data, and best effort delivery data; the intrusion traffic includes: sniffing intrusion traffic, denial-of-service traffic, Address Resolution Protocol (ARP) spoofing traffic, and replay intrusion traffic. Establishing a traffic model for the mixed transmission of normal traffic and intrusion traffic includes the following sub-steps:
[0015] S0201: In accordance with the International Electrotechnical Commission standard IEC 61375, the normal operating traffic of the train is configured in the train communication network as a normal sample in the normal traffic dataset; wherein the normal traffic includes monitoring data, process data, message data, streaming data, and best-effort delivery data. Among these data, monitoring data and process data are used to transmit critical information such as train control commands and train safety status information; these two types of data are real-time periodic data. Message data mainly transmits fault diagnosis information and has high priority and real-time requirements. Streaming data mainly transmits audio and video data related to train video monitoring and passenger entertainment activities; it is non-periodic, low real-time data. Best-effort delivery data involves data configuration and file transfer and has no real-time requirements.
[0016] S0202: Configure sniffing intrusion traffic (Probe) in the train communication network. Sniffing intrusion probe is a detection-type intrusion method. It sends request packets to the link indiscriminately through IP scanning and port scanning, obtains response packets from the link, and then parses these response packets to explore active hosts and available ports in the network, providing feasible intrusion targets for subsequent intrusions.
[0017] S0203: Configure Denial of Service (DoS) intrusion traffic in the train communication network. DoS is a flooding intrusion method that consumes the resources of the target device by sending a large number of invalid data packets, causing the target device to experience huge delays or even crashes, thus making it unable to provide services for normal requests, thereby achieving the purpose of "denying service" for normal requests.
[0018] S0204: Configure Address Resolution Protocol Spoofing (ARP Spoofing) and Replay intrusion traffic in the train communication network. Address Resolution Protocol Spoofing (ARP Spoofing) and Replay intrusion are man-in-the-middle intrusion methods.
[0019] S0205: Address Resolution Protocol (ARP) spoofing traffic achieves identity deception by sending false ARP replies to switching nodes, overwriting the target terminal's identity information in the switching node's cache. After receiving the data packet originally intended for the target host, the switching node forwards it to the intruder. The intruder intercepts the data packet and performs a replay attack, delaying or repeatedly injecting the intercepted data packet back into the network, thus enabling intrusion into the train communication network without deciphering the information in the data packet. In data transmission, the Address Resolution Protocol (ARP) broadcasts ARP requests containing the target IP address in the network. The physical address of the target terminal is determined by the ARP reply sent by the target terminal corresponding to the target IP address. ARP spoofing achieves identity deception by sending fake ARP replies to switching nodes, overwriting the target terminal's identity information in the switching node's cache. After receiving the data packet originally intended for the target host, the switching node forwards it to the intruder. The intruder can then intercept the data packet and launch a replay attack, which involves delaying or repeatedly re-injecting the intercepted data packet into the network, thus enabling intrusion into the train communication network without deciphering the information in the data packet.
[0020] In a further preferred embodiment of the above scheme, step S03, the preprocessing of the original data sample by packet parsing, normalization, and time windowing includes the following sub-steps:
[0021] S0301: Collect traffic data packets in the link, use the network packet parsing framework Scapy to parse the information in the data packets, and then encode the discrete features of each data packet with integers, such as IP address and port number; and normalize the continuous features, such as data packet length.
[0022] S0302: Time windowing is applied to data packet information in the link. The set of all data streams in the link is denoted by F:
[0023] F = {f1, f2, ..., f m},f i ∈F (1)
[0024] Where i = 1, 2, ..., m, m represents the number of data streams on the link, f i Represents a single data stream, each data stream f i It can be defined as a collection of data packets:
[0025]
[0026] in Represents data stream f i The j-th data packet in the data, for j∈1,2,...,n i n i This indicates the number of packets in each data stream, and each packet... They all have a timestamp
[0027] S0303: Transfer data stream f i It is divided into a series of time windows, each containing data packets within a time interval Δt. Define f. i The k-th time window is W (i,k) :
[0028]
[0029] in d i The starting timestamp of the k-th time window, where Δt represents the duration of the time window. Within the time window, data packets Only then was it included in W (i,k) middle.
[0030] S0304: Using a time window as the smallest unit, perform independent statistical analysis on the discrete characteristics of data packets within the time window. These discrete characteristics include the number of data packets, length, transmission time, transmission rate, and identifier count statistics. After statistical analysis, data samples containing statistical characteristics are obtained, and the statistical information is extracted as sample feature information for training the MSRNet-GLAM model.
[0031] S0305: Merge all data samples in chronological order to form an intrusion detection dataset, then randomly shuffle the dataset and divide it into a training set and a test set in a 6:4 ratio. The training set is used for training the MSRNet-GLAM model for intrusion detection in train communication networks, and the test set is used for evaluating the inspection performance of the MSRNet-GLAM model for intrusion detection in train communication networks.
[0032] In a further preferred embodiment of the above scheme, step S05, constructing the MSRNet-GLAM model for train communication network intrusion detection, specifically includes the following sub-steps:
[0033] S0501: As input to the MSRNet-GLAM model for train communication network intrusion detection, the training set data samples are first processed through a multi-scale residual network to obtain multi-scale feature information. The multi-scale residual network consists of multiple convolutional neural network layers with different kernel sizes. Let the kernel size be k. First, there is one pointwise convolutional layer, followed by three convolutional network branches with stacked convolutional layers of 1, 3, and 1 respectively. Specifically, the first branch's single-layer pointwise convolutional layer (k=1) is used to extract shallow input features; the second branch's three convolutional layers (k=3, 5, and 3) are used to extract deep input features; and the third branch's single-layer convolutional layer (k=3) is used to extract mid-level input features. Then, the outputs of the three branches are superimposed and input to another pointwise convolutional layer, and residually connected to the output of the first pointwise convolutional layer. All the above convolutional layers include one batch normalization layer to increase the stability of the MSRNet-GLAM model training by normalizing the input.
[0034] In the multi-scale residual module, the first pointwise convolutional layer is used to adjust the dimension of the features. The dimension-adjusted features are then input into various branches of the convolutional network, which have different numbers of convolutional layers and different kernel sizes, facilitating the extraction of multi-scale information of different granularities, such as deep and shallow layers, from high-dimensional features.
[0035] S0502: The input of the MSRNet-GLAM model for train communication network intrusion detection is processed by a multi-scale residual module to extract multi-scale feature information. The multi-scale feature information of the data samples is then input into the global and local attention modules for weighted processing to enhance the attention to important information in the multi-scale feature information. Inputting the multi-scale feature information into the global and local attention modules enhances the MSRNet-GLAM model's ability to capture key information between features and channels.
[0036] A further preferred embodiment of the above scheme involves inputting the multi-scale feature information of the data samples into the global and local attention modules for weighted processing, including the following steps:
[0037] First, a Local Attention Module (LAM) and a Global Attention Module (GAM) are constructed. Then, the multi-channel output feature information of the data samples is simultaneously input into the Local Attention Module (LAM) and the Global Attention Module (GAM). Finally, the outputs of the Local Attention Module (LAM) and the Global Attention Module (GAM) are concatenated through a concatenation layer.
[0038] The Local Attention Module (LAM) transforms the input using either pointwise convolution or dilation convolution to generate different representations of the input and maps them to a subspace defined as the "attention head." Specifically, the dilation convolution layer maps local features of the input to generate query and key vectors; the pointwise convolution layer maps the input to generate value vectors.
[0039]
[0040] Where q j k j and v j Let represent the query vector, key vector, and value vector obtained after transforming the input, respectively. The query vector, key vector, and value vector, as different representations of the input, are mapped to a subspace, which is defined as the "attention head." Each attention head's dilated convolutional layer uses a different dilation rate d. j =(d1,d2,...,d m ,d j ∈R), used to extract local correlations between different features of the input at multiple scales, where m is the number of attention heads.
[0041] By analyzing q j and k j Perform dot product and scaling operations to obtain the attention weight matrix of each feature in the input to the other features; after normalizing the weight matrix using the softmax activation function, it is then compared with v. j Perform weighted operations:
[0042]
[0043] Where: d q d k and d v They represent q respectively j k j and v j The vector dimension, head j The output of the attention head, head j ∈{head1,head2,...,head m Finally, the head is concatenated using a concatenation layer. j The output LAM of local attention is obtained. output :
[0044] LAM output =Concatenate{head1,head2,...,headm} (6)
[0045] Furthermore, in the multi-scale features output by the multi-scale residual network module, the output of each convolutional kernel represents a channel, and each channel carries feature maps of different granularities from the data samples. The Global Attention Module (GAM) inputs the feature maps of each channel into the Global Average Pooling (GAP) layer and the Global Max Pooling (GMP) layer respectively to extract the overall and key features of each channel, serving as the global features for each channel. The GAM module focuses on the correlation between global features at the global level: it uses pointwise convolution to perform convolution transformations on the feature maps with a kernel of 1, thereby expanding the channels.
[0046] x′=Pointwise Conv(x) (7)
[0047] Where: x′ represents the output after pointwise convolution; the average and maximum values of the feature information of each channel are output by the global average pooling layer (GAP) and the global max pooling layer (GMP), serving as the global features of the channel. Through a multilayer perceptron (MLP) composed of multiple fully connected layers, the MLP integrates the global features of the channels and extracts the correlation information between channels.
[0048] x′ p =MLP(GAP(x′))+MLP(GMP(x′)) (8)
[0049] x′ p This represents the correlation information between the overall features and key features in the channels after fusion by the multilayer perceptron (MLP); subsequently, the output is upsampled to restore the data dimension of the features in order to maintain the consistency of the data structure of the input and output of the attention module.
[0050] Finally, the global feature correlation information of each channel is normalized using the softmax activation function, serving as the correlation weights between global features in each channel. These weights are then element-wise multiplied and applied to each input channel to recalibrate the importance of each channel in the global features.
[0051] GAM output =σ(Upsample(x′) p ))⊙x′ (9)
[0052] GAM output σ represents the output of the global attention module; σ represents the sigmoid activation function; ⊙ represents element-wise multiplication.
[0053] S0506: Concatenate the outputs of local attention and global attention using a concatenation layer.
[0054] Attention = Concatenate{LAM output GAM output} (10)
[0055] Attention represents the output of concatenating the outputs of local attention and global attention. The output of Attention integrates local feature information between multi-scale features and global feature information across channels, which improves the MSRNet-GLAM model's attention to the key feature information of the data samples and reduces the fitting of non-key features.
[0056] S0507: The multi-scale feature information output by the multi-scale residual network is transformed through pointwise convolutional layers to match the data dimensions of the local feature information and the cross-channel global feature information output by the Attention; the output of the multi-scale residual network after dimension transformation is residually connected with the output of the Attention, and the original multi-scale features output by the multi-scale residual network are fused into the output of the Attention; the output after residual connection is subjected to layer normalization to prevent extreme values from appearing in the output after residual connection.
[0057] S0508: The output after layer normalization is integrated with the original multi-scale features, local feature information, and global feature information through a multilayer perceptron (MLP) with the softmax activation function, and outputs the probability distribution of each sample belonging to each category. These probability distributions will be used as the final output of the MSRNet-GLAM model.
[0058] In a further preferred embodiment of the above scheme, the multi-scale residual network consists of multiple layers of convolutional neural networks with different kernel sizes. The process of multi-scale learning through the multi-scale residual network includes the following steps:
[0059] First, let the kernel size be k, then the first layer is a pointwise convolutional layer with k=1, followed by 3 convolutional layer branches. In the multi-scale residual network, the first layer of pointwise convolutional layers is used to adjust the dimension of the features. The dimension-adjusted features are then input into each convolutional network branch to extract multi-scale information from the deep and shallow layers of the high-dimensional features.
[0060] Secondly, let the number of convolutional layers in the three stacked branches be 1, 3, and 1 respectively; where the single-layer pointwise convolutional layer with k=1 in the first branch is used to extract shallow input features; the three convolutional layers in the second branch with k=3, 5, and 3 respectively are used to extract deep input features; and the single-layer convolutional layer with k=3 in the third branch is used to extract middle-layer input features.
[0061] Finally, the outputs of the three convolutional layer branches are superimposed and input into another pointwise convolutional layer, and residual connections are made with the output of the first pointwise convolutional layer. All of the above convolutional layers contain one batch normalization layer to increase the stability of the MSRNet-GLAM model training by normalizing the input.
[0062] In a further preferred embodiment of the above scheme, step S05, inputting the data samples of the training set into the MSRNet-GLAM model for training includes performing multiple rounds of training optimization on the MSRNet-GLAM model using an improved focus loss function (IFL), with the minimization of the loss function output as the optimization objective. In each round of optimization training, the detection results of the MSRNet-GLAM model on the data samples are output. This invention constructs an improved focus loss function to measure the difference between the detection results of these data samples and the actual labels; with the minimization of the loss function output as the optimization objective, it gradually reduces the difference between the MSRNet-GLAM model output and the actual labels through multiple rounds of training, thereby optimizing the detection performance of the MSRNet-GLAM model.
[0063] A further preferred embodiment of the above scheme, which utilizes the improved focus loss function IFL to perform multi-round training and optimization of the MSRNet-GLAM model and performs intrusion detection on the train communication network traffic data to be detected, specifically includes the following sub-steps:
[0064] S0501: Input the training set data samples into the MSRNet-GLAM model for training. In each training round, output the detection results of the MSRNet-GLAM model on the data samples. Construct an improved focal loss function (IFL). The loss function uses minimizing the difference between the detection results of the MSRNet-GLAM model and the actual values as the optimization objective to optimize the training process of the MSRNet-GLAM model. To address the problem that some hidden intrusion traffic has high feature similarity and blurred boundaries with normal traffic, making it easy to confuse these hidden intrusion traffic with normal traffic during detection, this improved focal loss function IFL introduces a class balance factor α. tThe focus parameter γ is used to reduce the loss weight of easily classified samples, guiding the intrusion detection model to focus on difficult-to-classify minority class samples, thereby alleviating the problem of difficult classification of minority samples caused by class imbalance.
[0065] S0510: A similarity coefficient ω is introduced to increase the weight of easily confused samples, thereby alleviating the problem of blurred boundaries between some intrusion traffic and normal traffic, and satisfying the following:
[0066]
[0067] The MSRNet-GLAM model outputs the probability of each sample belonging to each class through the softmax activation function. These outputs are represented as probability distributions P, where P = {p1, p2, ..., p...} n}, where n is the number of categories. Where p j p represents the probability of the true class in probability distribution P; i,max This represents the maximum probability among the error categories. Let p... i,max and p j The relative distance is used as a measure of the degree of confusion in the classification results. α t γ represents the category balancing factor, used to adjust the weights between different categories; γ represents the adjustable focusing parameter, used to smoothly adjust the contribution of easily classified samples to the loss function.
[0068] S0511: Different labels are used to label each type of intrusion. The labeled intrusion detection dataset is randomly shuffled and then divided into training and test sets in a 6:4 ratio.
[0069] S0512: The training process of the MSRNet-GLAM model is optimized using an improved focus loss function to obtain a trained MSRNet-GLAM model;
[0070] S0513: Input the test set into the MSRNet-GLAM model to evaluate the intrusion performance index. If the MSRNet-GLAM intrusion detection model shows good intrusion detection evaluation results on both the training set and the test set, then the MSRNet-GLAM-based intrusion detection model will be used to implement intrusion detection in the train communication network.
[0071] In a further preferred embodiment of the above scheme, the performance of the MSRNet-GLAM intrusion detection model is evaluated using intrusion performance metrics including accuracy, recall, precision, and F1 score.
[0072] In summary, because the present invention adopts the above-described technical solution, the present invention has the following beneficial technical effects:
[0073] (1) In the process of data processing, this invention proposes a time windowing data processing method to extract the dynamic characteristics of train communication network traffic within a certain time window, thereby enhancing the expressive power of the data and making it easier for the system to detect hidden intrusion traffic from the perspective of spatiotemporal characteristics.
[0074] (2) In this invention, a multi-scale residual network is constructed to alleviate the network degradation problem that occurs during the stacking of deep neural networks and to promote the flow of gradients in deep networks.
[0075] (3) In this invention, an attention mechanism based on global and local feature attention is proposed, which enhances the MSRNet-GLAM model’s ability to capture key information at the local and global levels and reduces the interference of redundant feature information on the training of the MSRNet-GLAM model.
[0076] (4) In this invention, by constructing an improved focus loss function, the intrusion detection model’s ability to detect difficult-to-classify samples is enhanced, and the intrusion detection accuracy of the MSRNet-GLAM model is improved. Attached Figure Description
[0077] Figure 1 This is a schematic diagram of the train communication network intrusion detection method described in this invention.
[0078] Figure 2 This is a schematic diagram of the train communication network topology described in this invention.
[0079] Figure 3 This is a schematic diagram of time windowing for train communication network traffic data provided by the present invention.
[0080] Figure 4 This is a schematic diagram of the overall structure of the train communication network intrusion detection model provided by the present invention.
[0081] Figure 5 This is a schematic diagram of the multi-scale residual network module in the MSRNet-GLAM model for train communication network intrusion detection provided by the present invention.
[0082] Figure 6 This is a schematic diagram of the global and local attention mechanism module provided by the present invention. Detailed Implementation
[0083] To make the objectives, technical solutions, and advantages of this invention clearer, preferred embodiments will be listed below with reference to the accompanying drawings to provide a clear and complete description of the invention. However, it should be noted that the examples in the specification are preferred examples, not all examples. Furthermore, many details listed in the specification are merely to provide the reader with a deep understanding of the key issues of this invention; the invention can be implemented even without these specific details.
[0084] This invention provides a train communication network intrusion detection method based on multi-scale residual networks and global and local attention mechanisms. See the flowchart below. Figure 1 The method includes:
[0085] S01: Define the performance parameters of switching nodes, terminal devices, and links in the train communication network topology model; the train communication network topology consists of an interconnected train backbone network and a vehicle marshalling network, with each vehicle marshalling network's switching node connected to its subordinate terminal devices (EDs) for data transmission; the definition of the switching node mainly includes the number of switching nodes, the number of switching node ports, and IP address configuration; the definition of the terminal device mainly includes the number of terminal nodes and the IP address configuration of the terminal nodes; the definition of the link parameters mainly includes the number of links, link bandwidth, and link latency; the definition of performance parameters specifically includes the following sub-steps:
[0086] S0101: Establish the train communication network topology, which interconnects the Ethernet Train Backbone Node (ETBN) of the Train Backbone Network (ETB) with the Ethernet Consist Network Node (ECNN) of the Escherichial Train Network (ECN), forming a two-layer network structure. Each ECNN connects to its subordinate terminal devices (ED) and is responsible for data transmission. See the schematic diagram of the train communication network topology. Figure 2 ;
[0087] S0102: Define the performance parameters of switching nodes, terminal devices, and links in the train communication network topology model; the definition of the switching nodes mainly includes the number of switching nodes, the number of switching node ports, and IP address configuration; the definition of the terminal devices mainly includes the number of terminal nodes and the IP address configuration of the terminal nodes; the definition of the link parameters mainly includes the number of links, link bandwidth, and link latency.
[0088] S02: Under the train communication network topology model, establish a traffic model for the mixed transmission of normal traffic and intrusion traffic. Normal traffic includes: monitoring data, process data, message data, stream data, and best effort delivery data; intrusion traffic includes: probe intrusion, denial-of-service (DoS) intrusion, ARP spoofing, and replay intrusion traffic.
[0089] S03: Collect traffic data packets in the link as the original data samples in the intrusion detection dataset. After preprocessing the original data samples by parsing, normalizing and time windowing, a data sample containing statistical features is obtained.
[0090] S04: Merge all data samples in chronological order to form an intrusion detection dataset. Randomly shuffle the dataset and divide it into a training set and a test set in a 6:4 ratio. The training set is used for training the MSRNet-GLAM model for train communication network intrusion detection, while the test set is used to evaluate the inspection performance of the MSRNet-GLAM model. The MSRNet-GLAM model mainly consists of a multi-scale residual module and global and local attention modules, which are used to learn the key feature information of the data samples at multiple scales and output intrusion detection results based on the key feature information of the data samples.
[0091] S05: The MSRNet-GLAM model is trained. In each round of training, the MSRNet-GLAM model outputs the detection results of the data samples, and the test set data samples are input into the train communication network intrusion detection MSRNet-GLAM model for intrusion detection. This is to check and evaluate the performance of the train communication network intrusion detection MSRNet-GLAM model. If it shows good intrusion detection performance, intrusion detection of the train communication network is realized. The training set data samples are input into the MSRNet-GLAM model for training, and in each round of training, the MSRNet-GLAM model outputs the detection results of the data samples. During training, an improved focus loss function (IFL) is used to minimize the difference between the detection results and actual labels output by the MSRNet-GLAM model in each training round. This difference is then used as the optimization objective to optimize the training process of the MSRNet-GLAM model. The test set is then input into the trained MSRNet-GLAM model for intrusion detection, and the performance of the intrusion detection model is evaluated using accuracy, recall, precision, and F1 score. If the MSRNet-GLAM model can identify different types of intrusion traffic on the test set and demonstrates good intrusion detection performance, then the MSRNet-GLAM model can be used to implement intrusion detection in train communication networks.
[0092] In step S02 of this embodiment of the invention, establishing a traffic model for the mixed transmission of normal traffic and intrusion traffic under the train communication network topology model specifically includes the following sub-steps:
[0093] S0201: In accordance with the International Electrotechnical Commission standard IEC 61375, the operational traffic flow of the train in normal operation is configured in the topology of the train communication network as a normal sample in the data set. This includes monitoring data, process data, message data, streaming data, and best-effort delivery data. Monitoring data and process data are used to transmit critical information such as train control commands and train safety status information. Message data primarily transmits fault diagnosis information. Streaming data mainly transmits audio and video data related to train video monitoring and passenger entertainment activities. Best-effort delivery data involves data configuration and file transfer.
[0094] S0202: Configure sniffing (Probe) intrusion traffic in the train communication network. Sniffing Probe intrusion is a detection-type intrusion method. It sends request packets to the link indiscriminately through high-frequency or low-frequency IP scanning and port scanning to obtain response packets from the link. By parsing these response packets, it can detect active hosts and available ports in the network, providing feasible intrusion targets for subsequent intrusions.
[0095] S0203: Configure Denial of Service (DoS) intrusion traffic in the train communication network. DoS intrusion is a flooding-type intrusion method that consumes the resources of the target device by sending a large number of invalid data packets, thereby causing the target device to experience huge delays or even crash, and thus be unable to provide services for normal requests. Based on the type of attack packet protocol, DoS traffic based on UDP and TCP data was configured.
[0096] S0204: Configure Address Resolution Protocol Spoofing (ARP Spoofing) and Replay intrusion traffic in the train communication network; ARP spoofing and replay intrusion are man-in-the-middle intrusion methods. During data transmission, the Address Resolution Protocol (ARP) broadcasts ARP requests containing the target IP address across the network. The physical address of the target terminal is determined by the ARP reply sent by the target terminal corresponding to the target IP address.
[0097] S0205: Address Resolution Protocol (ARP) Spoofing achieves identity deception by sending fake ARP replies to the switching node, overwriting the target terminal's identity information in the switching node's cache. After receiving the data packet originally intended for the target host, the switching node will forward it to the intruder. After intercepting the data packet, the intruder can carry out a replay attack, that is, re-inject the intercepted data packet into the network with delay or repetition, thus achieving intrusion into the train communication network without cracking the information in the data packet.
[0098] In this invention, a train communication network data packet refers to the basic unit of data transmission in the train communication network link. It includes the packet's five-tuple (source IP address, destination IP address, source port number, destination port number, and transmission protocol) and header information such as data length, as well as payload information. The five-tuple information of the data packet forms a unique identifier for a data connection. The train communication network traffic data time windowing method provided by this invention extracts features from a series of data packets with the same five-tuple within a time window. A time windowing diagram is shown below. Figure 3 As shown. In this invention, the preprocessing of the original data sample, including packet parsing, normalization, and time windowing, includes the following sub-steps:
[0099] S0301: Collect traffic data packets in the link, use the network packet parsing framework Scapy to parse the information in the data packets, encode the discrete features of each data packet with integers, such as IP address and port number; and normalize the continuous features, such as data packet length.
[0100] S0302: Time windowing is applied to data packet information in the link. The set of all data streams in the link is denoted by F:
[0101] F = {f1, f2, ..., f m},f i ∈F (1)
[0102] Where i = 1, 2, ..., m, m represents the number of data streams on the link, f i Represents a single data stream, each data stream f i It can be defined as a collection of data packets:
[0103]
[0104] in Represents data stream f i The j-th data packet in the data, for j∈1,2,...,n i n i This indicates the number of packets in each data stream, and each packet... They all have a timestamp
[0105] S0303: Transfer data stream f i It is divided into a series of time windows, each containing data packets within a time interval Δt. Define f. i The k-th time window is W (i,k) :
[0106]
[0107] in d i The starting timestamp of the k-th time window, where Δt represents the duration of the time window. Within the time window, data packets Only then was it included in W (i,k) middle.
[0108] S0304: Using a time window as the smallest unit, perform independent statistical analysis on the discrete characteristics of data packets within the time window. These discrete characteristics include the number of data packets, length, transmission time, transmission rate, and identifier count. After statistical analysis, data samples containing statistical characteristics are obtained, and the specific characteristic information of these data samples is shown in Table 1.
[0109] Table 1 Sample Characteristics
[0110]
[0111] S0305: Merge all data samples in chronological order to form an intrusion detection dataset, then randomly shuffle the dataset and divide it into a training set and a test set in a 6:4 ratio. The training set is used for training the MSRNet-GLAM model for intrusion detection in train communication networks, and the test set is used for evaluating the inspection performance of the MSRNet-GLAM model for intrusion detection in train communication networks.
[0112] In this invention, Figure 4 This is a schematic diagram of the overall structure of the train communication network intrusion detection model provided by the present invention. The construction of the train communication network intrusion detection model specifically includes the following sub-steps:
[0113] S0401: The input of the MSRNet-GLAM model for train communication network intrusion detection is first processed through a multi-scale residual network to extract multi-scale features from the data samples. The multi-scale residual network consists of multiple layers of convolutional neural networks with different kernel sizes. Figure 5This is a schematic diagram of the multi-scale residual network in the train communication network intrusion detection model provided by this invention. The process of multi-scale learning through the multi-scale residual network includes the following steps: Let the kernel size be k. First, there is one pointwise convolutional layer, followed by three convolutional network branches. The number of convolutional layers stacked in these branches are 1, 3, and 1, respectively. In the multi-scale residual network, the first pointwise convolutional layer is used to adjust the dimension of the features. The dimension-adjusted features are input into each convolutional network branch to extract deep and shallow multi-scale information from the high-dimensional features. These branches have different numbers of convolutional layers and different sizes of convolutional kernels, which is conducive to extracting deep and shallow information from the high-dimensional features as multi-scale features of the output data samples. Specifically, the single pointwise convolutional layer with k=1 in the first branch is used to extract shallow input features; the three convolutional layers in the second branch have k=3, 5, and 3, respectively, to extract deep input features; and the single convolutional layer with k=3 in the third branch is used to extract mid-level input features. Then, the outputs of the three branches are stacked and fed into another pointwise convolutional layer, and residually connected to the output of the first pointwise convolutional layer. All convolutional layers include a batch normalization layer to increase the stability of the MSRNet-GLAM model training by normalizing the input.
[0114] S0402: The input to the MSRNet-GLAM model for train communication network intrusion detection is processed by a multi-scale residual module to extract multi-scale feature information. This multi-scale feature information is then input into global and local attention modules to enhance the MSRNet-GLAM model's ability to capture key information between features and channels.
[0115] First, construct the Local Attention Module (LAM) and the Global Attention Module (GAM). Figure 6 This is a schematic diagram of the global and local attention mechanism modules provided by the present invention. The multi-channel output feature information of the data samples is simultaneously input into the local attention module (LAM) and the global attention module (GAM). Finally, the outputs of the local attention module (LAM) and the global attention module (GAM) are concatenated by a concatenation layer.
[0116] The Local Attention Module (LAM) transforms the input using either pointwise convolution or dilation convolution. Specifically, it maps local features of the input through dilation convolution layers to generate query and key vectors; it maps the input through pointwise convolution to generate query and key vectors, and then maps the input again through pointwise convolution to generate value vectors.
[0117]
[0118] Where q j k j and v j Let represent the query vector, key vector, and value vector obtained after transforming the input, respectively. The query vector, key vector, and value vector, as different representations of the input, are mapped to a subspace, which is called the "attention head." Each attention head's dilated convolutional layer uses a different dilation rate d. j =(d1,d2,...,d m ,d j ∈R), used to extract local correlations between different features in the input at multiple scales, where m is the number of attention heads. The number of attention heads m is chosen to be 4, and the dilation rate d of the dilated convolutional layer in each attention head is... j The values were selected as 1, 2, 3, and 5 respectively.
[0119] By analyzing q j and k j Perform dot product and scaling operations to obtain the attention weight matrix of each feature in the data sample to the other features; after normalizing the weight matrix using the softmax activation function, it is then compared with v. j Perform weighted operations:
[0120]
[0121] Where: d q d k and d v They represent q respectively j k j and v j The vector dimension, head j The output of the attention head, head j ∈{head1,head2,...,head m Attention heads are concatenated using a concatenate layer. j The output of LAM (Local Attention) is obtained from the output of the local attention mechanism. output :
[0122] LAM output =Concatenate{head1,head2,...,head m} (6)
[0123] In the multi-scale features output by the multi-scale residual network module, the output of each convolutional kernel represents a channel, and each channel carries feature maps of different granularities from the data samples. The Global Attention Module (GAM) inputs the feature maps of each channel into the Global Average Pooling (GAP) layer and the Global Max Pooling (GMP) layer respectively to extract the overall and key features of each channel, and uses the overall and key features as the global features of each channel. The Global Attention Module (GAM) focuses on the correlation between global features at the global level: it uses pointwise convolution to perform convolution transformations on the feature maps with a kernel of 1 to expand the channels.
[0124] x′=Pointwise Conv(x) (7)
[0125] Where: x′ represents the output after pointwise convolution; the global average pooling layer (GAP) and the global max pooling layer (GMP) output the average and maximum values of the feature information of each channel, respectively, as the global features of the channel. Through a multilayer perceptron (MLP) composed of multiple fully connected layers, the MLP integrates the global features of the channels and extracts the correlation information between channels.
[0126] x′ p =MLP(GAP(x′))+MLP(GMP(x′)) (8)
[0127] x′ p This represents the correlation information between channels of the overall features and key features after fusion by the Multilayer Perceptron (MLP). Subsequently, the output is upsampled to restore the data dimensionality of the features, maintaining the consistency of the data structure before and after attention module processing. Finally, the global feature correlation information of each channel is normalized using the softmax activation function, serving as the correlation weights between global features in each channel. These weights are then element-wise multiplied and applied to each input channel to recalibrate the importance of each channel in the global features.
[0128] GAM output =σ(Upsample(x′) p ))⊙x′ (9)
[0129] GAM output σ represents the output of the Global Attention Module (GAM); σ represents the Sigmoid activation function; ⊙ represents element-wise multiplication.
[0130] S0403: Concatenate the outputs of the Local Attention Module (LAM) and the Global Attention Module (GAM) using a concatenation layer.
[0131] Attention = Concatenate{LAM output GAM output} (10)
[0132] Attention represents the output of concatenating the outputs of local attention and global attention. The output of Attention integrates local feature information between multi-scale features and global feature information across channels, which improves the MSRNet-GLAM model's attention to the key feature information of the data samples and reduces the fitting of non-key features.
[0133] S0404: The multi-scale feature information output from the multi-scale residual network in S0401 is transformed into data dimensions through pointwise convolutional layers to match the data dimensions of the Attention output. The multi-scale feature information output from the dimension-transformed multi-scale residual network is then residually connected to the Attention output, fusing the original multi-scale features from the multi-scale residual module output into the Attention output. Layer normalization is applied to the output after the residual connection to prevent extreme values. The normalized output is then passed through a multilayer perceptron (MLP) with a softmax activation function to integrate the original multi-scale features, local features, and global features, outputting the probability distribution of each sample belonging to each category. The category probability distribution of each data sample will be the final output of the MSRNet-GLAM model.
[0134] In this invention, inputting training set data samples into the MSRNet-GLAM model for training includes performing multiple rounds of training optimization on the MSRNet-GLAM model using the improved focus loss function IFL, with the optimization objective being the minimization of the loss function output. In each round of optimization training, the detection results of the MSRNet-GLAM model on the data samples are output. The multiple rounds of training optimization of the MSRNet-GLAM model using the improved focus loss function IFL specifically include the following sub-steps:
[0135] S0501: An improved focal loss function (IFL) is constructed. The focal loss function aims to minimize the difference between the detection results of the MSRNet-GLAM model and the actual values, thus optimizing the training process of the MSRNet-GLAM model. Some covert intrusion traffic has high feature similarity and blurred boundaries with normal traffic, making it easy to confuse with normal traffic and difficult to detect. To solve the problem of detecting covert intrusion traffic, the improved focal loss function IFL introduces a class balance factor α. t The focus parameter γ is used to reduce the loss weight of easily classified samples, guiding the intrusion detection model to focus on difficult-to-classify minority class samples, thereby alleviating the problem of difficult classification of minority samples caused by class imbalance; and a similarity coefficient ω is introduced to increase the weight of easily confused samples, so as to alleviate the problem of blurred boundaries between some intrusion traffic and normal traffic.
[0136]
[0137] The MSRNet-GLAM model outputs the probability of each sample belonging to each class through the softmax activation function. These outputs are represented as probability distributions P, where P = {p1, p2, ..., p...} n}, where n is the number of categories. Where p j p represents the probability of the true class in probability distribution P; i,max This represents the maximum probability among the error categories. Let p... i,max and p j The relative distance is used as a measure of the degree of confusion in the classification results. α t α represents the class balancing factor, used to adjust the weights between different classes; γ represents an adjustable focusing parameter, used to smoothly adjust the contribution of easily classified samples to the loss function. The value of the focusing parameter γ is chosen to be 2; the class balancing factor α... t Based on the degree of class imbalance after normalization, adjust the weight of each class from the base values:
[0138]
[0139] Where N i N represents the number of samples in the i-th class. max and N min N represents the maximum and minimum number of samples in each category. i Let ' represent the normalized value of the number of samples in class i; let the class balance factor be... Where n is the number of categories, i∈{1,2,...,n}, Let be the class weight of class i, and take . As Let the base value be 1-Ni As a compensation factor for the degree of imbalance, it is used to adjust the weight of each category on the base value.
[0140] S0502: Different labels are used to label each intrusion type. The labeled intrusion detection dataset is randomly shuffled and then divided into training and test sets in a 6:4 ratio. The MSRNet-GLAM model is trained and optimized using the training set and an improved focus loss function to obtain a trained MSRNet-GLAM model. The test set is input into the MSRNet-GLAM model to evaluate intrusion performance metrics. If the MSRNet-GLAM intrusion detection model shows good intrusion detection evaluation results on both the training and test sets, then the MSRNet-GLAM-based intrusion detection model will be used to implement intrusion detection in train communication networks. The intrusion performance metrics are mainly evaluated using accuracy, recall, precision, and F1 score. The calculation process for each evaluation metric is as follows:
[0141]
[0142] In this evaluation metric, TP (True Positive) represents the number of true positives, i.e., the number of correctly predicted positives. TN (True Negative) represents the number of true negatives, i.e., the number of correctly predicted negatives. FP (False Positive) represents the number of false positives, i.e., the number of incorrectly predicted positives. FN (False Negative) represents the number of false negatives, i.e., the number of incorrectly predicted negatives. Among these metrics, precision focuses on overall classification accuracy but can be misleading in imbalanced classes. Precision focuses on the false positive rate; high precision is suitable for scenarios where false positives need to be reduced. Recall focuses on the coverage of correct class samples; high recall is suitable for scenarios where false negatives need to be minimized. In intrusion detection of train communication networks, both false negatives and excessive false positives should be avoided. False negatives may lead to abnormal traffic intrusions, while excessive false positives result in wasted resources and alarm fatigue. The F1 score is the harmonic mean of precision and recall, comprehensively considering the accuracy and coverage of the classifier. If the MSRNet-GLAM model performs well on both the training and test sets in terms of accuracy, recall, precision, and F1 score, then the intrusion detection model based on MSRNet-GLAM will be used to implement intrusion detection in train communication networks.
[0143] For step S05 above, normal traffic and five types of intrusion traffic are labeled. Specifically, normal traffic is labeled as 0, UDP DoS intrusion traffic is labeled as 1, SYN DoS intrusion traffic is labeled as 2, ARP Spoofing intrusion traffic is labeled as 3, Replay intrusion traffic is labeled as 4, and Probe intrusion traffic is labeled as 5. The category distribution of the dataset after data preprocessing is shown in Table 2.
[0144] Table 2. Class Distribution of Dataset Samples
[0145]
[0146] The dataset was randomly shuffled and divided into training and test sets in a 6:4 ratio. To verify the effectiveness of the proposed train communication network intrusion detection method, it was compared with existing deep learning methods, namely AlexNet-GRU, CNN-LSTM, Temporal CNN, TGA, and Res-TranBiLSTM. Furthermore, to verify the effectiveness of the proposed improved focus loss function IFL, the intrusion detection method (MSRNet-GLAM-IFL) configured with the improved focus loss function IFL was compared with intrusion detection methods configured with Focal Loss and cross-entropy loss functions (With-CE and With-FL). Table 3 shows the performance comparison of each deep learning intrusion detection model in train communication network intrusion detection.
[0147] Table 3 Performance Comparison of Different Methods in Train Communication Network Intrusion Detection
[0148]
[0149] As can be seen from Table 3, in the process of train communication network intrusion detection, the MSRNet-GLAM-IFL model proposed in this invention has the highest detection accuracy, precision, recall and F1 score, which are 99.51%, 98.98%, 99.54% and 99.26% respectively, all of which are better than other deep learning-based intrusion detection models used for comparison.
[0150] The performance comparison of different methods in train communication network intrusion detection verifies that the train communication network intrusion detection method proposed in this invention, based on multi-scale residual networks and global and local attention mechanisms, has superiority. It can effectively detect train communication network intrusion traffic, and its accuracy meets the requirements of practical applications.
[0151] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. It should be noted that for those skilled in the art, any improvements, equivalent substitutions, and modifications made without departing from the principles of the present invention should also be considered within the scope of protection of the present invention.
Claims
1. A method for intrusion detection in train communication networks based on multi-scale residual networks, characterized in that: The intrusion detection method includes the following steps: S01: Establish a train communication network topology model and define the performance parameters of the switching nodes, terminal equipment, and links in the train communication network topology model; S02: Under the train communication network topology model, establish a traffic model for the mixed transmission of normal traffic and intrusion traffic; S03: Collect traffic data packets in the link as the original data samples in the intrusion detection dataset. After preprocessing the original data samples by parsing, normalizing and time windowing, a data sample containing statistical features is obtained. S04: Merge all data samples in chronological order to form an intrusion detection dataset, and randomly shuffle the dataset to divide it into a training set and a test set in a 6:4 ratio; S05: Construct the MSRNet-GLAM model for train communication network intrusion detection. Input the data samples of the training set into the MSRNet-GLAM model for training. In each round of training, the MSRNet-GLAM model outputs the detection results of the data samples. Input the data samples of the test set into the MSRNet-GLAM model for intrusion detection. Check and evaluate the performance of the MSRNet-GLAM model for train communication network intrusion detection. If it shows good intrusion detection performance, then the intrusion detection of the train communication network is realized. In step S05, the data samples of the training set are input into the MSRNet-GLAM model for training, which includes using the improved focus loss function IFL to perform multiple rounds of training optimization on the MSRNet-GLAM model, with the minimization of the loss function output as the optimization objective, and outputting the detection results of the MSRNet-GLAM model on the data samples in each round of optimization training. The optimization of the MSRNet-GLAM model through multiple rounds of training using the improved focus loss function IFL specifically includes the following sub-steps: S0509: Construct an improved focus loss function, which optimizes the difference between the detection results of the MSRNet-GLAM model and the actual values. The improved focus loss function IFL is introduced into the class balance factor. and focus parameters This reduces the loss weight of easily classified samples, guiding the MSRNet-GLAM model during training to focus on difficult-to-classify minority class samples, thereby alleviating the problem of difficult classification of minority samples caused by class imbalance. S0510: Introducing the similarity coefficient This is used to increase the weight of easily confused samples to alleviate the problem of blurred boundaries between some intrusive traffic and normal traffic, and satisfies: ,(11); The MSRNet-GLAM model uses an activation function. Output the probability that each sample belongs to each category; these outputs are represented as a probability distribution. , ,in Number of categories; Represents probability distribution The probability corresponding to the true category in the middle; This represents the maximum probability among the error categories; Indicates the category balance factor; Indicates adjustable focusing parameters; S0511: Different labels are used to label each type of intrusion. The labeled intrusion detection dataset is randomly shuffled and then divided into training and test sets in a 6:4 ratio. S0512: The training process of the MSRNet-GLAM model is optimized using an improved focus loss function to obtain a trained MSRNet-GLAM model; S0513: Input the test set into the MSRNet-GLAM model to evaluate the intrusion performance index. If the MSRNet-GLAM intrusion detection model shows good intrusion detection evaluation results on both the training set and the test set, then the MSRNet-GLAM-based intrusion detection model will be used to implement intrusion detection in the train communication network.
2. The train communication network intrusion detection method based on multi-scale residual networks according to claim 1, characterized in that: In step S01, the established train communication network topology consists of an interconnected train backbone network and a vehicle formation network. Each vehicle formation network's switching node is connected to its subordinate terminal equipment (ED) and is responsible for data transmission. The definition of the switching node includes the number of switching nodes, the number of switching node ports, and IP address configuration. The definition of the terminal equipment includes the number of terminal nodes and the IP address configuration of the terminal nodes. The definition of the link parameters includes the number of links, link bandwidth, and link latency.
3. The train communication network intrusion detection method based on multi-scale residual networks according to claim 1, characterized in that: In step S02, the normal traffic includes monitoring data, process data, message data, stream data, and maximum effort delivery data. Intrusion traffic includes: sniffing intrusion traffic, denial-of-service traffic, Address Resolution Protocol (ARP) spoofing traffic, and replay intrusion traffic. Establishing a traffic model that mixes normal traffic and intrusion traffic involves the following sub-steps: S0201: Configure the service traffic for normal train operation in the train communication network as normal samples in the normal traffic data set; S0202: Configure sniffing intrusion traffic in the train communication network. Sniffing intrusion is a detection-type intrusion method. It sends request packets to the link indiscriminately through IP scanning and port scanning, obtains response packets from the link, and then parses these response packets to explore active hosts and available ports in the network, providing feasible intrusion targets for subsequent intrusions. S0203: Configure denial-of-service intrusion traffic in the train communication network. Denial-of-service intrusion (DoS) is a flooding-type intrusion method that consumes the resources of the target device by sending a large number of invalid data packets, causing the target device to experience huge delays or even crash, thus making it unable to provide services for normal requests. S0204: Configure Address Resolution Protocol (ARP) spoofing and replay intrusion traffic in the train communication network; S0205: Address Resolution Protocol (ARP) spoofing traffic achieves identity deception by sending false ARP replies to the switching node, overwriting the target terminal's identity information in the switching node's cache; after receiving the data packet originally intended for the target host, the switching node forwards it to the intruder. The intruder intercepts the data packet and performs a replay attack, that is, re-injecting the intercepted data packet into the network with delay or repetition, thus achieving intrusion into the train communication network without cracking the information in the data packet.
4. The train communication network intrusion detection method based on multi-scale residual networks according to claim 1, characterized in that: In step S03, the preprocessing of the original data sample, including packet parsing, normalization, and time windowing, includes the following sub-steps: S0301: Collect traffic data packets in the link, use a network data packet parsing program framework to parse the information in the data packets, and encode the discrete features of each data packet with integers, including IP address and port number; then normalize the continuous features. S0302: Time windowing is applied to the data packet information in the link, using the set of all data streams in the link. express: ,(1); in , Indicates the number of data streams on the link. Represents a single data stream, each data stream Defined as a collection containing multiple data packets: ,(2); in Represents data stream The first in One data packet, for , This indicates the number of packets in each data stream, and each packet... They all have a timestamp ; S0303: Transfer data stream Divide into a series of time windows, such that each time window contains a time interval. Redefine the data packets within. The Middle The time window is : ,(3); in express The Middle The start timestamp of each time window Indicates the duration of the time window, when the timestamp Within the time window, data packets Included middle; S0304: Using a time window as the smallest unit, perform independent statistical analysis on the discrete characteristics of data packets within the time window. These discrete characteristics include the number of data packets, length, transmission time, transmission rate, and identifier counting statistical analysis to obtain a data sample containing statistical characteristics. S0305: Merge all data samples in chronological order to form an intrusion detection dataset, and randomly shuffle the dataset, dividing it into a training set and a test set in a 6:4 ratio. The training set is used for training the MSRNet-GLAM model for intrusion detection in train communication networks, and the test set is used for evaluating the inspection performance of the MSRNet-GLAM model for intrusion detection in train communication networks.
5. The train communication network intrusion detection method based on multi-scale residual networks according to claim 1, characterized in that: In step S05, constructing the MSRNet-GLAM model for train communication network intrusion detection specifically includes the following sub-steps: S0501: As the input to the MSRNet-GLAM model for intrusion detection in train communication networks, the data samples in the training set are first subjected to multi-scale learning through a multi-scale residual network to obtain multi-scale feature information of the data samples. S0502: Input the multi-scale feature information of the data samples into the global and local attention modules respectively for weighted processing to enhance the attention to important information in the multi-scale feature information.
6. The train communication network intrusion detection method based on multi-scale residual networks according to claim 5, characterized in that: The multi-scale feature information of the data samples is input into the global and local attention modules for weighted processing, including the following steps: S0503, construct a local attention module and a global attention module, then input the multi-scale feature information of the data sample into the local attention module and the global attention module at the same time, and then concatenate the outputs of the local attention module and the global attention module through the concatenation layer; The Local Attention Module (LAM) first transforms the input through a pointwise convolutional network or a dilated convolutional network to generate different representations of the input and map them to a subspace. In S0504, within each attention head, local feature mapping of the input is performed through dilated convolutional layers to generate query vectors and key vectors. Then, pointwise convolutions are used to map the input to generate value vectors. ,(4); in , and These represent the query vector, key vector, and value vector obtained after transforming the input in this attention head, respectively. The dilated convolutional layers in each attention head use different dilation rates. It is used to extract local correlations between different features at multiple scales, where Indicates the number of attention heads; Through the and Perform dot product and scaling operations to obtain the attention weight matrix of each feature in the data sample to the other features; then apply the activation function. After normalizing the attention weight matrix, and with Perform weighted operations: ,(5); in: , and They represent , and vector dimension, This indicates the output of the attention head. Attention heads are spliced together through the concatenation layer. The output of local attention is obtained. : ,(6); Furthermore, the Global Attention Module (GAM) includes a global average pooling layer and a global max pooling layer. In the output of the multi-scale residual network module, the output of each convolutional kernel represents a channel. Each channel carries feature maps of different granularities from the data samples. The GAM inputs the feature maps of each channel into the global average pooling layer and the global max pooling layer respectively, and uses the outputs of the global average pooling layer and the global max pooling layer as the overall and key features of each channel, i.e., the global features of each channel. The GAM focuses on the correlation of global features between channels at the global level, and uses pointwise convolution to perform convolution transformations with a kernel of 1 on the feature maps to expand the channels, specifically satisfying: ,(7); in: This represents the output after pointwise convolution; the average and maximum values of the feature information of each channel are output after global average pooling (GAP) and global max pooling (GMP) layers, serving as the global features of the channel; a multilayer perceptron composed of multiple fully connected layers integrates the global features of the channels, extracting the correlation information between channels: ,(8); This represents the correlation information between the overall features and key features in the channels after integration by the multilayer perceptron; subsequently, the output is upsampled to restore the data dimensionality of the features. S0505, through the activation function The correlation information between channels and the global features is normalized and used as the correlation weight between the global features in each channel. These weights are then weighted element-wise and applied to each input channel to recalibrate the importance of each channel in the global features. ,(9); in This represents the output of the Global Attention Module (GAM). This represents the Sigmoid activation function; Indicates element-wise multiplication; S0506: Concatenate the outputs of local attention and global attention using a concatenation layer, where the concatenation satisfies: ,(10); in This represents the output after concatenating the outputs of local attention and global attention. The output integrates local feature information between multi-scale features and global feature information across channels; S0507: The multi-scale feature information output by the multi-scale residual network is transformed into data dimensions through pointwise convolutional layers to match... Output data dimensions; combine the multi-scale feature information output by the multi-scale residual network after dimension transformation with... The output is subjected to residual concatenation, and the original multi-scale features output by the multi-scale residual module are fused into... In the output, the output after residual connection is subjected to layer normalization; S0508: The output after layer normalization integrates the original multi-scale features, local feature information and global feature information through a multilayer perceptron with the activation function Softmax, and outputs the probability distribution of each sample belonging to each category. That is, the category probability distribution of each data sample will be used as the final output of the MSRNet-GLAM model.
7. A train communication network intrusion detection method based on a multi-scale residual network according to claim 5 or 6, characterized in that: The multi-scale residual network consists of multiple layers of convolutional neural networks with different kernel sizes. The process of multi-scale learning through the multi-scale residual network includes the following steps: First, let the kernel size be k, then the first layer is a pointwise convolutional layer with k=1, followed by 3 convolutional layer branches. In the multi-scale residual network, the first layer of pointwise convolutional layers is used to adjust the dimension of the features. The dimension-adjusted features are then input into each convolutional network branch to extract multi-scale information from the deep and shallow layers of the high-dimensional features. Secondly, let the number of convolutional layers in the three stacked branches be 1, 3, and 1 respectively; where the single-layer pointwise convolutional layer with k=1 in the first branch is used to extract shallow input features; the three convolutional layers in the second branch with k=3, 5, and 3 respectively are used to extract deep input features; and the single-layer convolutional layer with k=3 in the third branch is used to extract middle-layer input features. Finally, the outputs of the three convolutional layer branches are superimposed and input into another pointwise convolutional layer, and residual connections are made with the output of the first pointwise convolutional layer. All of the above convolutional layers contain one batch normalization layer to increase the stability of the MSRNet-GLAM model training by normalizing the input.
8. The train communication network intrusion detection method based on multi-scale residual networks according to claim 1, characterized in that: The performance of the MSRNet-GLAM intrusion detection model was evaluated using intrusion performance metrics, including accuracy, recall, precision, and F1 score.