Intelligent intrusion detection method and device based on bidirectional adaptive fusion and deformable attention
By introducing high-order deformable attention module and bidirectional adaptive fusion module in the RT-DETR model, the problem of insufficient detection performance in complex industrial scenarios is solved, and higher detection accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202510305402.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-07-01
AI Technical Summary
Traditional intrusion detection methods are difficult to provide reliable security guarantees in complex industrial scenarios, are susceptible to environmental interference, resulting in missed and missed detection, and have low small object detection performance in complex scenarios.
Using intelligent intrusion detection methods based on bidirectional adaptive fusion and deformable attention, we replace the BasicBlock module in the RT-DETR model as a high-order deformable attention module, and use the bidirectional adaptive fusion module instead of the Concat module to improve the detection accuracy and robustness of the model in complex environments.
It improves the accuracy and robustness of intrusion detection in complex scenarios, can detect and promptly remind intruders, and provide more reliable security guarantees.
Smart Images

Figure CN120236237A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of object detection, and particularly to an intelligent intrusion detection method based on bidirectional adaptive fusion and deformable attention. Background Art
[0002] With the continuous development of industrial automation and intelligence, the safety protection of dangerous areas in factories has become an important link in ensuring production safety and the safety of personnel's lives and property. Therefore, it is very necessary to detect intrusions in dangerous areas within the factory.
[0003] Traditional intrusion detection methods mainly rely on physical isolation means (such as fences, warning signs) and simple sensor technologies (such as infrared sensors, pressure sensors, and optoelectronic induction devices). Although these methods can play a warning and protection role to a certain extent, they gradually expose many deficiencies in their complex and changeable industrial scenarios: 1. Sensors are easily interfered by environmental factors, resulting in serious degradation of performance and may even completely fail, unable to provide reliable safety protection; 2. Traditional detection methods are difficult to cope with complex scenarios. In industrial scenarios, the environment is usually complex and changeable, and traditional detection methods are easily affected by the external environment, resulting in missed detections and false detections. In addition, the detection performance of small targets in complex scenarios is usually low, further restricting practical applications.
[0004] Therefore, in order to overcome the above deficiencies, there is an urgent need for an intelligent method that can detect intrusions in dangerous areas of factories in complex scenarios. Summary of the Invention
[0005] Object of the Invention: Aiming at the above problems, an intelligent intrusion detection method and device based on bidirectional adaptive fusion and deformable attention effectively improve the accuracy and robustness of the model in complex environments by replacing the BasicBlock module in the Backbone layer of RT-DETR with a high-order deformable attention module and using a bidirectional adaptive fusion module to replace the Concat module in the Head layer of the RT-DETR model.
[0006] Technical Solution: The present invention proposes an intelligent intrusion detection method based on bidirectional adaptive fusion and deformable attention, including the following steps:
[0007] Step 1: Obtain the surveillance video of personnel intrusion in the dangerous area of the factory, and semi-automatically annotate the personnel intrusion behavior in the video image to form a personnel intrusion behavior dataset;
[0008] Step 2: Replace the BasicBlock module in the Backbone layer of the RT-DETR model with the High-Order Deformable Attention module HDA; the High-Order Deformable Attention module HDA divides the input features into two parts, each part is convolved, then processed through the deformable attention module, and finally the two parts of features are fused to obtain the final output features;
[0009] Step 3: Replace the Concat module in the Head layer of the RT-DETR model with the Bidirectional Adaptive Fusion module BAFN;
[0010] Step 4: Use the training set to train the improved RT-DETR model, and the RT-DETR model maps the input data to the output space to produce predicted results;
[0011] Step 5: Evaluate the performance of the trained improved RT-DETR model on the test set.
[0012] Furthermore, the specific method of Step 1 is as follows:
[0013] Step 1.1: Collect the monitoring videos of the dangerous areas in the factory, randomly extract multiple frames of pictures containing personnel intrusion from the videos as the dataset SData;
[0014] Step 1.2: Use LabelImg to label part of the dataset. For the target, use a rectangular label box to mark the position of the target and mark the category of the target;
[0015] Step 1.3: Train the labeled data to obtain a weight file, and use the weight file and the automatic annotation script to automatically annotate the remaining dataset, and manually correct the annotation after annotation;
[0016] Step 1.4: Randomly divide the labeled images into a training set, a validation set and a test set according to a certain proportion.
[0017] Furthermore, the specific method of the High-Order Deformable Attention module HDA in Step 2 is as follows:
[0018] Step 2.1: Input feature map Divide it into two parts according to the channel dimension, namely x1 and x2, where H represents the height of the feature map, W represents the width of the feature map, and C represents the number of channels of the feature map;
[0019] Step 2.2: Process x1 and x2 respectively through the convolutional layer to generate enhanced feature maps x1' and x'2;
[0020] Step 2.3: Process the enhanced feature map to generate a uniform point grid where H g = H / r, W g = W / r, and the coordinates of each grid point are defined in the range from (0, 0) to (H g - 1, W g - 1), and finally normalized to the interval [-1, 1]. Here, H g represents the grid height of the feature map, W g represents the grid width of the feature map, and r is a hyperparameter;
[0021] Step 2.4: Perform a linear transformation on the enhanced feature map to generate a query vector q = x1'W q , where W q is a learnable weight matrix;
[0022] Step 2.5: Input the generated query vector into the lightweight offset network θ offset to calculate the offset Δp = θ offset (q). Constrain the offset through the hyperbolic tangent function: Δp ← s × tanh(Δp), where s is a predefined scaling factor used to control the magnitude of the offset;
[0023] Step 2.6: Adjust the position of the reference point according to the offset Δp to obtain a new sampling point p + Δp;
[0024] Step 2.7: Use the adjusted sampling point p + Δp to extract features from the input feature map x1' to generate a key vector and a value vector: where φ is the bilinear interpolation sampling function, W k and W v represent the learnable weight matrices for the key and the value respectively;
[0025] Step 2.8: Input the query vector q, the key vector and the value vector into the multi - head attention mechanism. In each attention head, calculate the weighted result of the three of them: where z (m) represents the output feature of the m - th attention head, is the square root of the feature dimension, used to stabilize the calculation of the attention weights, B represents the reference point position, R is the relative position offset, and θ(.) is an activation function that normalizes the input value;
[0026] Step 2.9: Combine the output features of all attention heads as the final output feature;
[0027] Step 2.10: Execute Steps 2.3 to 2.9 for the enhanced feature map x'2;
[0028] Step 2.11: Fuse the output features generated by the two parts to obtain the output of the high-order deformable attention module.
[0029] Further, the specific method of the bidirectional adaptive fusion module BAFN in step 3 is as follows:
[0030] Step 3.1: Extract multi-scale features from the backbone network, denoted as F1, F2, F3, representing feature maps of different resolutions respectively;
[0031] Step 3.2: The high-level feature F2 is processed by a depthwise separable convolutional layer to enhance the feature representation, and then through a flattening operation, it is transformed into two tensors and where H represents the height of the feature map, W represents the width of the feature map, and C represents the number of channels of the feature map;
[0032] Step 3.3: The low-level feature F1 is processed by a depthwise separable convolutional layer to enhance the feature representation, and then through a flattening operation, it is transformed into a tensor
[0033] Step 3.4: Linear projections are respectively performed on the three tensors to generate a query vector q = W q (σ(φ(x q , w, s))), a key vector k = W k (σ(φ(x k , w, s))) and a value vector v = W v (σ(φ(x v , w, s))), where φ represents a 3×3 convolution, σ represents a reshaping function, w represents the convolution parameter, s represents the stride, W q , W k , W v represent the projection weights of different vectors;
[0034] Step 3.5: Generate an attention weight matrix by calculating the dot product between the query vector and the key vector to represent the mutual relationship between different features, and use Softmax for normalization so that the sum of the weights at each query position is 1;
[0035] Step 3.6: Use the attention weight matrix to perform weighted summation on the value vector to obtain the updated feature, and reshape the new feature into the original dimension H×W×C;
[0036] Step 3.7: Further consolidate and enhance the feature representation by passing the updated feature through a residual connection to obtain the finally output feature d2;
[0037] Step 3.8: Use Feature F3 and the generated Feature d2 as the high-level feature and low-level feature respectively, and re-execute Steps 3.2 to 3.7 to obtain the final output feature d3;
[0038] Step 3.9: Upsample Feature F3 and then perform 1×1 convolution to obtain the extended feature F3' = ReLu(BN(f 1×1 (U(F3)))), where ReLu represents the activation function, BN represents normalization, f 1×1 represents 1×1 convolution, and U() represents upsampling;
[0039] Step 3.10: Perform max pooling and average pooling operations on the obtained extended feature F3' to obtain the spatial weight matrix for capturing the spatial relationship of the features. The formula is: where Sigmoid represents the activation function, f 7×7 represents 7×7 convolution, P a represents average pooling, and P m represents max pooling;
[0040] Step 3.11: Perform a weighted operation on the spatial weight matrix and Feature F2. The formula is: where ⊙ represents element-wise multiplication, and s2 represents the finally generated feature;
[0041] Step 3.12: Use s2 and F1 as the high-level feature and low-level feature respectively, and re-execute Steps 3.9 to 3.11 to obtain the final output feature s1;
[0042] Step 3.13: Fuse the generated Feature d2 and s2, and fuse d3 and s1 to obtain the fused feature. The formula is: F = f 3×3 (ε(s, d), w, t), where F represents the final output feature, f 3×3 represents 3×3 convolution, ε represents the feature fusion for verifying the channel dimension, w represents the convolution parameter, and t represents the stride;
[0043] Step 3.14: Fuse all the fused features again to obtain the final output feature.
[0044] Furthermore, the specific method of Step 4 is as follows:
[0045] Step 4.1: Perform forward propagation through the training set, input the image into the improved RT-DETR model, and obtain the predicted result;
[0046] Step 4.2: Calculate the loss between the model prediction and the true label using cross-entropy loss;
[0047] Step 4.3: Calculate the gradient of the loss function with respect to the model parameters using the backpropagation algorithm;
[0048] Step 4.4: Update the weights of the model using the gradient descent optimization algorithm;
[0049] Step 4.5: Repeat steps 4.1 - 4.4 for multiple iterations until the model converges or reaches a previously specified number of training epochs.
[0050] The present invention also discloses an intelligent intrusion detection device based on bidirectional adaptive fusion and deformable attention, including a memory, a processor, and a computer program stored on the memory and executable on the processor. It is characterized in that when the computer program is loaded into the processor, it implements the above - mentioned intelligent intrusion detection method based on bidirectional adaptive fusion and deformable attention.
[0051] Beneficial effects:
[0052] 1. For the intelligent detection of intrusion in the dangerous areas of factories, the present invention pays more attention to improving the accuracy and robustness of intrusion detection in complex scenarios, and can perform real - time detection, and can give voice reminders to the intruders in a timely manner. The dangerous areas of factories are diverse and complex. Using the method of bidirectional adaptive fusion can extract more accurate features, thereby improving the accuracy and robustness of detection.
[0053] 2. The present invention uses an improved RT - DETR model, replaces the BasicBlock module in the Backbone layer of the model with a high - order deformable attention module, and realizes dynamic sampling of key regions by learning the offset, which can adapt to targets of different sizes and shapes.
[0054] 3. The present invention uses a bidirectional adaptive fusion module to replace the Concat module in the Head layer of the RT - DETR model, extracts global background information and saliency information through average pooling and max pooling, and can capture multi - scale context information, thereby improving the detection accuracy and robustness of the model in complex scenarios. Description of the drawings
[0055] Figure 1 is the overall flowchart of the present invention;
[0056] Figure 2 is the model structure diagram of the present invention;
[0057] Figure 3 is the working diagram of the deformable attention;
[0058] Figure 4 is the fusion process of the output features d and s. Detailed implementation manners
[0059] The present invention will be further clarified below in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. After reading the present invention, various equivalent forms of modification by those skilled in the art fall within the scope defined by the appended claims of this application.
[0060] The present invention discloses an intelligent intrusion detection method and device based on bidirectional adaptive fusion and deformable attention, which is applicable to the problem that the intrusion detection of personnel in the dangerous area of the factory is affected by the surrounding complex environment. First, obtain the monitoring video of personnel intrusion in the dangerous area of the factory, and semi-automatically annotate the personnel intrusion behavior in the video image to form a personnel intrusion behavior data set; then improve the deformable attention module to form a high-order deformable attention module, and replace the BasicBlock module in the Backbone layer of the RT-DETR model with the high-order deformable attention module; use the bidirectional adaptive fusion module to replace the Concat module in the Head of the RT-DETR model; finally, use the training set to train the improved RT-DETR model, and the model maps the input data to the output space to generate a prediction result. The specific steps are as follows:
[0061] Step 1: Obtain the monitoring video of personnel intrusion in the dangerous area of the factory, and semi-automatically annotate the personnel intrusion behavior in the video image to form a personnel intrusion behavior data set:
[0062] Step 1.1: Collect the monitoring video of the dangerous area of the factory, and randomly extract multiple frames of pictures containing personnel intrusion from the video as the data set SData.
[0063] Step 1.2: Use LabelImg to annotate part of the data set. For the target, use a rectangular label box to mark the position of the target and mark the category of the target.
[0064] Step 1.3: Use the model trained with the annotated data set and the automatic annotation script to automatically annotate the remaining data set, and manually correct the annotation after annotation.
[0065] Step 1.4: Randomly divide the labeled images into three groups according to a certain proportion: training set, validation set and test set.
[0066] Step 2: Replace the BasicBlock module in the Backbone layer of the RT-DETR model with the high-order deformable attention module HDA:
[0067] Step 2.1: Replace the BasicBlock module in the Backbone layer of the RT-DETR model with the high-order deformable attention module.
[0068] Step 2.2: Input feature map It is divided into two parts along the channel dimension, namely x1 and x2, to reduce the mutual interference between channels, where H represents the height of the feature map, W represents the width of the feature map, and C represents the number of channels of the feature map.
[0069] Step 2.3: Process x1 and x2 respectively through a convolutional layer to generate enhanced feature maps x1' and x'2.
[0070] Step 2.4: Process the enhanced feature maps to generate a uniform point grid where H g = H / r, W g = W / r, and the coordinates of each grid point are defined in the range from (0, 0) to (H g - 1, W g - 1), and finally normalized to the interval [-1, 1]. H g represents the grid height of the feature map, W g represents the grid width of the feature map, and r is a hyperparameter.
[0071] Step 2.5: Perform a linear transformation on the enhanced feature maps to generate a query vector q = x1'W q , where W q is a learnable weight matrix.
[0072] Step 2.6: Input the generated query vector into the lightweight offset network θ offset , and calculate the offset Δp = θ offset (q). To prevent the offset from being too large and affecting the sampling stability, the offset is constrained by the hyperbolic tangent function: Δp ← s × tanh(Δp), where s is a predefined scaling factor used to control the amplitude of the offset.
[0073] Step 2.7: Adjust the position of the reference point according to the offset Δp to obtain a new sampling point p + Δp.
[0074] Step 2.8: Use the adjusted sampling point p + Δp to extract features from the input feature map x1' to generate key and value vectors: where φ is a bilinear interpolation sampling function, and W k and W v represent the learnable weight matrices for the key and value respectively.
[0075] Step 2.9: Input the query vector q, the key vector and the value vector into the multi - head attention mechanism. In each attention head, calculate the weighted result of the three of them: where z (m) represents the output feature of the m-th attention head, is the square root of the feature dimension, used to stabilize the calculation of the attention weights, B represents the reference point position, R is the relative position offset, and θ(.) is an activation function that normalizes the input value.
[0076] Step 2.10: Combine the output features of all attention heads as the final output feature.
[0077] Step 2.11: Perform Steps 2.4 to 2.10 on the enhanced feature map x'2.
[0078] Step 2.12: Fuse the output features generated by the two parts to obtain the output of the high-order deformable attention module.
[0079] Step 3: Replace the Concat module in the Head layer of the RT-DETR model with a bidirectional adaptive fusion module BAFN:
[0080] Step 3.1: Extract multi-scale features from the backbone network, labeled as F1, F2, F3, representing feature maps of different resolutions respectively.
[0081] Step 3.2: Feature F2 is processed through a depthwise separable convolutional layer to enhance the feature representation, and then through a flattening operation to convert it into two tensors and where H represents the height of the feature map, W represents the width of the feature map, and C represents the number of channels of the feature map.
[0082] Step 3.3: Feature F1 is processed through a depthwise separable convolutional layer to enhance the feature representation, and then through a flattening operation to convert it into a tensor
[0083] Step 3.4: Linear projections are performed on the three tensors respectively to generate the query vector q = W q (σ(φ(x q , w, s))), the key vector k = W k (σ(φ(x k , w, s))) and the value vector v = W v (σ(φ(x v , w, s))), where φ represents a 3×3 convolution, σ represents a reshaping function, w represents the convolution parameter, s represents the stride, W q , W k , W v represent the projection weights of different vectors.
[0084] Step 3.5: Generate an attention weight matrix by calculating the dot product between the query vector and the key vector to represent the mutual relationship between different features, and normalize it using Softmax so that the sum of the weights at each query position is 1.
[0085] Step 3.6: Use the attention weight matrix to perform weighted summation on the value vector to obtain the updated feature, and reshape the new feature into the original dimension H×W×C.
[0086] Step 3.7: Pass the updated feature through a residual connection to further consolidate and enhance the feature representation, obtaining the finally output feature d2.
[0087] Step 3.8: Use the feature F3 and the generated feature d2 as the high-level feature and the low-level feature respectively, and re-execute Steps 3.2 to 3.7 to obtain the finally output feature d3.
[0088] Step 3.9: Upsample the feature F3 and then perform 1×1 convolution to obtain the extended feature F3' = ReLu(BN(f 1×1 (U(F3)))), where ReLu represents the activation function, BN represents normalization, f 1×1 represents 1×1 convolution, and U() represents upsampling.
[0089] Step 3.10: Perform max pooling and average pooling operations on the obtained extended feature F3' to obtain the spatial weight matrix for capturing the spatial relationship of the feature, and the formula is: where represents the generated spatial weight matrix, Sigmoid represents the activation function, f 7×7 represents 7×7 convolution, P a represents average pooling, and P m represents max pooling.
[0090] Step 3.11: Perform a weighted operation on the spatial weight matrix and the feature F2, and the formula is: ⊙ represents element-wise multiplication, and s2 represents the finally generated feature.
[0091] Step 3.12: Use s2 and F1 as the high-level feature and the low-level feature respectively, and re-execute Steps 3.9 to 3.11 to obtain the finally output feature s1.
[0092] Step 3.13: Fuse the generated features d2 and s2, and fuse d3 and s1 to obtain the fused feature, and the formula is: F = f 3×3 (ε(s,d),w,t), where F represents the finally output feature, f 3×3A 3×3 convolution is denoted as, ε represents the feature fusion of the verification channel dimension, w represents the convolution parameter, and t represents the stride.
[0093] Step 3.14: Fuse all the features obtained after fusion again to obtain the final output features.
[0094] Step 4: Use the training set to train the improved RT-DETR model. The RT-DETR model maps the input data to the output space to produce the predicted results:
[0095] Step 4.1: Perform forward propagation through the training set obtained by dividing the dataset. Input the image into the improved RT-DETR model to obtain the predicted results;
[0096] Step 4.2: Calculate the loss between the model prediction and the true label using the cross-entropy loss;
[0097] Step 4.3: Use the backpropagation algorithm to calculate the gradient of the loss function with respect to the model parameters;
[0098] Step 4.4: Update the weights of the model using the gradient descent optimization algorithm;
[0099] Step 4.5: Repeat Steps 4.1 - 4.4 and iterate multiple times until the model converges or reaches the previously predetermined number of training epochs.
[0100] The present invention can be combined with a computer system to form an intelligent auxiliary device for detecting the intrusion of personnel in the dangerous area of the factory. The device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the computer program is loaded into the processor, it implements the above-mentioned intelligent intrusion detection method based on bidirectional adaptive fusion and deformable attention.
[0101] The above embodiments are only for illustrating the technical concept and features of the present invention, and their purpose is to enable those familiar with this technology to understand the content of the present invention and implement it accordingly, and cannot be used to limit the protection scope of the present invention. Any equivalent transformation or modification made according to the spirit and essence of the present invention should be covered within the protection scope of the present invention.
Claims
1. An intelligent intrusion detection method based on bidirectional adaptive fusion and deformable attention, characterized in that: The steps include: Step 1: Obtain surveillance videos of personnel intrusion into dangerous areas of the factory, and semi-automatically annotate the personnel intrusion behaviors in the video images to form a personnel intrusion behavior dataset; Step 2: Use a high-order deformable attention module HDA to replace the BasicBlock module in the Backbone layer of the RT-DETR model; the high-order deformable attention module HDA divides the input features into two parts, convolves each part, and then processes it through the deformable attention module. The two parts of the features are finally fused to obtain the final output features; Step 3: Use the bidirectional adaptive fusion module BAFN to replace the Concat module of the Head layer in the RT-DETR model; Step 4: Use the training set to train the improved RT-DETR model. The RT-DETR model maps the input data to the output space and generates predicted results. Step 5: Evaluate the performance of the trained improved RT-DETR model on the test set.
2. The intelligent intrusion detection method based on bidirectional adaptive fusion and deformable attention according to claim 1 is characterized in that: The specific method of step 1 is: Step 1.1: Collect surveillance videos of dangerous areas in the factory and randomly extract multiple frames of images containing human intrusion from the video as the data set SData; Step 1.2: Use LabelImg to label part of the data set. For the target, use a rectangular label box to mark the location of the target and label the target category. Step 1.3: Train the labeled data to obtain the weight file, use the weight file and the automatic labeling script to automatically label the remaining data sets, and manually correct the labels after labeling; Step 1.4: Randomly divide the labeled images into training set, validation set and test set in proportion.
3. The intelligent intrusion detection method based on bidirectional adaptive fusion and deformable attention according to claim 1 is characterized in that: The specific method of the high-order deformable attention module HDA in step 2 is: Step 2.1: Input feature map It is divided into two parts according to the channel dimension, namely x1 and x2, where H represents the height of the feature map, W represents the width of the feature map, and C represents the number of channels of the feature map; Step 2.2: Process x1 and x2 through convolutional layers to generate enhanced feature maps x1' and x'2; Step 2.3: The enhanced feature map Processing to generate a uniform grid of points Among them, H g =H / r,W g =W / r, the coordinates of each grid point are defined between (0,0) and (H g -1,W g -1), and finally normalized to the interval [-1,1], H g Represents the grid height of the feature map, W g Represents the grid width of the feature map, r is a hyperparameter; Step 2.4: The enhanced feature map Perform linear transformation to generate query vector q = x1'W q , where W q is a learnable weight matrix; Step 2.5: Input the generated query vector into the lightweight offset network θ offset , calculate the offset Δp = θ offset (q), constrain the offset through the hyperbolic tangent function: Δp←s×tanh(Δp), where s is a predefined scaling factor used to control the magnitude of the offset; Step 2.6: Adjust the position of the reference point according to the offset Δp to obtain a new sampling point p+Δp; Step 2.7: Use the adjusted sampling point p+Δp to extract features from the input feature map x1' and generate the key vector and value vector: Among them, φ is the bilinear interpolation sampling function, W k and W v Learnable weight matrices representing keys and values respectively; Step 2.8: Substitute query vector q, key vector Sum value vector Input into the multi-head attention mechanism, and in each attention head, calculate the weighted result of the three: where z (m) represents the output feature of the mth attention head, is the square root of the feature dimension, used to stabilize the calculation of attention weights, B represents the reference point position, R is the relative position offset, and θ(.) is an activation function that normalizes the input value; Step 2.9: Combine the output features of all attention heads as the final output features; Step 2.10: Perform steps 2.3 to 2.9 for the enhanced feature map x2; Step 2.11: Fuse the output features generated by the two parts to obtain the output of the high-order deformable attention module.
4. The intelligent intrusion detection method based on bidirectional adaptive fusion and deformable attention according to claim 1 is characterized in that: The specific method of the bidirectional adaptive fusion module BAFN in step 3 is: Step 3.1: Extract multi-scale features from the backbone network, marked as F1, F2, and F3, which represent feature maps of different resolutions respectively; Step 3.2: The high-level feature F2 is processed by a depth-wise separable convolutional layer to enhance the feature representation, and then converted into two tensors through a flattening operation. and Where H represents the height of the feature map, W represents the width of the feature map, and C represents the number of channels of the feature map; Step 3.3: The underlying feature F1 is processed through a depthwise separable convolutional layer to enhance the feature representation, and then converted into a tensor through a flattening operation. Step 3.4: Perform linear projection on the three tensors to generate the query vector q = W q (σ(φ(x q ,w,s))), key vector k=W k (σ(φ(x k ,w,s))) and value vector v = W v (σ(φ(x v ,w,s))), where φ represents a 3×3 convolution, σ represents a reshaping function, w represents a convolution parameter, s represents a step size, and W q , W k , W v Represents the projection weights of different vectors; Step 3.5: Generate an attention weight matrix by calculating the dot product between the query vector and the key vector to represent the relationship between different features, and use Softmax to normalize so that the sum of the weights of each query position is 1; Step 3.6: Use the attention weight matrix to perform weighted summation on the value vector to obtain the updated features, and reshape the new features into the original dimensions H×W×C; Step 3.7: The updated features are further consolidated and enhanced through residual connections to obtain the final output feature d2; Step 3.8: Use feature F3 and generated feature d2 as high-level feature and low-level feature input respectively and re-execute steps 3.2 to 3.7 to obtain the final output feature d3; Step 3.9: Upsample feature F3 and perform 1×1 convolution to obtain extended feature F3'=ReLu(BN(f 1×1 (U(F3)))), where ReLu represents the activation function, BN represents normalization, and f 1×1 represents 1×1 convolution, U() represents upsampling; Step 3.10: Perform maximum pooling and average pooling operations on the obtained extended features F3' to obtain the spatial weight matrix M F3 , used to capture the spatial relationship of features, the formula is: M F3 =Sigmoid(f 7×7 ((P a (F3'); P m (F3')))), Sigmoid represents the activation function, f 7×7 represents a 7×7 convolution, P a represents average pooling, P m represents maximum pooling; Step 3.11: Transform the spatial weight matrix M F3 The weighted operation is performed with feature F2, and the formula is: s2 = F2-(M F3 ⊙F2), ⊙ represents element-by-element multiplication, and s2 represents the final generated feature; Step 3.12: Use s2 and F1 as high-level features and low-level features respectively and re-execute steps 3.9 to 3.11 to obtain the final output feature s1; Step 3.13: Fuse the generated features d2 and s2, and d3 and s1 to obtain the fused features. The formula is: F = f 3×3 (ε(s,d),w,t), where F represents the final output feature, f 3×3 represents 3×3 convolution, ε represents the feature fusion of the verification channel dimension, w represents the convolution parameter, and t represents the step size; Step 3.14: Fuse all the fused features again to get the final output features.
5. The intelligent intrusion detection method based on bidirectional adaptive fusion and deformable attention according to claim 1 is characterized in that: The specific method of step 4 is: Step 4.1: Forward propagate through the training set, input the image into the improved RT-DETR model, and obtain the predicted result; Step 4.2: Use cross entropy loss to calculate the loss between the model prediction and the true label; Step 4.3: Use the back-propagation algorithm to calculate the gradient of the loss function with respect to the model parameters; Step 4.4: Update the model weights using the gradient descent optimization algorithm; Step 4.5: Repeat steps 4.1-4.4 for multiple iterations until the model converges or reaches the previously predetermined training rounds.
6. An intelligent intrusion detection device based on bidirectional adaptive fusion and deformable attention, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the computer program is loaded into a processor, the intelligent intrusion detection method based on bidirectional adaptive fusion and deformable attention according to any one of claims 1 to 5 is implemented.