A network intrusion detection method based on double-layer BiLSTM knowledge distillation

By optimizing network intrusion detection through a two-layer BiLSTM knowledge distillation and time attention mechanism, the high computational cost and low detection accuracy of existing methods are solved, achieving efficient and real-time intrusion detection results.

CN122137655APending Publication Date: 2026-06-02KUNMING UNIV OF SCI & TECH

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
KUNMING UNIV OF SCI & TECH
Filing Date
2026-03-17
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing BiLSTM-based network intrusion detection methods suffer from high computational overhead, long inference latency, low utilization of temporal features, and degradation of detection accuracy under extremely imbalanced data.

Method used

We employ a two-layer BiLSTM knowledge distillation method, combining a time attention mechanism with a distillation training approach that integrates soft and hard labels. Through knowledge transfer between the two-layer BiLSTM teacher and student models, we optimize feature extraction and classification, reduce the number of model parameters, and improve detection efficiency.

Benefits of technology

It achieved a reduction of approximately 4.9 times in the number of model parameters, a reduction in inference latency to 12ms, and an improvement in detection accuracy to 99.86% and 99.32%, respectively. It meets the real-time detection requirements in resource-constrained environments and significantly improves the accuracy of identifying variant attacks and zero-day attacks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122137655A_ABST
    Figure CN122137655A_ABST
Patent Text Reader

Abstract

This invention relates to a network intrusion detection method based on bilayer BiLSTM knowledge distillation, belonging to the fields of deep learning and network security technology. The invention includes the following steps: First, a bilayer bidirectional long short-term memory (BiLSTM) network is used to extract temporal features of network traffic data to capture bidirectional contextual information in traffic packets. Second, a knowledge distillation architecture is introduced, using a pre-trained teacher model to pass soft-label knowledge to a lightweight student model, thereby compressing model parameters and enhancing the model's generalization performance in imbalanced data scenarios. Furthermore, by combining an attention mechanism to dynamically weight key spatial features, the feature extraction process is optimized, and deep fusion of multi-dimensional features is achieved through a fully connected layer. This invention significantly reduces system computational overhead and storage requirements while effectively improving the classification accuracy and stability of the lightweight detection network.
Need to check novelty before this filing date? Find Prior Art