Threat signaling stream data generation and detection system based on generative adversarial network

Through the threat signaling flow data generation and detection system based on the generative adversarial network, the problems of shortage of threat signaling data sets and imbalance in the category of 5G communication network are solved, the identification capability of the threat detection system is improved, and the security and reliability of the communication network are ensured.

CN120074897AInactive Publication Date: 2025-05-30CHONGQING UNIV OF POSTS & TELECOMM
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510165445.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-05-30
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art does not work well in processing and classifying some types of threat stream data in 5G communication networks, resulting in poor performance of machine learning or deep packet inspection methods in threat detection.

Method used

The threat signaling flow data generation and detection system based on the generative adversarial network is adopted. Through data management, data generation, threat detection and data visualization modules, threat signaling flow samples that fit the real data distribution are generated, the data set is expanded, and the identification capabilities of the threat detection system are improved.

Benefits of technology

It enhances the diversity of threat signaling flow data, solves the problems of shortage of threat signaling data sets and category imbalance in 5G communication networks, improves the ability to identify threats of a few categories, and ensures the security and reliability of the communication network.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120074897A_ABST
    Figure CN120074897A_ABST
Patent Text Reader

Abstract

The invention requests to protect a threat signaling flow data generation and detection system based on a generative adversarial network. Comprising a data generation module which is used for helping a user to quickly train a model and call the model to generate different types of data samples required by the user through two task modes of a training algorithm model and a calling algorithm model, and analyzing and comparing threat data generation results; the threat detection module is used for identifying threat traffic and normal traffic types, comprises two tasks of detection model training and threat detection, and is used for detecting a threat signaling flow type and tracking and positioning IP and port number information of an abnormal terminal; the data visualization module is used for visually displaying the information in a chart form; and the user management module is used for managing accounts and comprises a mailbox login registration module and a mailbox password retrieving module. According to the invention, the generation and detection capability of the threat signaling stream data can be effectively improved, and the stable operation and data quality of the system are ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of communication network security, and specifically to a threat signaling flow data generation and detection system based on a generative adversarial network. The present invention also provides a usage method of the system. Background Art

[0002] With the rapid increase in the number of network elements and terminals in the 5G communication network and the continuous update of terminal device types, the security threats of the 5G core network show a trend of diversification and heterogeneity. In this case, ensuring the security and reliability of the 5G core network has become a core issue that urgently needs to be solved. Many studies have explored the security issues of the 5G core network from aspects such as the physical layer and the network layer, and various technical solutions have been proposed. Among them, the intrusion detection technology at the network layer has been proven to be an effective means to deal with threat signaling flows. However, with the development of technology and the wide application of technologies such as machine learning, training a machine learning model to identify threat traffic has become a new idea. This method can effectively detect threat signaling flows and help communication network operators monitor the network security status by classifying user network behaviors, ensuring the stability of communication services. However, due to technical limitations, some types of threat flow data are missing, resulting in poor performance of methods based on machine learning or deep packet inspection in processing and classifying this type of threat traffic. Therefore, it is particularly important to use adversarial learning technology to quantitatively characterize threat signaling flows and expand the data set by generating threat signaling flow samples to improve the recognition ability of the threat detection system for minority-class threats. Through threat signaling flow generation, the diversity of threat signaling flow data can be enhanced, and the problems of shortage and class imbalance of the current 5G communication network threat signaling data set can be solved, enabling limited threat data to generate equivalent data value, and providing a possible solution for the construction of future network threat data sets.

[0003] Compared with the patent method of the publication number CN113114673A (a network intrusion detection method and system based on a generative adversarial network), this method focuses on the generation and detection of threat signaling flow data, and the feature extraction and data preprocessing methods are different. The actual structure and optimization objective of the generative adversarial network model are also different. The design objective of this method is to cover a wide range of network intrusion detection scenarios and extract relatively general features. In specific attack types or complex network environments, it is easy to have problems of low detection accuracy. And this method designs a traffic feature extraction and data preprocessing method according to the characteristics of the 5G signaling network protocol, enabling the generative model to generate new samples that fit the real data distribution, so that the model can be used to enhance the sample latent space and expand the data set in the communication network anomaly detection system, providing more effective data support for threat signaling flow detection. Summary of the Invention

[0004] The present invention aims to solve the above problems of the prior art. A threat signaling flow data generation and detection system based on a generative adversarial network is proposed. The technical solution of the present invention is as follows:

[0005] A threat signaling flow data generation and detection system based on a generative adversarial network, which includes:

[0006] A data management module for preprocessing and managing dataset information, including data collection and upload, data cleaning, dynamically updating the data processing status and results, and statistics of the existing datasets in the system;

[0007] A data generation module for helping users quickly train models and call models to generate different types of data samples required by users through two task methods of training algorithm models and calling algorithm models, and analyzing and comparing the results of generating threat data;

[0008] A threat detection module for identifying threat traffic and normal traffic types, including two tasks of detection model training and threat detection, for detecting threat signaling flow categories, and tracking and locating the IP and port number information of abnormal terminals;

[0009] A data visualization module for visually displaying dataset statistics information, sample distribution similarity, sample structure similarity, sample data features, and anomaly detection result information in the form of charts;

[0010] A user management module for managing accounts, including an email login and registration module and an email password retrieval module.

[0011] Further, the data management module includes a data preprocessing module, a module for dynamically updating the data preprocessing status, a data preprocessing result module, a database selection module, and a database preview module; wherein,

[0012] The data preprocessing module is used to provide users with custom data preprocessing methods, including selecting feature extraction, outlier cleaning, normalization processing, and feature dimensionality reduction operations to formulate data preprocessing tasks;

[0013] The module for dynamically updating the data preprocessing status is used to display the execution progress of the data preprocessing tasks to users and prompt task completion in the form of a pop-up window;

[0014] The data preprocessing result module is used for users to preview the feature results of data preprocessing and display the comparison of the data volumes of various types of traffic before and after processing;

[0015] The database selection module is used for users to custom-select and preview the threat dataset mode, customize dataset selection, threat type selection, and attack scenario selection schemes, and query data information in the database;

[0016] The database preview module is used for the system to display the overall situation of the dataset in the custom preview mode to the user, including the country to which the attacking IP belongs, the proportion of each type of traffic in the dataset, and the quantity information of each type of traffic.

[0017] Furthermore, the data generation module includes a generation model selection and training module, a dynamic update model training status module, a generation model management module, a sample distribution similarity calculation module, a sample relationship structure calculation module, a data generation task module, a dynamic update sample generation task status module, a view all task information module, and a view generated sample information module; among them,

[0018] The generation model selection and training module is used for the user to select the type of threat data to be trained and generated, the algorithm model it depends on, and the hyperparameters of model training, and automatically obtain the start time of the current training task in the system to add a timestamp to the training task;

[0019] The dynamic update model training status module is used for the system to obtain the progress and results of the backend model training, and display the information of whether the training is successful to the user in the form of a message prompt window;

[0020] The generation model management module is used for the user to preview and delete the information of the existing models in the system, facilitating the user to manage the system model library files;

[0021] The sample distribution similarity calculation module is used for the system to calculate the overall distribution of the samples generated by the generation model and the overall distribution of the real samples, and verify the similarity between the two;

[0022] The sample relationship structure calculation module is used for the system to calculate the Pearson correlation coefficient between the internal features of the generated sample set. The calculation formula of the Pearson correlation coefficient r is as follows:

[0023]

[0024] In the formula, x i , y i respectively represent the i-th value of the sample variables x and y; respectively represent the means of the sample variables x and y. Then, the correlation coefficient of the generated samples is compared with the correlation coefficient of the features of the real sample set to verify the consistency between the internal relationships of the two features;

[0025] The data generation task module is used for the user to formulate a sample generation task, select the attack category to which the sample belongs, the name of the generation model to be called, and the required number of samples. After the system obtains the task parameters, it calls the corresponding model to complete the data generation task;

[0026] The dynamic update sample generation task status module is used for the system to display the execution progress of the data generation task to the user and show the execution results, including the total number of samples generated this time, the number of compliant samples, the number of invalid samples, and the success rate;

[0027] The view all task information module is used for the system to display the total number of tasks, the number of failed tasks, and the number of successful tasks to the user;

[0028] The view sample information module is used for the user to view the specific values of all feature values of the generated samples and the label categories to which they belong.

[0029] Furthermore, the threat detection module includes a detection model selection and training module, a dynamic update model training status module, a detection model management module, a dynamic update model training result module, a threat traffic detection task module, and a traffic detection and analysis module; among them,

[0030] The detection model selection and training module is used for the user to select the dataset, algorithm model, and hyperparameters for model training that need to be relied on, and automatically obtain the start time of the current training task in the system to add a time stamp to the training task;

[0031] The dynamic update model training status module is used for the system to obtain the progress and results of the backend model training, and the information on whether the training is successful is presented to the user in the form of a message prompt window;

[0032] The detection model management module is used for the user to preview and delete the information of the existing models in the system, which is convenient for the user to manage the system model library files;

[0033] The dynamic update model training result module is used for the system to display the effect of model training to the user, and show the improvement of the model in terms of performance under different types of datasets, including accuracy, precision, recall, and f1-score;

[0034] The threat traffic detection task module is used for the system to detect the uploaded traffic file, and the user uploads the traffic file to be tested and selects the model to be tested to formulate a threat detection task;

[0035] The traffic detection and analysis module is used for the system to display threat traffic and normal traffic to the user, and track and locate the location and information of the threat traffic, specifically including traffic number, source IP, source port, destination IP, destination port, and traffic type information.

[0036] Furthermore, the data visualization module includes a data preprocessing quantity comparison bar chart module, an attack IP country quantity pie chart module, a dataset different type traffic proportion pie chart module, a dataset different type traffic quantity bar chart module, a generation model training loss curve chart module, a data distribution comparison curve chart module, and a data structure relationship heat map module.

[0037] Furthermore, the system architecture is mainly based on FastAPI 0.x + Vue 2.x. Python technology is used to implement the training and execution of the backend algorithm model, data processing, condition selection, and providing interface services. Vue is used to implement the functions of generating data on the front end, detecting abnormal traffic, displaying dataset information, and calling interfaces.

[0038] The advantages and beneficial effects of the present invention are as follows:

[0039] 1. Novel functions: Currently, threat detection tasks rely on traditional pattern recognition and feature selection, which have high requirements for the professional knowledge reserve and technical capabilities of relevant staff. The software system proposed by the present invention combines deep learning technology and system platform technology, which can help network management professionals generate the required threat data more quickly and conveniently and analyze the threat information existing in the captured network traffic.

[0040] 2. Comprehensive system architecture: The system proposed by the present invention adopts a modular design, including system architecture, data management, data generation, threat detection, data visualization, and user management modules. This design provides a complete solution that can meet the needs of different users and scenarios and has good scalability and flexibility.

[0041] 3. Openness and stability: The software system proposed by the present invention uses FastAPI and Vue technologies as the main architecture, uses Python technology to implement functions such as backend algorithm model training and execution, data processing, etc., and at the same time realizes the functions of the front-end interface through Vue. This technical solution can provide an efficient, stable, and user-friendly system experience.

[0042] 4. Flexibility: The data management module proposed by the present invention provides rich functions, including data preprocessing, dynamically updating data status, database selection and preview, etc. Users can manage and process data according to their needs to ensure the quality and consistency of the data. Users can select a model for training according to their needs, and through functions such as model selection, training status display, and training result evaluation provided by the system, the controllability and adjustability of the model training process are realized.

[0043] 5. High efficiency and accuracy: The data generation module proposed in the present invention adopts a generative adversarial network model. Through two methods, namely training the algorithm model and invoking the algorithm model, it can quickly generate various types of threat samples, ensure the consistency of the overall distribution and feature structure of the generated data, and the generated simulated data is highly similar to real data. Such a design can improve the efficiency and accuracy of data generation. The threat detection module proposed in the present invention has the ability to efficiently and accurately identify threat traffic and normal traffic types. By performing two tasks, namely training the detection model and conducting threat detection, it ensures the class accuracy of identifying threat signaling flows and traces and locates key information such as the IP and port of the abnormal terminal.

[0044] 6. Intuitiveness: The data visualization module proposed in the present invention provides a variety of chart modules, such as a bar chart for quantity comparison, a pie chart for traffic proportion, a training loss curve chart, etc. The design proposed in the present invention can help users analyze data more deeply and provide a scientific basis for decision-making. Each module in the system proposed in the present invention has a dynamic update function, which can display the task execution progress and results in a timely manner. Such a design can provide users with real-time feedback and status prompts, and improve users' perception and understanding of the system status.

[0045] 7. Authenticity: The system experiment of the present invention is based on a real 5G threat data set, and the system effect is real and reliable. The open-source data set comes from the 5G test network of the University of Oulu in Finland. The attacker nodes attack the servers deployed in the 5GTN MEC environment, and the attack scenarios include DoS attacks and port scans; another part of the data comes from the data collected by the XPRO Replay professional traffic simulation instrument during attack simulation in a simulated real 5G environment, and the attack scenarios include vulnerability attacks and DoS attacks. Through experimental verification, both types of data can cause highly destructive effects on the core network.

[0046] The innovation of the present invention is mainly reflected in the realization of a data management module that specifically extracts customized features for 5G protocols, ensuring the effectiveness of the system in learning traffic sequence features and generating real distribution samples. In addition, the verification of the system in a real 5G communication network environment and a simulation environment ensures its effectiveness in practical applications, meets the security analysis and data management needs of 5G network operation and maintenance personnel, and has high innovation and application value. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 is a schematic diagram of the deployment solution of the preferred embodiment provided by the present invention;

[0048] Figure 2 is a schematic diagram of the overall architecture of the system of the present invention;

[0049] Figure 3 is a schematic diagram of the software architecture of the system of the present invention;

[0050] Figure 4 Schematic diagram of the system function module of the present invention;

[0051] Figure 5 Schematic diagram of the system home page of the present invention;

[0052] Figure 6 Schematic diagram of the data management module of the present invention;

[0053] Figure 7 Schematic diagram of the data generation module of the present invention;

[0054] Figure 8 Performance comparison chart of the data generation effect of the present invention;

[0055] Figure 9 Schematic diagram of the threat detection module of the present invention;

[0056] Figure 10 Schematic diagram of the data visualization module of the present invention;

[0057] Figure 11 Schematic diagram of the user management module of the present invention. Detailed implementation manners

[0058] Next, the technical solutions in the embodiments of the present invention will be clearly and detailedly described in conjunction with the accompanying drawings in the embodiments of the present invention. The described embodiments are only a part of the embodiments of the present invention.

[0059] The technical solution for the present invention to solve the above technical problems is:

[0060] Embodiment 1

[0061] This embodiment provides a threat signaling flow data generation and detection system based on a generative adversarial network. The system can be deployed at each network element of the core network. The system is based on the B / S architecture, and the relationship between software and hardware in the system is realized through the designed layering and service separation. The relationship between software and hardware is reflected in that the software layer design can make full use of hardware resources to complete specified tasks. The service layer of the system assigns tasks to corresponding computing servers for processing according to the task type, and uses the computing resources of the servers to improve the speed and efficiency of task processing. At the same time, through the idea of microservices, tasks such as scheduling, computing, and storage are separated, and each service runs on the most suitable hardware.

[0062] As Figure 2As shown, based on the microservices concept, the scheduling, computing, and storage tasks are separated to achieve decoupling of system modules. Specifically, the scheduling task is responsible for the system server, including the allocation and scheduling of user tasks. The computing service processes two types of complex computing tasks: model training and data processing, concentrating computing resources to improve processing speed and efficiency. Ensure that compute-intensive tasks do not degrade in performance due to resource contention. The storage task is managed by a dedicated storage server, responsible for storing and managing user-uploaded data files and other data required for system operation.

[0063] As Figure 3 shown, the software part of the system is oriented to actual needs, using Python and FastAPI as backend development tools and Vue2 as the frontend development framework. The system is divided into four layers according to the functional hierarchy: the presentation layer, the interaction layer, the service layer, and the data layer.

[0064] Presentation layer: The presentation layer of the present invention is the user interface of the system, mainly responsible for the visualization tasks of the system and data, and is developed using the Vue2.0 framework and Element UI. The presentation layer sends the user's operation requests to the interaction layer for processing and displays the results processed by the interaction layer to the user. The user interacts with the system frontend through the access layer to initiate operation command requests such as data upload and task assignment.

[0065] Interaction layer: The interaction layer of the present invention is responsible for processing requests sent by the frontend and forwarding the requests to the backend layer for processing. In the system, the interaction layer is mainly implemented by Vue-router route control and Axios interface requests. Vue-router is used to manage the routes of the frontend pages, implementing page jump control, parameter passing, and route interception functions. The Axios interface request is responsible for data transfer and exchange between the presentation layer and the service layer. By defining a unified Axios interface and a standardized data format, data interaction between the front and backend is achieved, and the accuracy and integrity of the data are ensured. Security verification and data encryption and decryption of user requests are implemented in the interaction layer.

[0066] Service layer: The service layer of the present invention is responsible for implementing the business logic and function realization of the system. The service layer is the core part of the system and is developed using the Python and FastAPI frameworks. After the system backend service of the present invention receives the request sent by the presentation layer, it performs corresponding task allocation and scheduling according to the content of the request. Generally, it is divided into system services, data analysis services, model training services, and model invocation services. The system service processes system messages, business requests, etc., and maintains the overall operation of the system. The data analysis service is responsible for completing data preprocessing, sample similarity analysis, and structural relationship analysis. The model invocation is responsible for processing the tasks sent by the system service and invoking the existing generation model to complete the sample generation task or the detection model to complete the threat detection task. The model training service is responsible for completing the training task of the specified model according to the user-defined model and dataset-related parameters. The system and computing server designed by the present invention will carry the service layer to maintain the operation of the entire system.

[0067] Data layer: The data layer of the present invention is responsible for storing and managing the data of the system, mainly including the storage of dataset files uploaded by users. In the system, the data layer uses the MySQL database for data storage and management, providing operations such as adding, deleting, modifying, and querying data.

[0068] The advantages of the threat signaling flow data generation and detection system architecture based on the generative adversarial network are summarized as follows:

[0069] 1) Full utilization of software and hardware: The system of the present invention can fully utilize hardware resources. By allocating tasks to the corresponding computing servers for processing, the performance and efficiency of the system are improved. This design method enables the system to better meet the requirements of large-scale task processing while ensuring the stability and reliability of the system.

[0070] 2) Decoupled design: The present invention adopts the microservice architecture concept to split the system into independent service units, and each service unit can be independently deployed and extended. This design method makes the system easier to understand and maintain. At the same time, it also improves the flexibility and scalability of the system, and new services can be quickly adjusted and deployed according to needs. The decoupled design of the system modules reduces the dependency relationship between modules, making the system more flexible and maintainable. When a certain module changes, it will not affect the normal operation of other modules, reducing the risk and maintenance cost of the system.

[0071] 3) High stability and maintainability: The combined software and hardware design and decoupled design of the present invention improve the overall stability and maintainability of the system. The modular design of the system enables each module to be independently tested and debugged, facilitating the location and solution of problems, thereby improving the reliability and maintainability of the system.

[0072] 4) High scalability: The decoupled design and microservices concept of the present invention improve the scalability of the system, enabling the system to better cope with new requirements and changes that may arise in the future. The system can quickly add or delete service units as needed without affecting the overall operation of the system.

[0073] 5) Clear business logic: The system of the present invention is divided into a presentation layer, an interaction layer, a service layer, and a data layer according to functional levels, making the business logic of the system clearly visible. Each layer has clear responsibilities and functions, facilitating understanding and management. At the same time, it is also conducive to teamwork and code maintenance.

[0074] 6) Front-end and back-end separation: The present invention adopts a front-end and back-end separation design pattern, enabling the development and deployment of the front-end and back-end of the system to be carried out independently, reducing the coupling degree of the system, and improving the flexibility and maintainability of the system. Front-end and back-end separation can also improve the concurrent processing ability of the system and enhance the user experience.

[0075] 7) Security and data integrity: The interaction layer of the present invention implements security verification of user requests and data encryption and decryption, ensuring the security of the system and the integrity of the data. This design method can effectively prevent malicious attacks and data tampering, protecting user privacy and data security.

[0076] 8) Flexible development tools: Using Python and FastAPI as back-end development tools and Vue2 as the front-end development framework makes the development process of the system more flexible and efficient. As a simple and readable programming language, Python can improve development efficiency and code readability. The FastAPI framework provides the function of quickly building APIs, making back-end development simpler and more efficient. As a lightweight front-end framework, Vue2 provides rich components and functions, making front-end development more convenient and flexible.

[0077] The functions of the threat signaling flow data generation and detection system based on the generative adversarial network are generally described as follows:

[0078] As Figure 4As shown in the figure, the threat signaling flow data generation and detection system based on the generative adversarial network includes a data management module, a data generation module, a threat detection module, a data visualization module, and a user management module. Among them, the data management module is used for data preprocessing and data management of dataset information, including data collection and upload, data cleaning, dynamically updating the data processing status and results, and statistics of the existing datasets in the system; the data generation module is used to help users quickly train models and call models to generate different types of data samples required by users through two task methods: training algorithm models and calling algorithm models, and analyze and compare the results of the generated data. This algorithm not only ensures the consistency of the overall distribution of data samples but also the consistency of the feature structure relationships of data samples; the threat detection module is used to identify threat traffic and normal traffic types, including two tasks: detection model training and threat detection. This algorithm can improve the detection effect of the detection model, ensure the efficient and accurate detection of threat signaling flow categories, and track and locate information such as the IP and port numbers of abnormal terminals; the data visualization module is used to visually display information such as dataset statistics, sample distribution similarity, and sample structure similarity in the form of charts; the user management module is used to manage accounts, including an email login and registration module, an email password recovery module, and a personal information modification module.

[0079] 1. Data Management Module

[0080] 1) Data Preprocessing Module: Users can customize data preprocessing methods, and can select options such as feature extraction, outlier cleaning, normalization processing, and feature dimensionality reduction to formulate data preprocessing tasks;

[0081] 2) Dynamically Update Data Preprocessing Status Module: The system shows the execution progress of data preprocessing tasks to users and prompts task completion in the form of a pop-up window;

[0082] 3) Data Preprocessing Result Module: The system dynamically updates the data preprocessing results, shows the feature results of preview data preprocessing to users, and shows the comparison of the data volumes of various types of traffic before and after processing;

[0083] 4) Database Selection Module: Users can customize and select the preview threat dataset mode, customize dataset selection, attack type selection, and attack scenario selection schemes, and query data information in the database;

[0084] 5) Database Preview Module: The system shows users the overall situation of the dataset in the custom preview mode, including the country to which the attack IP belongs, the proportion of various types of threats in the dataset, and the quantity information of various types of threats.

[0085] 2. Data Generation Module

[0086] 1) Generation model selection and training module: The user selects the type of traffic samples to be trained and generated, the algorithm model to be relied on, and the hyperparameters for model training, and automatically obtains the start time of the current training task in the system as the training task plus a timestamp;

[0087] 2) Dynamic update of model training status module: The system obtains the progress and results of the backend model training, and the information on whether the training is successful is presented to the user in the form of a message prompt window;

[0088] 3) Generation model management module: The user performs information preview and deletion operations on the existing models in the system, facilitating the user to manage the system model library files;

[0089] 4) Sample distribution similarity calculation module: The system calculates the overall distribution of the samples generated by the generation model and the overall distribution of the real samples, and verifies the similarity between the two;

[0090] 5) Sample relationship structure calculation module: The system calculates the correlation coefficient between the internal features of the generated sample set, and compares the correlation coefficient with the correlation coefficient of the real sample set features to verify the consistency between the internal relationships of the two features;

[0091] 6) Data generation task module: The user formulates a sample generation task, selects the attack category to which the sample belongs, the name of the generation model to be called, and the number of samples required. After the system obtains the task parameters, it calls the corresponding model to complete the data generation task;

[0092] 7) Dynamic update of sample generation task status module: The system shows the execution progress of the data generation task to the user, and shows the execution results, including the total number of samples generated this time, the number of compliant samples, the number of invalid samples, and the success rate;

[0093] 8) View all task information module: The system shows the total number of tasks, the number of failed tasks, and the successful tasks to the user;

[0094] 9) View sample information module: Used for the user to view the specific values of all feature values of the generated samples and the label categories to which they belong.

[0095] 3. Threat detection module

[0096] 1) Detection model selection and training module: The user selects the dataset to be relied on, the algorithm model, and the hyperparameters for model training, and automatically obtains the start time of the current training task in the system as the training task plus a timestamp;

[0097] 2) Dynamic update of model training status module: The system obtains the progress and results of the backend model training, and the information on whether the training is successful is presented to the user in the form of a message prompt window;

[0098] 3) Detection Model Management Module: Users can preview and delete information of the existing models in the system, facilitating the management of the system model library files by users.

[0099] 4) Dynamic Update of Model Training Results Module: The system shows the training effect of the model to users, demonstrating the improvement in model performance under different types of datasets, including accurancy, precision, recall, and f1-score.

[0100] 5) Threat Traffic Detection Task Module: Users upload the traffic files to be tested and select the models to be tested to formulate threat detection tasks, and the system detects the uploaded traffic files according to the task parameters.

[0101] 6) Traffic Detection and Analysis Module: The system updates the detection results, shows threat traffic and normal traffic to users, and traces and locates the position and information of the threat traffic, specifically including traffic number, source IP, source port, destination IP, destination port, and traffic type information.

[0102] 4. Data Visualization Module

[0103] 1) Bar Chart Module for Comparing the Quantity of Data Preprocessing: This module is used to compare the quantity changes before and after data preprocessing using different methods. In the form of a bar chart, it intuitively shows the increase and decrease of the data volume during the data preprocessing process, helping users understand the impact degree of various preprocessing methods on the dataset.

[0104] 2) Pie Chart Module for the Quantity of Attack IPs Belonging to Each Country: This module is used to show the quantity distribution of attack IPs belonging to each country. In the form of a pie chart, it clearly shows the proportion of the number of attack IPs in each country, helping users quickly understand the geographical distribution of attack behaviors.

[0105] 3) Pie Chart Module for the Proportion of Each Type of Traffic in the Dataset: This module is used to show the proportion of different types of threat traffic in the dataset. In the form of a pie chart, it intuitively shows the proportion of each threat type in the dataset, helping users understand the category distribution of the dataset.

[0106] 4) Bar Chart Module for the Quantity of Each Type of Traffic in the Dataset: This module is used to show the quantity of different threat types in the dataset. In the form of a bar chart, it clearly shows the quantity of each threat type in the dataset, helping users quickly understand the distribution of various types of traffic in the dataset.

[0107] 5) Curve Chart Module for Generating the Training Loss of the Model: This module is used to show the change of the loss function during the training of the generative model. In the form of a curve chart, it intuitively shows the change trend of the loss function of the generative model under different training rounds, helping users evaluate the training effect and convergence of the model.

[0108] 6) Data distribution comparison curve graph module: This module is used to compare the data distribution between the original data and the generated threat data. In the form of a curve graph, it clearly shows the differences in data distribution, helping users analyze the characteristics and similarities of the generated threat data.

[0109] 7) Heat map module of data structure relationship: This module is used to display the correlation and relationship between different features in the dataset. In the form of a heat map, it intuitively shows the correlation coefficients between data features, helping users understand the structure of the data and the mutual relationship between features.

[0110] The systems, devices, modules or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions.

[0111] It should also be noted that the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent in such process, method, commodity or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of another identical element in the process, method, commodity or device including the said element.

[0112] The above embodiments should be understood as only illustrative of the present invention and not restrictive of the scope of protection of the present invention. After reading the content described in the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent changes and modifications also fall within the scope defined by the claims of the present invention.

Claims

1. A threat signaling flow data generation and detection system based on a generative adversarial network, characterized in that: include: The data management module is used to pre-process and manage data set information, including data collection and uploading, data cleaning, dynamic update of data processing status and data processing results, and statistics of existing data sets in the system; The data generation module is used to help users quickly train models and call models to generate different types of data samples required by users through two task modes: training algorithm models and calling algorithm models, and analyze and compare the results of generating threat data; The threat detection module is used to identify threat traffic and normal traffic types. It includes two tasks: detection model training and threat detection. It is used to detect threat signaling flow categories and track and locate the IP and port number information of abnormal terminals. Data visualization module, used to visualize the statistical information of data sets, sample distribution similarity, sample structure similarity, sample data characteristics, and anomaly detection result information in the form of charts; The user management module is used to manage accounts, including the email login and registration module and the email password retrieval module.

2. According to claim 1, a threat signaling flow data generation and detection system based on a generative adversarial network is characterized in that: The data management module includes a data preprocessing module, a dynamic update data preprocessing status module, a data preprocessing result module, a database selection module, and a database preview module; wherein, The data preprocessing module is used for the system to provide users with customized data preprocessing methods, including selecting feature extraction, outlier cleaning, standardization, and feature dimension reduction operations to formulate data preprocessing tasks; The dynamic update data preprocessing status module is used for the system to display the execution progress of the data preprocessing task to the user, prompting the task completion in the form of a pop-up window; The data preprocessing result module is used for users to preview the characteristic results of data preprocessing, and to display the data volume comparison of each type of traffic before and after the processing is completed; The database selection module is used for users to customize the selection of preview threat data set modes, customize data set selection, threat type selection, attack scenario selection schemes, and query data information in the database; The database preview module is used by the system to show the user the overall situation of the data set in the custom preview mode, including the country to which the attacking IP belongs, the proportion of each type of traffic in the data set, and the quantity information of each type of traffic.

3. According to claim 1, a threat signaling flow data generation and detection system based on a generative adversarial network is characterized in that: The data generation module includes a generation model selection and training module, a dynamic update model training status module, a generation model management module, a sample distribution similarity calculation module, a sample relationship structure calculation module, a data generation task module, a dynamic update sample generation task status module, a viewing all task information module and a viewing generation sample information module; wherein, The generation model selection and training module is used by the user to select the type of threat data that needs to be trained, the algorithm model it relies on, and the hyperparameters of the model training, and automatically obtains the start time of the current training task in the system to add a timestamp to the training task; The dynamic update model training status module is used by the system to obtain the progress and results of the backend model training, and the information on whether the training is successful or not is displayed to the user in the form of a message prompt window; The generated model management module is used by the user to preview and delete the existing models in the system, so as to facilitate the user to manage the system model library files; The sample distribution similarity calculation module is used to systematically calculate the overall distribution of samples generated by the generation model and the overall distribution of real samples, and verify the similarity between the two; The sample relationship structure calculation module is used to calculate and generate the Pearson correlation coefficient between the internal features of the sample set. The calculation formula of the Pearson correlation coefficient r is as follows: In the formula, x i ,y i Represent the i-th value of sample variables x and y respectively; Represent the mean of sample variables x and y respectively. Then compare the generated sample correlation coefficient with the feature correlation coefficient of the real sample set to verify the consistency between the internal relationship of the two features; The data generation task module is used by the user to formulate a sample generation task, select the attack category to which the sample belongs, the name of the generation model to be called, and the number of samples required. After the system obtains the task parameters, it calls the corresponding model to complete the data generation task; The dynamically updated sample generation task status module is used for the system to display the execution progress of the data generation task to the user, and display the execution results, including the total number of samples generated this time, the number of compliant samples, the number of invalid samples, and the success rate; The view all task information module is used for the system to display the total number of tasks, the number of failed tasks and the number of successful tasks to the user; The sample information viewing module is used for users to view the specific values ​​of all feature values ​​of the generated samples and the label categories to which they belong.

4. According to claim 1, a threat signaling flow data generation and detection system based on a generative adversarial network is characterized in that: The threat detection module includes a detection model selection and training module, a dynamic update model training status module, a detection model management module, a dynamic update model training result module, a threat flow detection task module and a flow detection analysis module; wherein, The detection model selection and training module is used by the user to select the data set, algorithm model and hyperparameters of model training that need to be relied on, and automatically obtains the start time of the current training task in the system to add a timestamp to the training task; The dynamic update model training status module is used by the system to obtain the progress and results of the backend model training, and the information on whether the training is successful or not is displayed to the user in the form of a message prompt window; The detection model management module is used by users to preview and delete existing models in the system, making it convenient for users to manage system model library files; The dynamic update model training result module is used for the system to show the effect of model training to users, showing the performance improvement of the model under different types of data sets, including accuracy, precision, recall, and f1-score; The threat traffic detection task module is used for the system to detect the uploaded traffic files, and the user uploads the traffic files to be tested and selects the model to be tested to formulate the threat detection task; The traffic detection and analysis module is used by the system to display threat traffic and normal traffic to users, and track and locate the location and information of threat traffic, including traffic number, source IP, source port, destination IP, destination port, and traffic type information.

5. According to claim 1, a threat signaling flow data generation and detection system based on a generative adversarial network is characterized in that: The data visualization module includes a data preprocessing quantity comparison bar chart module, a country number ring chart module for attacking IPs, a fan chart module for the proportion of different types of traffic in the data set, a bar chart module for the number of different types of traffic in the data set, a model training loss curve chart module, a data distribution comparison curve chart module, and a data structure relationship heat map module.

6. A threat signaling flow data generation and detection system based on a generative adversarial network according to any one of claims 1 to 5, characterized in that: The system architecture is mainly based on FastAPI0.x+Vue2.x. Python technology is used to implement back-end algorithm model training and execution, data processing, condition selection and interface service provision. Vue implements the front-end data generation, abnormal traffic detection, data set information display and interface calling functions.

Citation Information

Patent Citations

  • Network intrusion detection method and system based on generative adversarial network

    CN113114673A

  • Intrusion detection method and system based on multi-generator GAN data enhancement

    CN117118718A

  • Network boundary threat detection system based on AI drive

    CN118827227A

  • Network traffic threat detection method based on artificial intelligence model

    CN119094149A