Network attack sample intelligent identification and defense method, system, device and medium
Patent Information
- Application Number
- CN202610932292.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-26
- Publication Date
- 2026-09-22
AI Technical Summary
[0003]然而,相关技术方案通常在封闭静态环境中训练,模型参数和判别边界固化,攻击者易构造对抗样本实施绕过,并且系统训练完成后便固化下来,识别能力不再随攻击技术演变而更新
[0010]本公开实施例中所提供的网络攻击样本智能识别与防御方法、系统、装置及介质,通过动态环境模拟器监测攻击样本生成器与攻击样本判别器在演化周期内的对抗表现数据并计算适应性分数,结合演化指导信号对生成器和判别器进行周期性调整,实现了生成器和判别器在网络攻击样本识别过程中的持续演化与协同优化,提升了网络攻击防御对变异样本的适应能力以及分层防御动作的执行准确性。
Smart Images

Figure CN122802215A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of cybersecurity technology, and more specifically, to a method, system, device, and medium for intelligent identification and defense of network attack samples. Background Technology
[0002] As cyberattack techniques continue to evolve, traditional detection methods based on fixed signatures or rule bases are struggling to cope with rapidly evolving malware and advanced threats, and their lag often leaves defense systems in a reactive state. Artificial intelligence technologies, such as generative adversarial networks (GANs), have been introduced into the cybersecurity field. These GANs enhance detection capabilities through adversarial training that generates simulated attack samples and distinguishes between real traffic and generated samples.
[0003] However, related technical solutions are usually trained in a closed and static environment, where model parameters and discrimination boundaries are fixed. Attackers can easily construct adversarial examples to bypass them, and once the system is trained, it becomes fixed, and the recognition capability is no longer updated with the evolution of attack techniques. Summary of the Invention
[0004] This disclosure provides at least one method, system, device, and medium for intelligent identification and defense of network attack samples. By periodically evolving and adjusting the generator and discriminator through a dynamic environment simulator, continuous adaptive optimization of network attack sample identification and defense is achieved.
[0005] This disclosure provides a method for intelligent identification and defense against network attack samples, including: Obtain the trained attack sample generator and attack sample discriminator; Real-time capture of network data streams and parsing of the network data streams to obtain real-time traffic feature vectors; The attack sample generator generates a synthetic attack feature vector, and the attack sample discriminator outputs a probability discrimination result based on the real-time traffic feature vector and the synthetic attack feature vector. Based on the probability discrimination results, layered defense actions are performed on the network data stream; wherein, the layered defense actions include real-time interception, deep data analysis, and normal traffic passage. At the end of each evolution cycle, the adversarial performance data of the attack sample generator and the attack sample discriminator during the evolution cycle are monitored using a dynamic environment simulator, and the generation fitness score of the attack sample generator and the discrimination fitness score of the attack sample discriminator are calculated based on the adversarial performance data. The dynamic environment simulator determines a generation evolution guidance signal for the attack sample generator based on the generation fitness score, and adjusts the attack sample generator using the generation evolution guidance signal; and the dynamic environment simulator determines a discrimination evolution guidance signal for the attack sample discriminator based on the discrimination fitness score, and adjusts the attack sample discriminator using the discrimination evolution guidance signal. Return to the step of capturing network data streams in real time, and continue execution using the adjusted attack sample generator and attack sample discriminator.
[0006] This disclosure provides an intelligent identification and defense system for network attack samples, including: The data capture module is used to capture network data streams in real time and parse the network data streams to obtain real-time traffic feature vectors; The sample recognition engine includes a trained attack sample generator, an attack sample discriminator, and a dynamic environment simulator. The attack sample generator is used to generate a synthetic attack feature vector; the attack sample discriminator is used to output a probability discrimination result based on the real-time traffic feature vector and the synthetic attack feature vector. The dynamic environment simulator is configured to monitor the adversarial performance data of the attack sample generator and the attack sample discriminator during the evolution cycle at the end of each evolution cycle, and calculate the generation fitness score of the attack sample generator and the discrimination fitness score of the attack sample discriminator based on the adversarial performance data; determine a generation evolution guidance signal for the attack sample generator based on the generation fitness score, and adjust the attack sample generator using the generation evolution guidance signal; and determine a discrimination evolution guidance signal for the attack sample discriminator based on the discrimination fitness score, and adjust the attack sample discriminator using the discrimination evolution guidance signal. The sample recognition engine uses the adjusted attack sample generator and the adjusted attack sample discriminator to perform intelligent recognition of network attack samples in the next evolution cycle; The defense execution module is used to perform layered defense actions on the network data stream based on the probability discrimination result; wherein, the layered defense actions include real-time interception, deep data analysis, and normal traffic passage.
[0007] This disclosure provides an intelligent identification and defense device for network attack samples, including: The model acquisition module is used to acquire trained attack sample generators and attack sample discriminators. The data acquisition module is used to capture network data streams in real time and parse the network data streams to obtain real-time traffic feature vectors. The model application module is used to generate a synthetic attack feature vector using the attack sample generator, and to output a probability discrimination result using the attack sample discriminator based on the real-time traffic feature vector and the synthetic attack feature vector. The defense execution module is used to perform layered defense actions on the network data stream based on the probability discrimination result; wherein, the layered defense actions include real-time interception, deep data analysis, and normal traffic passage. The model adjustment module is configured to: at the end of each evolution cycle, monitor the adversarial performance data of the attack sample generator and the attack sample discriminator during the evolution cycle using a dynamic environment simulator, and calculate the generation fitness score of the attack sample generator and the discrimination fitness score of the attack sample discriminator based on the adversarial performance data; determine a generation evolution guidance signal for the attack sample generator based on the generation fitness score using the dynamic environment simulator, and adjust the attack sample generator using the generation evolution guidance signal; determine a discrimination evolution guidance signal for the attack sample discriminator based on the discrimination fitness score using the dynamic environment simulator, and adjust the attack sample discriminator using the discrimination evolution guidance signal; and return to the step of real-time capture of network data streams, and continue execution using the adjusted attack sample generator and the attack sample discriminator.
[0008] This disclosure provides a computer device, including a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the computer device is running, the processor communicates with the memory via the bus. When the machine-readable instructions are executed by the processor, the intelligent identification and defense method for network attack samples as described in any of the above possible embodiments is executed.
[0009] This disclosure provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the intelligent identification and defense method for network attack samples as described in any of the possible embodiments above.
[0010] The intelligent identification and defense method, system, device, and medium for network attack samples provided in this disclosure monitor the adversarial performance data of the attack sample generator and the attack sample discriminator during the evolution cycle through a dynamic environment simulator and calculate the fitness score. Combined with the evolution guidance signal, the generator and discriminator are periodically adjusted, realizing the continuous evolution and collaborative optimization of the generator and discriminator in the process of network attack sample identification. This improves the adaptability of network attack defense to mutated samples and the accuracy of the execution of layered defense actions.
[0011] To make the above-mentioned objects, features and advantages of this disclosure more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description
[0012] To more clearly illustrate the technical solutions of the embodiments of this disclosure, the accompanying drawings referenced in the embodiments will be briefly described below. These drawings are incorporated in and constitute a part of this specification. They illustrate embodiments conforming to this disclosure and, together with the specification, serve to explain the technical solutions of this disclosure. It should be understood that the following drawings only show some embodiments of this disclosure and should not be considered as limiting the scope. Those skilled in the art can obtain other related drawings based on these drawings without creative effort.
[0013] Figure 1 A flowchart of a method for intelligent identification and defense of network attack samples provided in an embodiment of this disclosure is shown; Figure 2 A flowchart of a triple authentication task processing method provided by an embodiment of this disclosure is shown; Figure 3 A flowchart of an adversarial memory pool construction and update method provided by an embodiment of this disclosure is shown; Figure 4 This diagram illustrates the structure of an intelligent identification and defense system for network attack samples provided in an embodiment of this disclosure. Figure 5 This diagram illustrates the structure of a network attack sample intelligent identification and defense device provided in an embodiment of the present disclosure; Figure 6 A schematic diagram of the structure of a computer device provided in an embodiment of this disclosure is shown. Detailed Implementation
[0014] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this disclosure, and not all of them. The components of the embodiments of this disclosure described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this disclosure provided in the accompanying drawings is not intended to limit the scope of the claimed disclosure, but merely represents selected embodiments of this disclosure. All other embodiments obtained by those skilled in the art based on the embodiments of this disclosure without inventive effort are within the scope of protection of this disclosure.
[0015] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.
[0016] In this document, the term "and / or" merely describes a relationship, indicating that three relationships can exist. For example, A and / or B can represent three cases: A alone, A and B simultaneously, and B alone. Furthermore, the term "at least one" in this document means any combination of at least two of any one or more elements. For example, including at least one of A, B, and C can mean including any one or more elements selected from the set consisting of A, B, and C.
[0017] With the continuous evolution of cyberattack techniques, traditional security defenses face severe challenges. Detection methods based on fixed signatures or rule bases, such as traditional firewalls, intrusion detection systems, and antivirus software, heavily rely on pre-set virus signatures, attack signatures, or static rules to identify threats by matching known characteristics. However, such solutions struggle to cope with rapidly evolving malware, advanced evasion attacks, and unknown threats, and their inherent lag often leaves defense systems in a reactive state. Against this backdrop, artificial intelligence technologies, represented by generative adversarial networks (GANs), have been introduced into the cybersecurity field, aiming to achieve more intelligent threat identification by learning the deep characteristics of data. A typical application involves using a generative network to simulate attack samples, while a discriminant network distinguishes real traffic from generated samples. The adversarial training between the two improves the discriminator's detection capabilities, and this approach is commonly used for intrusion detection, malicious traffic analysis, malware detection, and data augmentation.
[0018] Research has revealed fundamental flaws in existing defense schemes based on generative adversarial networks (GANs) and broader traditional defense systems. Firstly, statically trained GANs are typically trained in relatively closed and fixed environments, with the generator and discriminator engaging in a game around a limited dataset or a pre-defined objective. Once training is complete, the model parameters and discrimination boundaries of the entire system become fixed, and its recognition capabilities no longer update with the evolution of attack techniques. This static adversarial model is severely inconsistent with the dynamic evolution of attack and defense in real cyberspace. Attackers can easily discover the fixed decision boundaries of the discriminator through analysis or probing, and construct targeted adversarial samples to bypass them, quickly rendering the defense effects gained in the early training ineffective. A deeper problem lies in the lack of effective guidance and global judgment of the game process in traditional adversarial training frameworks: the generator's goal is often limited to deceiving the discriminator within the current period, easily falling into generating repetitive samples with a single generation pattern; the discriminator's goal is limited to improving the classification accuracy of the existing sample set, potentially leading to overfitting and loss of generalization ability. The game between the two sides exhibits a local and short-sighted competitive state, failing to simulate the long-term dynamic process of continuous strategy innovation and spiraling capability improvement in real attack and defense. On the other hand, existing technology systems generally separate detection models from defense responses. In systems represented by traditional PDR / P2DR models, the discriminator output is usually just an isolated probability value or a simple alarm. Subsequent interception or decision-making is based on a single threshold judgment, which has problems such as strong subjectivity in threshold setting, high false alarm and false negative rates, and lack of dynamic adaptive capabilities.
[0019] Based on the above research, this disclosure provides a method, system, device, and medium for intelligent identification and defense of network attack samples. Specifically, firstly, a trained attack sample generator and attack sample discriminator are acquired. Then, network data streams are captured and parsed in real time to extract real-time traffic feature vectors. Next, the attack sample generator synthesizes attack feature vectors, and the attack sample discriminator outputs a probability discrimination result based on the real-time traffic feature vectors and the synthesized attack feature vectors. This result guides the implementation of layered defense actions on the network data streams, including real-time interception, deep data analysis, and normal traffic passage. At the end of each evolution cycle, the dynamic environment simulator monitors the adversarial performance data of the generator and discriminator during that cycle, and calculates the generation fitness score and discrimination fitness score accordingly. This generates corresponding generation evolution guidance signals and discrimination evolution guidance signals to adjust the attack sample generator and attack sample discriminator. Afterward, the process returns to the real-time network data stream capture step, and the adjusted generator and discriminator continue to execute the above defense and evolution process.
[0020] In this embodiment, the adversarial performance data of the attack sample generator and the attack sample discriminator during the evolution cycle are monitored by a dynamic environment simulator and the fitness score is calculated. The generator and discriminator are periodically adjusted in combination with the evolution guidance signal, thereby realizing the continuous evolution and collaborative optimization of the generator and discriminator in the process of network attack sample identification. This improves the adaptability of network attack defense to mutated samples and the accuracy of the execution of layered defense actions.
[0021] To facilitate understanding of this embodiment, the executing entity of the intelligent identification and defense method for network attack samples provided in this disclosure will first be described in detail. The executing entity of the intelligent identification and defense method for network attack samples provided in this disclosure is a computer device. This computer device can be a terminal device or a server. The terminal device can also be a mobile device, a user terminal, a terminal, a handheld device, a computing device, an in-vehicle device, a wearable device, etc. The server can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud storage, big data, and artificial intelligence platforms. Optionally, this method can also be applied to an implementation environment composed of computer devices and servers.
[0022] The intelligent identification and defense method for network attack samples provided in this application embodiment will be described in detail below with reference to the accompanying drawings. See also Figure 1 The diagram shows a flowchart of a method for intelligent identification and defense of network attack samples provided in this disclosure, which includes the following steps S101-S105: S101, Obtain the trained attack sample generator and attack sample discriminator.
[0023] Understandably, the attack sample generator is a pre-trained deep neural network model whose input is a random noise vector and output is a synthetic feature vector simulating network attack behavior. Its role is to proactively construct data samples with attack characteristics within the cyberspace security defense system for adversarial training and defense capability verification. The attack sample discriminator is another pre-trained deep neural network model whose function is to classify and identify real-time network data streams. This model's input is a network traffic feature vector, and its output is a probability judgment of whether the traffic belongs to normal traffic, real attack traffic, or synthetic attack traffic.
[0024] Here, the attack sample generator and attack sample discriminator mentioned above have been trained in advance through supervised pre-training and unsupervised adversarial training. In the initial stage, the two models have the basic ability to identify and simulate common network attack types, which can provide reliable initial parameters and basic performance for subsequent co-evolution.
[0025] For example, taking intrusion detection in the field of cybersecurity as an example, publicly available malicious traffic datasets, such as the CICIDS2017 dataset or the UNSW-NB15 dataset, can be used for initial training of the attack sample generator and attack sample discriminator. During the initial training process, the attack sample discriminator first learns to distinguish between normal traffic and attack traffic using real network traffic data, while the attack sample generator simulates and generates attack samples that can initially deceive the discriminator through adversarial learning. After multiple rounds of alternating training, the two models reach an initial game equilibrium state, thus obtaining the trained attack sample generator and attack sample discriminator, which can be used for subsequent online detection and co-evolutionary updates.
[0026] S102, capture network data stream in real time, and parse the network data stream to obtain real-time traffic feature vector.
[0027] Here, a network data stream refers to a continuous sequence of data packets transmitted in a computer network. Each data packet may contain information such as the source Internet Protocol (IP) address, destination Internet Protocol (IP) address, Transmission Control Protocol (DCP) or User Datagram Protocol (UDP) port number, protocol type, and payload. These data packets can be organized into a stream in chronological order; for example, all data packets in a DCP connection or a UDP session can form a stream. At the enterprise network egress or bypass of a data center core switch, all network data streams passing through that node can be captured in real time using pre-deployed port mirroring probes or network splitters.
[0028] Understandably, since the raw network data stream is in binary form and cannot be directly input into machine learning models for computation, protocol parsing and feature extraction can be performed on the network data stream. This converts key fields in each data packet or stream into numerical forms, such as calculating stream duration, average packet length, packet rate, and flag distribution statistics. The result is a fixed-dimensional numerical array called a real-time traffic feature vector. This vector can represent the behavioral patterns of a network session. For example, the traffic feature vector of normal web browsing contains short durations and a large number of small uplink packets, while the traffic feature vector of a distributed denial-of-service attack is characterized by extremely high packet rates and extremely short packet intervals.
[0029] S103, the attack sample generator generates a synthetic attack feature vector, and the attack sample discriminator outputs a probability discrimination result based on the real-time traffic feature vector and the synthetic attack feature vector.
[0030] Furthermore, after obtaining the real-time traffic feature vector, an attack sample generator can first generate a synthetic attack feature vector based on a randomly sampled latent spatial noise vector. Then, an attack sample discriminator can jointly discriminate between the real-time traffic feature vector and the synthetic attack feature vector, calculating the probability discrimination result based on the activation values of the discriminator's output layer. The probability discrimination result refers to the probability value vector output by the attack sample discriminator after classifying and predicting the input feature vector. It is a multi-dimensional floating-point array that can be used to express the probability that the input traffic belongs to different categories, including the probability of belonging to normal traffic, the probability of belonging to real attack traffic, and the probability of belonging to synthetic attack traffic. It can also include the probability distribution on the attack category label and the corresponding discrimination confidence.
[0031] Here, the synthetic attack feature vector refers to the simulated attack data obtained by the attack sample generator through forward computation of a deep neural network. The attack sample generator can generate one or more synthetic attack feature vectors based on different sampling points in the latent space, which can be used as new attack simulation samples for adversarial training to challenge the recognition boundary of the attack sample discriminator and prompt it to improve its discrimination ability.
[0032] In some possible implementations, because there is a blurred boundary between normal behavior and attack behavior in network traffic, a single binary classification task may not be able to fully utilize attack type information. If only a binary result is output, the feature differences between different attack families may be ignored, thereby reducing the robustness of the model. Therefore, in order to improve the fine-grained recognition capability and generalization performance of the attack sample discriminator, refer to... Figure 2 As shown, when using the attack sample discriminator to output the probability discrimination result, the following steps S201~S203 may be included: S201, the real-time traffic feature vector and the synthetic attack feature vector are input together into the attack sample discriminator.
[0033] Here, after extracting the real-time traffic feature vector and generating the synthetic attack feature vector, the real-time traffic feature vector and the synthetic attack feature vector can be input together into the input layer of the attack sample discriminator as samples to be processed in the same batch. In this way, the attack sample discriminator can process real traffic and synthetic traffic at the same time and compare the distribution differences between the two.
[0034] S202, the attack sample discriminator is used to perform a triple discrimination task on each input feature vector.
[0035] Specifically, the triple discrimination task refers to the attack sample discriminator simultaneously completing the classification or regression of three sub-tasks in the same forward propagation process. By performing multi-task prediction on each input feature vector, richer discriminative features can be extracted and the generalization ability of the model can be enhanced. Here, the triple identification task can include: First, identifying whether the input feature vector belongs to normal traffic or attack traffic. This is a binary classification subtask used to distinguish between normal business behavior and suspicious malicious behavior. For example, identifying Hypertext Transfer Protocol requests as normal and synchronous flood attack traffic as an attack. Second, identifying whether the input feature vector belongs to real network traffic or synthetic traffic. This is also a binary classification subtask used to detect whether the input sample comes from a real network environment or is artificially constructed by an attack sample generator, thereby determining whether the attack sample discriminator can effectively identify the fake samples simulated by the generator. Third, when the input feature vector is determined to be attack traffic, the attack category label of the input feature vector can be identified. This is a multi-classification subtask that can be used to further identify specific attack types. For example, classifying attack traffic as distributed denial-of-service attacks, botnet control commands, ransomware communication, or port scanning behavior, thereby providing more refined attack information for subsequent defense strategies.
[0036] S203, Based on the output of the triple discrimination task, generate the probability discrimination result.
[0037] Furthermore, based on the output probability values of the three sub-tasks in the triple identification task, a comprehensive probability discrimination result can be calculated and generated. This result typically includes three parts: binary classification probabilities of normal versus attack, binary classification probabilities of real versus synthetic traffic, and probability distribution across attack categories. The attack sample discriminator provides a final judgment based on these probability results. For example, when the attack probability is higher than a set threshold and the probability of real traffic is high, the output is real attack traffic along with a specific attack category label and confidence score.
[0038] S104, Perform layered defense actions on the network data stream based on the probability discrimination result.
[0039] Here, after obtaining the probability discrimination result output by the attack sample discriminator, layered defense actions can be performed on the currently captured network data stream based on the network traffic category and confidence level indicated by the probability discrimination result. This avoids applying the same blocking or allowing strategy to all suspicious traffic, thereby improving the utilization efficiency of defense resources and reducing the business impact caused by false alarms. Layered defense actions refer to classifying network data streams into different threat levels based on the attack probability value and confidence value in the probability discrimination result, and performing differentiated processing operations for each level. Specifically, this can include three basic actions: real-time blocking, in-depth data analysis, and allowing normal traffic.
[0040] For example, when the attack sample discriminator outputs a high-confidence real attack probability for a network data stream (i.e., the attack probability is greater than a preset high threshold and the discriminator's confidence in the result exceeds a set value), real-time interception can be performed to directly block the network connection of the traffic, discard related data packets, and send an alarm message to the network administrator. When the discriminator's result is a suspected attack or a low-confidence attack (i.e., the attack probability is between the low and high thresholds, or the discriminator's output confidence is too low to reliably determine), the network data stream can be directed to an isolation analysis system independent of the main network environment, such as a sandbox. The sandbox is a virtual execution environment isolated from the main network, capable of performing in-depth behavioral analysis. In this isolated system, the payload carried by the traffic or the session reconstruction process can be simulated to observe whether malicious behavioral characteristics exist. When the discriminator's result is normal traffic, it indicates that the network data stream's behavioral characteristics highly match the normal traffic model, and there are no known or suspected attack characteristics. In this case, the traffic can be directly allowed to pass through network nodes and reach its destination without any blocking or delay processing.
[0041] In some other embodiments, different additional defense actions can be set according to the business needs and security policies of the actual network environment. For example, in addition to real-time interception of high-confidence attack traffic, the complete attack payload can be recorded for post-event evidence collection. Suspected attack traffic can be rate-limited while being directed to the sandbox to reduce potential harm. Or, a temporary blacklist can be executed to block multiple low-confidence attack traffic from specific Internet Protocol addresses. As long as differentiated threat response can be achieved based on probability judgment results, no specific limitations are made here.
[0042] Understandably, since the attack sample generator and attack sample discriminator described in this application are continuously updated, their parameters are adjusted at the end of each evolutionary cycle based on guidance signals from the dynamic environment simulator. Therefore, the recognition and simulation capabilities of both models can gradually improve as adversarial training progresses. To further improve the efficiency and stability of co-evolution and prevent the generator from forgetting historically successful attack simulation patterns and the discriminator from forgetting previously difficult-to-identify samples, a memory pool can be constructed based on historical experience during adversarial training to store high-value samples and reuse them repeatedly in subsequent cycles. (Refer to...) Figure 3 As shown, the steps S301 to S303 may be included: S301, build and maintain an adversarial memory pool.
[0043] Here, the adversarial memory pool is a data structure capable of storing and retrieving sample data. It can be implemented using a queue or a cache database. It stores high-value synthetic attack samples that have successfully deceived the attack sample discriminator. These are synthetic samples constructed by the attack sample generator during a certain evolutionary cycle and incorrectly classified as real attack traffic or normal traffic by the attack sample discriminator, as well as real traffic samples misclassified by the attack sample discriminator. These samples include false negatives (attack traffic misclassified as normal traffic) and false positives (normal traffic misclassified as attack traffic). Initially, the adversarial memory pool can be empty or pre-store a batch of known typical attack samples and difficult samples, such as high-difficulty attack traffic selected from public datasets or representative false positive / false negative samples accumulated historically. This provides the generator and discriminator with effective historical experience references in the early stages of evolution, accelerating the convergence of the co-evolution process. As each evolution cycle progresses, the dynamic environment simulator writes eligible samples into the memory pool. At the same time, it can use a first-in-first-out or elimination strategy based on sample contribution to control the upper limit of the memory pool capacity, ensuring that the stored samples always have high training value.
[0044] S302, in each evolution cycle, samples are sampled from the adversarial memory pool, and the sampled samples are provided as negative samples to the attack sample generator and as retraining samples to the attack sample discriminator; in the step of adjusting the attack sample generator and the attack sample discriminator, the negative samples are used to update the model parameters of the attack sample generator and the retraining samples are used to update the model parameters of the attack sample discriminator.
[0045] Specifically, the evolution cycle refers to a complete iterative process from the completion of one model adjustment, through online detection and defense, adversarial performance monitoring, and adaptive score calculation, to the end of the next model adjustment. It represents the time interval during which the attack sample generator and attack sample discriminator parameters remain unchanged and operate online. The duration can be set to hours or days, depending on the attack frequency and computational resources of the actual network environment. At the end of each evolution cycle, after the adaptive score is calculated and the evolution guidance signal is generated in the dynamic environment simulator, samples are randomly drawn from the adversarial memory pool according to a preset resampling ratio. These samples are provided as negative samples to the attack sample generator and simultaneously as retraining samples to the attack sample discriminator. For the attack sample generator, negative samples refer to synthetic attack samples that have historically successfully deceived the discriminator. During the learning process, the generator needs to mimic the feature distribution of these samples to generate new attack feature vectors that are more difficult for the discriminator to recognize. For the attack sample discriminator, retraining samples refer to difficult samples that have historically been misjudged. During retraining, the discriminator needs to use these samples, along with their correct labels, as supervisory signals for relearning.
[0046] Furthermore, in the steps of adjusting the attack sample generator and the attack sample discriminator, negative samples can be used to perform additional gradient descent updates on the attack sample generator. That is, the loss of the generator on negative samples is calculated and backpropagated to adjust its network weights so that the feature vector generated by the generator is closer to the distribution of these high-value attack samples. At the same time, the attack sample discriminator can be retrained and updated using retrained samples. That is, the classification loss of the discriminator on these historically difficult samples is calculated and its parameters are updated to strengthen the discriminator's ability to identify such error-prone samples.
[0047] S303, the results of the deep data analysis are fed back to the dynamic environment simulator, stored in the adversarial memory pool as new key attack variant samples, and used to update the key variant sample library for the next evolution cycle.
[0048] Meanwhile, since new and variant attacks constantly emerge in real-world network environments, and the probability discrimination results output by the attack sample discriminator when processing these unknown traffic may have low confidence, after performing deep data analysis on low-confidence suspected attack traffic, the clear conclusions obtained from the deep data analysis can be fed back to the dynamic environment simulator as new key attack variant samples. These samples are then stored in the adversarial memory pool and used to update the key variant sample library used in calculating the discriminant fitness score in the next evolutionary cycle. Here, the initial key variant sample library can use publicly available malicious sample sets or be annotated by experts. In this way, when the dynamic environment simulator calculates the discriminator's key variant sensitivity in subsequent cycles, it can evaluate based on the latest real attack samples, more accurately reflecting the discriminator's defense capability against currently active attack variants, thereby improving the response speed and adaptability of the co-evolutionary framework to new attacks.
[0049] S105, at the end of each evolution cycle, the dynamic environment simulator is used to monitor the adversarial performance data of the attack sample generator and the attack sample discriminator during the evolution cycle, and the generation fitness score of the attack sample generator and the discrimination fitness score of the attack sample discriminator are calculated based on the adversarial performance data; the dynamic environment simulator is used to determine the generation evolution guidance signal for the attack sample generator based on the generation fitness score, and the attack sample generator is adjusted using the generation evolution guidance signal; and the dynamic environment simulator is used to determine the discrimination evolution guidance signal for the attack sample discriminator based on the discrimination fitness score, and the attack sample discriminator is adjusted using the discrimination evolution guidance signal; the process returns to the step of real-time capture of network data streams, and continues execution using the adjusted attack sample generator and the attack sample discriminator.
[0050] Understandably, at the end of each evolutionary cycle, it indicates that the attack sample generator and attack sample discriminator have completed online detection and defense within that cycle, accumulating sufficient adversarial interaction data. At this point, a dynamic environment simulator can be used to monitor the adversarial performance data of the attack sample generator and attack sample discriminator within that cycle. Based on this adversarial performance data, fitness scores and discriminant fitness scores are calculated respectively, serving as quantitative criteria for evaluating the game state and evolutionary direction of both sides. The dynamic environment simulator is a control module independent of the generator and discriminator, not participating in the forward computation and backpropagation of the neural network. It can quantitatively evaluate the performance of the generator and discriminator based on a preset dynamic game judgment function and policy exploration algorithm, thereby generating evolutionary guidance signals to guide the training of both sides in the next cycle.
[0051] Here, adversarial performance data refers to various performance metrics generated by the attack sample generator and attack sample discriminator during the adversarial game within the current evolution cycle. This can include both generative adversarial performance data and discriminative adversarial performance data. Generative adversarial performance data is a set of metrics evaluating the attack sample generator's ability to simulate attacks. It can include the success rate of the synthetic attack feature vectors generated by the attack sample generator in deceiving the current attack sample discriminator, the diversity measure of the generated synthetic attack feature vectors, and the novelty measure of the attacks simulated by the generated synthetic attack feature vectors. Discriminative adversarial performance data is a set of metrics evaluating the attack sample discriminator's classification and recognition capabilities. It can include the attack sample discriminator's baseline classification accuracy for normal traffic and attack traffic in real-time network data streams, the current recognition rate of high-value synthetic attack samples in the adversarial memory pool, and the sensitivity to key attack variant samples in the key variant sample library.
[0052] Specifically, the success rate of the synthetic attack feature vector generated by the attack sample generator in deceiving the current attack sample discriminator refers to the proportion of the attack sample discriminator that incorrectly classifies the synthetic attack feature vector as real attack traffic or normal traffic. This can be obtained by dividing the number of synthetic samples generated by the attack sample generator that were misclassified by the discriminator within the current evolution cycle by the total number of synthetic samples. The diversity measure of the generated synthetic attack feature vector represents the dispersion of the generated samples in the feature space. It is a quality indicator that measures whether the generator produces repetitive or similar attack patterns. This can be obtained by calculating the maximum mean difference between the current generated sample distribution and the historical generated sample distribution. The novelty measure of the attack simulated by the generated synthetic attack feature vector refers to the degree of deviation of the generated sample from the attack types that the discriminator has learned. It can indicate whether the generator has explored new attack features that have not been seen or are difficult to classify. This can be measured by calculating the reconstruction error of the generated sample using an independent novelty detection network. The larger the reconstruction error, the more novel the sample. The baseline classification accuracy of the attack sample discriminator for normal traffic and attack traffic in real-time network data streams can be obtained by dividing the number of correctly classified real network traffic by the discriminator within the current evolution cycle by the total number of real traffic, reflecting the discriminator's basic classification capability. The current recognition rate of high-value synthetic attack samples in the adversarial memory pool represents the proportion of synthetic attack samples that the discriminator can correctly identify in the current cycle that have historically successfully deceived it, reflecting the discriminator's ability to retain historically difficult examples. The sensitivity to identifying key attack variant samples in the key variant sample library refers to the discriminator's recall rate for a small number of key attack variant samples obtained from external threat intelligence or deep data analysis results from the dynamic environment simulator; it is a core indicator for measuring the discriminator's ability to detect new variant attacks.
[0053] Understandably, when calculating the generation fitness score and discriminant fitness score based on the generative adversarial performance data and the discriminant adversarial performance data, a weighted sum can be performed based on their respective indicators. The weight coefficients of each indicator can be dynamically adjusted by the dynamic environment simulator according to the current defense requirements, thus obtaining a quantitative score that comprehensively reflects the overall performance of the generator and discriminator within the current evolutionary cycle. A higher generation fitness score indicates that the attack sample generator performs more balancedly and excellently in the three dimensions of deception ability, sample diversity, and attack novelty; a higher discriminant fitness score indicates that the attack sample discriminator performs more robustly in the three dimensions of basic classification ability, historical difficult case retention ability, and sensitivity to new variants.
[0054] Furthermore, after obtaining the generation fitness score and the discrimination fitness score, the dynamic environment simulator can be used to generate evolutionary guidance signals for the attack sample generator and the attack sample discriminator based on the numerical values of these two scores and the changing trends of various sub-indicators. This will guide both parties to adjust their training objectives and learning strategies in the next evolutionary cycle, so as to obtain a continuously optimized co-evolutionary effect.
[0055] For example, during the generation of evolutionary guidance signals, when the novelty metric in the fitness score is lower than a preset novelty threshold, it indicates that the synthetic attack feature vectors currently generated by the attack sample generator are too similar to attack types already familiar to the discriminator, and have not explored enough new attack patterns. In this case, a dynamic environment simulator can be used to generate instructions that increase the weight of the policy exploration reward term, i.e., increase the coefficient of the policy exploration reward term in the generator's loss function. This serves as an evolutionary guidance signal to guide the attack sample generator to try more unexplored regions in the latent space in the next cycle. Here, the policy exploration reward term encourages the attack sample generator to explore unexplored regions in the latent space. When the feature vectors generated by the generator are far from the feature distribution center of existing attack samples, the reward value increases, thus making the generator more inclined to generate more diverse and novel synthetic attack feature vectors.
[0056] For example, during the process of discriminating evolution guidance signals, when the current recognition rate of high-value synthetic attack samples in the discriminating fitness score is lower than a preset recognition rate threshold, it indicates that the attack sample discriminator has forgotten the attack patterns that have successfully deceived it in the past. To combat memory decay, a dynamic environment simulator can be used to generate instructions to increase the resampling ratio, i.e., increase the proportion of samples drawn from the adversarial memory pool and added to the training batch, serving as a discriminating evolution guidance signal. The resampling ratio indicates the proportion of samples drawn from the adversarial memory pool and added to the attack sample discriminator's training data. A higher ratio means the discriminator will see more historically difficult samples during training, thereby strengthening its ability to remember historical attack patterns.
[0057] Here, during the model tuning process for the attack sample generator and attack sample discriminator, corresponding loss functions can be defined for each. The loss function for the attack sample generator is as follows: It can be defined as: ; in, This represents the probability that the attack sample discriminator will misclassify the synthesized attack feature vector G(z) as a real attack. The larger this value is, the stronger the generator's deception ability. Therefore, taking the negative logarithm in the loss function reduces the loss value when the deception is successful. This is a strategy exploration reward, whose value is the negative logarithm of the average distance between the synthetic attack feature vector and the mainstream cluster centers in the feature space. The greater the distance, the more novel the region explored by the generator. The larger the value, the more it guides the generator to move towards unexplored areas; γ is a weighting coefficient controlled by the dynamic environment simulator. When the novelty metric is low, the dynamic environment simulator will increase the value of γ to increase the proportion of the strategy exploration reward in the total loss. Loss function of attack sample discriminator It can be defined as: ; in, This represents the attack sample discriminator's analysis of real network traffic samples. When performing classification, the probability distribution of its output With real labels The cross-entropy loss between them is used to train the discriminator to accurately distinguish between real normal traffic and real attack traffic; This is represented by the attack sample discriminator's view on the synthetic attack feature vector. When performing classification, the probability distribution of its output With synthetic sample labels The cross-entropy loss between them is used to train the discriminator to distinguish between real and synthetic traffic; This means that when an input sample is determined to be attack traffic. At that time, the attack category probability distribution output by the attack sample discriminator With attack category real label The cross-entropy loss is used to train the discriminator to perform fine-grained family classification of attack traffic.
[0058] Meanwhile, when the generator's adaptive score or the discriminator's adaptive score falls below a preset adaptive score threshold for multiple consecutive evolutionary cycles, it indicates that the generator's or discriminator's performance has reached a bottleneck and cannot be effectively improved through conventional strategy adjustments. In this case, a dynamic environment simulator can be used to generate a partial reset instruction for the model parameters. This instruction performs random perturbations on the model parameters of the generator or discriminator, that is, it randomly initializes only the weights of the last few layers or a certain proportion of neurons in the network. This introduces new optimization directions without completely destroying the model's existing capabilities, helping the model escape local optima. The preset adaptive score threshold is a pre-set score boundary value used to determine whether the generator's or discriminator's performance is below an acceptable range. It can be set based on the average score over historical evolutionary cycles or expert experience; for example, it can be set to 0.6 or 0.65. When the score falls below this threshold for three consecutive cycles, a parameter reset operation is triggered.
[0059] In other embodiments, when generating evolutionary guidance signals, joint adjustments can be made based on the relative magnitudes between the generated fitness score and the discriminant fitness score, as well as the changing trends of their respective sub-indicators. Specifically, the attention bias weights of the discriminator for various attack samples can be dynamically adjusted. For example, when the discriminator's sensitivity to identifying a certain type of attack variant remains low, the dynamic environment simulator can generate instructions to prioritize selecting samples related to that type of attack from the adversarial memory pool to participate in the next training cycle. No specific limitations are made here.
[0060] Furthermore, after adjusting the model parameters of the attack sample generator and the attack sample discriminator, the process can return to step S102 above based on the adjusted model parameters. The adjusted attack sample generator and the attack sample discriminator can then be used to continue to perform real-time capture, parsing, synthesis of attack feature vectors, probability discrimination output, and layered defense actions of the network data stream in the next evolution cycle. This enables the defense model to achieve autonomous and collaborative evolution and long-term adaptive optimization in the dynamic environment of continuous evolution of network attack technologies.
[0061] In this way, through the aforementioned periodic adversarial performance monitoring, adaptive score calculation, evolutionary guidance signal generation, model parameter adjustment, and adversarial memory pool-assisted retraining, this method enables the attack sample generator and attack sample discriminator to alternately improve their simulation and recognition capabilities in each evolutionary cycle, thereby forming a spiral upward process in which the generator continuously explores novel attack patterns and the discriminator continuously strengthens the classification boundary.
[0062] The following describes in detail the intelligent identification and defense method for network attack samples disclosed in this application, using several specific application scenarios and numerical calculations as examples. The specific process may include the following: In the first embodiment, a simulated enterprise network environment is set up for four weeks to verify the actual process of this method in defending against variants of distributed denial-of-service (DDoS) attacks. In the initial stage, the attack sample discriminator is trained using historical conventional DDoS attack samples, and the identification rate for novel slow application-layer DDoS attacks is only 82%. After co-evolutionary training is initiated, within the first week, the attack sample generator successfully simulates various slow attack modes, achieving an initial deception success rate of 68% for the attack sample discriminator. The dynamic environment simulator calculates a generation fitness score of 0.71 and a discriminant fitness score of 0.65. Because the discriminant fitness score was significantly lower than the generator fitness score, and the novelty metric in the generator fitness score was only 0.55, the dynamic environment simulator initiated a strategy adjustment: increasing the weight of the strategy exploration reward term in the attack sample generator from 0.1 to 0.25, and increasing the proportion of difficult samples from the adversarial memory pool in the attack sample discriminator's training data from 20% to 40%. By the end of week 2, the attack sample discriminator's recognition rate for slow attacks improved to 61%, while the attack sample generator subsequently generated attack variants with mixed protocol features. In week 3, the dynamic environment simulator detected that the growth in the sensitivity of key variant sample recognition in the discriminant fitness score stagnated, so a 15% random perturbation of the neuron weights in the last hidden layer of the attack sample discriminator was applied. By the end of week 4, the system reached a relatively optimal game equilibrium. Ultimately, the attack sample discriminator achieved an overall recognition rate of 96.5% for various distributed denial-of-service attacks, and the recognition rate for novel slow attack variants stabilized at 89.3%. In subsequent continuous online evolution, when the defense strategy execution module provides feedback on the in-depth data analysis results of a suspicious traffic that was not immediately blocked, the system generates samples with similar characteristics within 8 hours and completes two cycles of reinforcement training, thereby increasing the subsequent identification rate of this type of attack to over 95%.
[0063] In another application scenario, this method is specifically optimized for identifying encrypted traffic during the data outflow phase of advanced persistent threat (APS) attacks. Initially, the attack sample discriminator's detection rate for encrypted outflow traffic was only 47%. After co-evolution began, the dynamic environment simulator set the weight of the sensitivity for identifying key variant samples in the discriminative fitness score to the highest value of 0.5. During the first five training cycles, the attack sample generator focused on simulating encrypted outflow behavior, but the diversity metric of its generated samples remained below 0.6. The dynamic environment simulator, after analysis, determined that the generation strategy was limited and added an additional diversity constraint term to the attack sample generator. This term directly affects the generator's output layer, encouraging an increase in the standard deviation of the output features. Simultaneously, the dynamic environment simulator introduced 50 newly labeled encrypted outflow samples from an external threat intelligence platform, adding them to the key variant sample library. After 20 evolution cycles, the diversity metric of the encrypted outflow samples generated by the attack sample generator improved to 0.82. Under enhanced training, the attack sample discriminator's sensitivity to identifying key variant samples in its discriminative fitness score improved from 0.48 to 0.79. Specifically, the attack sample discriminator no longer relies solely on the single feature of payload size, but instead integrates multi-dimensional temporal features such as connection duration and traffic burst intervals for decision-making. In actual testing, this method improved the detection rate of this type of encrypted outgoing traffic from the initial 47% to 81.5%, while maintaining a false positive rate below 0.3%.
[0064] This method further illustrates the decision-making logic of the dynamic environment simulator through a specific numerical calculation: Assume that after the t-th training period, the dynamic environment simulator collects the following judgment indicators: the deception success rate of the attack sample generator is 0.75, the diversity metric is 0.80, and the novelty metric is 0.60. The baseline classification accuracy of the attack sample discriminator is 0.93, the current recognition rate of high-value synthetic attack samples in the adversarial memory pool is 0.78, and the recognition sensitivity of key variant samples is 0.70. The weight configuration adopted by the dynamic environment simulator in the current period is as follows: in the generation fitness score, the weight of the deception success rate is 0.5, the weight of the diversity metric is 0.25, and the weight of the novelty metric is 0.25; in the discriminant fitness score, the weight of the baseline classification accuracy is 0.4, the weight of the historical high-deception sample recognition rate is 0.3, and the weight of the key variant recognition sensitivity is 0.3. Based on the above weights, the generated fitness score is calculated as follows: 0.5×0.75+0.25×0.80+0.25×0.60=0.375+0.20+0.15=0.725. The discriminative fitness score is calculated as follows: 0.4×0.93+0.3×0.78+0.3×0.70=0.372+0.234+0.21=0.816. The dynamic environment simulator compares the generated fitness score and the discriminative fitness score, and finds that the discriminative fitness score is higher than the generated fitness score, and the novelty metric in the generated fitness score is the lowest among the three. According to the preset policy exploration algorithm rules, when the novelty metric is lower than the threshold of 0.65 and is the lowest-scoring item, the dynamic environment simulator generates the following evolutionary guidance signal: First, increase the weight of the policy exploration reward item in the attack sample generator loss function by 30%. Second, in the next cycle, normal samples that were previously misclassified with high confidence by the attack sample discriminator are preferentially selected from the adversarial memory pool and added to the training targets of the attack sample generator at a mixing ratio of 15%, forcing the attack sample generator to learn more covert attack features. In this way, through this quantitative, rule-based feedback control, this method achieves targeted and efficient co-evolution.
[0065] The intelligent identification and defense method, system, device, and medium for network attack samples provided in this disclosure monitor the adversarial performance data of the attack sample generator and the attack sample discriminator during the evolution cycle through a dynamic environment simulator and calculate the fitness score. Combined with the evolution guidance signal, the generator and discriminator are periodically adjusted, realizing the continuous evolution and collaborative optimization of the generator and discriminator in the process of network attack sample identification. This improves the adaptability of network attack defense to mutated samples and the accuracy of the execution of layered defense actions.
[0066] Those skilled in the art will understand that, in the above-described method of the specific implementation, the order in which each step is written does not imply a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined by its function and possible internal logic.
[0067] Based on the same inventive concept, this disclosure also provides a network attack sample intelligent identification and defense system corresponding to the network attack sample intelligent identification and defense method. Since the principle of the system in this disclosure for solving the problem is similar to the network attack sample intelligent identification and defense method described above in this disclosure, the implementation of the system can refer to the implementation of the method, and the repeated parts will not be described again.
[0068] Reference Figure 4 The diagram shown is a schematic of an intelligent identification and defense system for network attack samples provided in an embodiment of this disclosure. The system includes: The data capture module is used to capture network data streams in real time and parse the network data streams to obtain real-time traffic feature vectors; The sample recognition engine includes a trained attack sample generator, an attack sample discriminator, and a dynamic environment simulator. The attack sample generator is used to generate a synthetic attack feature vector; the attack sample discriminator is used to output a probability discrimination result based on the real-time traffic feature vector and the synthetic attack feature vector. The dynamic environment simulator is configured to monitor the adversarial performance data of the attack sample generator and the attack sample discriminator during the evolution cycle at the end of each evolution cycle, and calculate the generation fitness score of the attack sample generator and the discrimination fitness score of the attack sample discriminator based on the adversarial performance data; determine a generation evolution guidance signal for the attack sample generator based on the generation fitness score, and adjust the attack sample generator using the generation evolution guidance signal; and determine a discrimination evolution guidance signal for the attack sample discriminator based on the discrimination fitness score, and adjust the attack sample discriminator using the discrimination evolution guidance signal. The sample recognition engine uses the adjusted attack sample generator and the adjusted attack sample discriminator to perform intelligent recognition of network attack samples in the next evolution cycle; The defense execution module is used to perform layered defense actions on the network data stream based on the probability discrimination result; wherein, the layered defense actions include real-time interception, deep data analysis, and normal traffic passage.
[0069] In some possible embodiments, the sample recognition engine is specifically used for: The real-time traffic feature vector and the synthetic attack feature vector are input together into the attack sample discriminator; The attack sample discriminator performs a triple identification task on each input feature vector. The triple identification task includes: identifying whether the input feature vector belongs to normal traffic or attack traffic, identifying whether the input feature vector belongs to real network traffic or synthetic traffic, and identifying the attack category label of the input feature vector when the input feature vector is determined to be attack traffic. The probability discrimination result is generated based on the output of the triple discrimination task.
[0070] In some possible embodiments, the defense execution module is further configured to: Construct and maintain an adversarial memory pool, which is used to store high-value synthetic attack samples that have successfully deceived the attack sample discriminator, as well as real traffic samples that have been misjudged by the attack sample discriminator. In each evolution cycle, samples are sampled from the adversarial memory pool. The sampled samples are provided as negative samples to the attack sample generator and as retraining samples to the attack sample discriminator. In the step of adjusting the attack sample generator and the attack sample discriminator, the negative samples are used to update the model parameters of the attack sample generator and the retraining samples are used to update the model parameters of the attack sample discriminator. The results of the deep data analysis are fed back to the dynamic environment simulator, stored in the adversarial memory pool as new key attack variant samples, and used to update the key variant sample library for the next evolution cycle.
[0071] In some possible embodiments, the dynamic environment simulator is specifically used for: Collect the adversarial performance data generated by the attack sample generator within the evolution cycle; wherein, the adversarial performance data includes: the success rate of the synthetic attack feature vector generated by the attack sample generator in deceiving the current attack sample discriminator, the diversity measure of the generated synthetic attack feature vector, and the novelty measure of the attack simulated by the generated synthetic attack feature vector. Based on the generated adversarial performance data, the generation fitness score of the attack sample generator is calculated; Collect the adversarial performance data of the attack sample discriminator within the evolution cycle; wherein, the adversarial performance data includes: the baseline classification accuracy of the attack sample discriminator for normal traffic and attack traffic in real-time network data stream, the current recognition rate of high-value synthetic attack samples in the adversarial memory pool, and the recognition sensitivity of key attack variant samples in the key variant sample library. Based on the discriminant performance data, the discriminant fitness score of the attack sample discriminator is calculated.
[0072] In some possible embodiments, the dynamic environment simulator is specifically used for: When the novelty metric in the generated fitness score is lower than a preset novelty threshold, an instruction to increase the weight of the strategy exploration reward item is generated as a guidance signal for the generation evolution; the strategy exploration reward item is used to encourage the attack sample generator to explore unexplored regions in the potential space to generate synthetic attack feature vectors with higher novelty.
[0073] In some possible embodiments, the dynamic environment simulator is specifically used for: When the current recognition rate of the high-value synthetic attack sample in the discrimination fitness score is lower than a preset recognition rate threshold, an instruction to increase the resampling ratio is generated as a discrimination evolution guidance signal; the resampling ratio indicates the proportion of samples drawn from the adversarial memory pool and added to the attack sample discriminator training data; Furthermore, when the generated fitness score and / or the discriminant fitness score are lower than a preset fitness score threshold for multiple consecutive evolution cycles, a partial reset instruction for model parameters is generated to perform random perturbation on the model parameters of the attack sample generator and / or the attack sample discriminator.
[0074] Based on the same inventive concept, this disclosure also provides a network attack sample intelligent identification and defense device corresponding to the network attack sample intelligent identification and defense method. Since the principle of the device in this disclosure for solving the problem is similar to the network attack sample intelligent identification and defense method described above in this disclosure, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be described again.
[0075] Reference Figure 5 The diagram shown is a schematic of a network attack sample intelligent identification and defense device 500 provided in an embodiment of this disclosure. The device includes: The model acquisition module 501 is used to acquire the trained attack sample generator and attack sample discriminator. The data acquisition module 502 is used to capture network data streams in real time and parse the network data streams to obtain real-time traffic feature vectors; The model application module 503 is used to generate a synthetic attack feature vector using the attack sample generator, and to output a probability discrimination result using the attack sample discriminator based on the real-time traffic feature vector and the synthetic attack feature vector. The defense execution module 504 is used to perform layered defense actions on the network data stream based on the probability discrimination result; wherein, the layered defense actions include real-time interception, deep data analysis, and normal traffic passage. The model adjustment module 505 is configured to: at the end of each evolution cycle, monitor the adversarial performance data of the attack sample generator and the attack sample discriminator during the evolution cycle using a dynamic environment simulator, and calculate the generation fitness score of the attack sample generator and the discrimination fitness score of the attack sample discriminator based on the adversarial performance data; determine a generation evolution guidance signal for the attack sample generator based on the generation fitness score using the dynamic environment simulator, and adjust the attack sample generator using the generation evolution guidance signal; determine a discrimination evolution guidance signal for the attack sample discriminator based on the discrimination fitness score using the dynamic environment simulator, and adjust the attack sample discriminator using the discrimination evolution guidance signal; and return to the step of real-time capture of network data streams, and continue execution using the adjusted attack sample generator and the attack sample discriminator.
[0076] In some possible embodiments, the model application module 503 is specifically used for: The real-time traffic feature vector and the synthetic attack feature vector are input together into the attack sample discriminator; The attack sample discriminator performs a triple identification task on each input feature vector. The triple identification task includes: identifying whether the input feature vector belongs to normal traffic or attack traffic, identifying whether the input feature vector belongs to real network traffic or synthetic traffic, and identifying the attack category label of the input feature vector when the input feature vector is determined to be attack traffic. The probability discrimination result is generated based on the output of the triple discrimination task.
[0077] In some possible embodiments, the defense execution module 504 is further configured to: Construct and maintain an adversarial memory pool, which is used to store high-value synthetic attack samples that have successfully deceived the attack sample discriminator, as well as real traffic samples that have been misjudged by the attack sample discriminator. In each evolution cycle, samples are sampled from the adversarial memory pool. The sampled samples are provided as negative samples to the attack sample generator and as retraining samples to the attack sample discriminator. In the step of adjusting the attack sample generator and the attack sample discriminator, the negative samples are used to update the model parameters of the attack sample generator and the retraining samples are used to update the model parameters of the attack sample discriminator. The results of the deep data analysis are fed back to the dynamic environment simulator, stored in the adversarial memory pool as new key attack variant samples, and used to update the key variant sample library for the next evolution cycle.
[0078] In some possible embodiments, the model adjustment module 505 is specifically used for: Collect the adversarial performance data generated by the attack sample generator within the evolution cycle; wherein, the adversarial performance data includes: the success rate of the synthetic attack feature vector generated by the attack sample generator in deceiving the current attack sample discriminator, the diversity measure of the generated synthetic attack feature vector, and the novelty measure of the attack simulated by the generated synthetic attack feature vector. Based on the generated adversarial performance data, the generation fitness score of the attack sample generator is calculated; Collect the adversarial performance data of the attack sample discriminator within the evolution cycle; wherein, the adversarial performance data includes: the baseline classification accuracy of the attack sample discriminator for normal traffic and attack traffic in real-time network data stream, the current recognition rate of high-value synthetic attack samples in the adversarial memory pool, and the recognition sensitivity of key attack variant samples in the key variant sample library. Based on the discriminant performance data, the discriminant fitness score of the attack sample discriminator is calculated.
[0079] In some possible embodiments, the model adjustment module 505 is specifically used for: When the novelty metric in the generated fitness score is lower than a preset novelty threshold, an instruction to increase the weight of the strategy exploration reward item is generated as a guidance signal for the generation evolution; the strategy exploration reward item is used to encourage the attack sample generator to explore unexplored regions in the potential space to generate synthetic attack feature vectors with higher novelty.
[0080] In some possible embodiments, the model adjustment module 505 is specifically used for: When the current recognition rate of the high-value synthetic attack sample in the discrimination fitness score is lower than a preset recognition rate threshold, an instruction to increase the resampling ratio is generated as a discrimination evolution guidance signal; the resampling ratio indicates the proportion of samples drawn from the adversarial memory pool and added to the attack sample discriminator training data; Furthermore, when the generated fitness score and / or the discriminant fitness score are lower than a preset fitness score threshold for multiple consecutive evolution cycles, a partial reset instruction for model parameters is generated to perform random perturbation on the model parameters of the attack sample generator and / or the attack sample discriminator.
[0081] Based on the same technical concept, this disclosure also provides a computer device. (See also...) Figure 6 The diagram shows the structure of a computer device 600 provided in this embodiment of the present disclosure, including a processor 601, a memory 602, and a bus 603. The memory 602 stores execution instructions and includes a main memory 6021 and an external memory 6022. The main memory 6021, also called internal memory, is used to temporarily store computational data in the processor 601 and data exchanged with external memory 6022 such as a hard disk. The processor 601 exchanges data with the external memory 6022 through the main memory 6021.
[0082] In this embodiment, the memory 602 is specifically used to store application code that executes the solution of this application, and its execution is controlled by the processor 601. That is, when the computer device 600 is running, the processor 601 communicates with the memory 602 through the bus 603, so that the processor 601 executes the application code stored in the memory 602, and then executes the method described in any of the foregoing embodiments.
[0083] The memory 602 may be, but is not limited to, random access memory (RAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), etc.
[0084] Processor 601 may be an integrated circuit chip with signal processing capabilities. The aforementioned processor can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this invention. The general-purpose processor can be a microprocessor or any conventional processor.
[0085] It is understood that the structures illustrated in the embodiments of this application do not constitute a specific limitation on the computer device 600. In other embodiments of this application, the computer device 600 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0086] This disclosure also provides a computer-readable storage medium storing a computer program. When executed by a processor, the computer program performs the steps of the intelligent identification and defense method for network attack samples described in the above-described method embodiments. The storage medium can be a volatile or non-volatile computer-readable storage medium.
[0087] This disclosure also provides a computer program product carrying program code. The program code includes instructions that can be used to execute the steps of the intelligent identification and defense method for network attack samples described in the above method embodiments. For details, please refer to the above method embodiments, which will not be repeated here.
[0088] The aforementioned computer program product can be implemented through hardware, software, or a combination thereof. In one optional embodiment, the computer program product is specifically embodied in a computer storage medium; in another optional embodiment, the computer program product is specifically embodied in a software product, such as a software development kit (SDK), etc.
[0089] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems and devices described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. In the several embodiments provided in this disclosure, it should be understood that the disclosed systems and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division; in actual implementation, there may be other division methods. Furthermore, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection may be through some communication interfaces; the indirect coupling or communication connection of devices or units may be electrical, mechanical, or other forms.
[0090] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0091] In addition, the functional units in the various embodiments of this disclosure can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0092] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a processor-executable, non-volatile, computer-readable storage medium. Based on this understanding, the technical solution of this disclosure, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this disclosure. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0093] Finally, it should be noted that the above-described embodiments are merely specific implementations of this disclosure, used to illustrate the technical solutions of this disclosure, and not to limit them. The protection scope of this disclosure is not limited thereto. Although this disclosure has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some of the technical features within the scope of the technology disclosed in this disclosure. Such modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this disclosure, and should all be covered within the protection scope of this disclosure.
Claims
1. A method for intelligent identification and defense of network attack samples, characterized in that, include: Obtain the trained attack sample generator and attack sample discriminator; Real-time capture of network data streams and parsing of the network data streams to obtain real-time traffic feature vectors; The attack sample generator generates a synthetic attack feature vector, and the attack sample discriminator outputs a probability discrimination result based on the real-time traffic feature vector and the synthetic attack feature vector. Based on the probability discrimination results, layered defense actions are performed on the network data stream; wherein, the layered defense actions include real-time interception, deep data analysis, and normal traffic passage. At the end of each evolution cycle, the adversarial performance data of the attack sample generator and the attack sample discriminator during the evolution cycle are monitored using a dynamic environment simulator, and the generation fitness score of the attack sample generator and the discrimination fitness score of the attack sample discriminator are calculated based on the adversarial performance data. The dynamic environment simulator determines a generation evolution guidance signal for the attack sample generator based on the generation fitness score, and adjusts the attack sample generator using the generation evolution guidance signal; and the dynamic environment simulator determines a discrimination evolution guidance signal for the attack sample discriminator based on the discrimination fitness score, and adjusts the attack sample discriminator using the discrimination evolution guidance signal. Return to the step of capturing network data streams in real time, and continue execution using the adjusted attack sample generator and attack sample discriminator.
2. The method according to claim 1, characterized in that, The step of using the attack sample discriminator to output a probability discrimination result based on the real-time traffic feature vector and the synthetic attack feature vector includes: The real-time traffic feature vector and the synthetic attack feature vector are input together into the attack sample discriminator; The attack sample discriminator performs a triple identification task on each input feature vector. The triple identification task includes: identifying whether the input feature vector belongs to normal traffic or attack traffic, identifying whether the input feature vector belongs to real network traffic or synthetic traffic, and identifying the attack category label of the input feature vector when the input feature vector is determined to be attack traffic. The probability discrimination result is generated based on the output of the triple discrimination task.
3. The method according to claim 2, characterized in that, After performing layered defense actions on the network data stream based on the probability discrimination result, the method further includes: Construct and maintain an adversarial memory pool, which is used to store high-value synthetic attack samples that have successfully deceived the attack sample discriminator, as well as real traffic samples that have been misjudged by the attack sample discriminator. In each evolution cycle, samples are sampled from the adversarial memory pool. The sampled samples are provided as negative samples to the attack sample generator and as retraining samples to the attack sample discriminator. In the step of adjusting the attack sample generator and the attack sample discriminator, the negative samples are used to update the model parameters of the attack sample generator and the retraining samples are used to update the model parameters of the attack sample discriminator. The results of the deep data analysis are fed back to the dynamic environment simulator, stored in the adversarial memory pool as new key attack variant samples, and used to update the key variant sample library for the next evolution cycle.
4. The method according to claim 3, characterized in that, The step of calculating the generation fitness score of the attack sample generator and the discrimination fitness score of the attack sample discriminator based on the adversarial performance data includes: Collect the adversarial performance data generated by the attack sample generator within the evolution cycle; wherein, the adversarial performance data includes: the success rate of the synthetic attack feature vector generated by the attack sample generator in deceiving the current attack sample discriminator, the diversity measure of the generated synthetic attack feature vector, and the novelty measure of the attack simulated by the generated synthetic attack feature vector. Based on the generated adversarial performance data, the generation fitness score of the attack sample generator is calculated; Collect the adversarial performance data of the attack sample discriminator within the evolution cycle; wherein, the adversarial performance data includes: the baseline classification accuracy of the attack sample discriminator for normal traffic and attack traffic in real-time network data stream, the current recognition rate of high-value synthetic attack samples in the adversarial memory pool, and the recognition sensitivity of key attack variant samples in the key variant sample library. Based on the discriminant performance data, the discriminant fitness score of the attack sample discriminator is calculated.
5. The method according to claim 4, characterized in that, The step of determining the generation evolution guidance signal for the attack sample generator based on the generation fitness score using the dynamic environment simulator includes: When the novelty metric in the generated fitness score is lower than a preset novelty threshold, an instruction to increase the weight of the strategy exploration reward item is generated using a dynamic environment simulator as a guidance signal for the generation evolution. The strategy exploration reward item is used to encourage the attack sample generator to explore unexplored areas in the potential space to generate synthetic attack feature vectors with higher novelty.
6. The method according to claim 4, characterized in that, The step of using the dynamic environment simulator to determine the discriminative evolution guidance signal for the attack sample discriminator based on the discriminative fitness score includes: When the current recognition rate of the high-value synthetic attack sample in the discrimination fitness score is lower than the preset recognition rate threshold, the dynamic environment simulator is used to generate an instruction to increase the resampling ratio as a discrimination evolution guidance signal; the resampling ratio indicates the proportion of samples drawn from the adversarial memory pool and added to the attack sample discriminator training data; Furthermore, when the generated fitness score and / or the discriminant fitness score are lower than the preset fitness score threshold for multiple consecutive evolution cycles, the dynamic environment simulator generates a partial reset instruction for the model parameters, which is used to perform random perturbation on the model parameters of the attack sample generator and / or the attack sample discriminator.
7. A network attack sample intelligent identification and defense system, characterized in that, include: The data capture module is used to capture network data streams in real time and parse the network data streams to obtain real-time traffic feature vectors; The sample recognition engine includes a trained attack sample generator, an attack sample discriminator, and a dynamic environment simulator. The attack sample generator is used to generate synthetic attack feature vectors; The attack sample discriminator is used to output a probability discrimination result based on the real-time traffic feature vector and the synthetic attack feature vector; The dynamic environment simulator is used to monitor the adversarial performance data of the attack sample generator and the attack sample discriminator during the evolution cycle at the end of each evolution cycle, and to calculate the generation fitness score of the attack sample generator and the discrimination fitness score of the attack sample discriminator based on the adversarial performance data. Furthermore, a generation evolution guidance signal for the attack sample generator is determined based on the generation fitness score, and the attack sample generator is adjusted using the generation evolution guidance signal; Furthermore, a discrimination evolution guidance signal is determined for the attack sample discriminator based on the discrimination fitness score, and the attack sample discriminator is adjusted using the discrimination evolution guidance signal; The sample recognition engine uses the adjusted attack sample generator and the adjusted attack sample discriminator to perform intelligent recognition of network attack samples in the next evolution cycle; The defense execution module is used to perform layered defense actions on the network data stream based on the probability discrimination result; wherein, the layered defense actions include real-time interception, deep data analysis, and normal traffic passage.
8. A network attack sample intelligent identification and defense device, characterized in that, include: The model acquisition module is used to acquire trained attack sample generators and attack sample discriminators. The data acquisition module is used to capture network data streams in real time and parse the network data streams to obtain real-time traffic feature vectors. The model application module is used to generate a synthetic attack feature vector using the attack sample generator, and to output a probability discrimination result using the attack sample discriminator based on the real-time traffic feature vector and the synthetic attack feature vector. The defense execution module is used to perform layered defense actions on the network data stream based on the probability discrimination result; wherein, the layered defense actions include real-time interception, deep data analysis, and normal traffic passage. The model adjustment module is configured to: at the end of each evolution cycle, monitor the adversarial performance data of the attack sample generator and the attack sample discriminator during the evolution cycle using a dynamic environment simulator, and calculate the generation fitness score of the attack sample generator and the discrimination fitness score of the attack sample discriminator based on the adversarial performance data; determine a generation evolution guidance signal for the attack sample generator based on the generation fitness score using the dynamic environment simulator, and adjust the attack sample generator using the generation evolution guidance signal; determine a discrimination evolution guidance signal for the attack sample discriminator based on the discrimination fitness score using the dynamic environment simulator, and adjust the attack sample discriminator using the discrimination evolution guidance signal; and return to the step of real-time capture of network data streams, and continue execution using the adjusted attack sample generator and the attack sample discriminator.
9. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 6.
10. A computer device, comprising a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method of any one of claims 1 to 6.