Computer network security detection method and system
By generating priority classification numbers through quantum hashing and reinforcement learning models, and combining quantum neural networks and lightweight convolutional neural networks for detection, the problem of feature forgery and insecure transmission in traditional network security detection under quantum computing environments is solved, achieving efficient and intelligent network security detection.
Patent Information
- Application Number
- CN202511104131.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-07
- Publication Date
- 2025-11-07
AI Technical Summary
When faced with the challenges of quantum computing, traditional computer network security detection systems suffer from weakened collision resistance of hash algorithms, and the loss of security foundation for feature extraction and transmission encryption, allowing attackers to easily forge features or eavesdrop on data.
Feature extraction is performed using a quantum hashing algorithm, which is combined with a reinforcement learning model to generate priority classification numbers. Detection is then performed using a quantum neural network and a lightweight convolutional neural network to construct a hierarchical detection mechanism. This mechanism is then combined with an adaptive response mechanism to form a closed-loop system.
It enhances the resistance and uniqueness of features, achieves efficient resource allocation, improves the intelligence and dynamic adaptability of network security detection, reduces false positives and false negatives, and strengthens the system's practical defense capabilities.
Smart Images

Figure CN120915531A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer network security, and relates to a computer network security detection method and system. BACKGROUND
[0002] With the acceleration of the global digitalization process, cyberspace has become the "fifth territory" of national sovereignty extension, but the technical evolution of network attacks and the breakthrough development of quantum computing are making the traditional computer network security detection system fall into multiple dilemmas. The current network attacks show three trends: the normalization of APT (Advanced Persistent Threat) attacks, the quantumization of attack tools, and the upgrading of threat stealth technology.
[0003] However, the traditional detection system exposes fatal short boards due to the limitations of the underlying architecture: quantum computing subverts the traditional security cornerstone, the traditional hash algorithm (such as SHA-256) greatly weakens the collision resistance in the face of Grover algorithm of quantum computer, Shor algorithm threatens the RSA encryption system, resulting in the loss of security foundation of feature extraction and transmission encryption, and attackers can easily fake features or eavesdrop data. SUMMARY
[0004] Therefore, the embodiments of the present application provide a computer network security detection method and system, which at least solve the problem of feature forgery and insecure transmission caused by the threat of quantum algorithm to the traditional encryption and feature extraction mechanism.
[0005] The technical scheme of the embodiments of the present application is as follows: In a first aspect, the embodiments of the present application provide a computer network security detection method applied to a computer network security detection system, the system comprising a data source layer, a data collection layer, a data processing layer, a priority classification layer, and a detection execution layer, and the method comprising: The data collection layer acquires network data of the data source layer and encrypts, and sends the encrypted network data to the data processing layer; the data processing layer extracts features from the encrypted network data through a quantum hash coding algorithm to obtain a target quantum feature vector, and sends the target quantum feature vector to the priority classification layer; the priority classification layer generates a priority classification number based on the target quantum feature vector, event type, historical response record, and system load through a reinforcement learning model, and sends the priority classification number and the target quantum feature vector to the detection execution layer, the priority classification number comprising a priority level and a sub-classification number; and the detection execution layer calls a corresponding detection model based on the priority classification number to detect the target quantum feature vector to obtain a detection result and a confidence level, the detection result being used to represent the security attribute of the network data.
[0006] In a second aspect, the embodiments of the present application provide a computer network security detection system, comprising: a data source layer, a data collection layer, a data processing layer, a priority classification layer and a detection execution layer, wherein: The data collection layer is configured to acquire network data of the data source layer and encrypt the network data, and send the encrypted network data to the data processing layer; the data processing layer is configured to perform feature extraction on the encrypted network data by using a quantum hash coding algorithm to obtain a target quantum feature vector, and send the target quantum feature vector to the priority classification layer; the priority classification layer is configured to generate a priority classification number based on the target quantum feature vector, an event type, a historical response record and a system load by using a reinforcement learning model, and send the priority classification number and the target quantum feature vector to the detection execution layer, wherein the priority classification number comprises a priority level and a sub-classification number; and the detection execution layer is configured to perform detection on the target quantum feature vector based on the priority classification number by using a corresponding detection model to obtain a detection result and a confidence level, wherein the detection result is used to represent a security attribute of the network data.
[0007] The technical scheme provided by the embodiments of the present application has at least the following beneficial effects: The quantum hash coding technology is used to realize secure extraction of data features, thereby improving the attack resistance and uniqueness of the features; the reinforcement learning algorithm is used to dynamically generate a priority classification number, thereby constructing a hierarchical detection mechanism, realizing efficient allocation of resources, improving response efficiency and classification efficiency, and improving the intelligence and dynamic adaptability of network security detection. DETAILED DESCRIPTION
[0008] In order to more clearly illustrate the technical scheme in the embodiments of the present application, the following will briefly introduce the drawings needed in the embodiment description. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor. Figure 1 A flowchart of a computer network security detection method provided by the embodiments of the present application; Figure 2 An architecture diagram of a computer network security detection system provided by the embodiments of the present application; Figure 3 A working flowchart of a computer network security detection system provided by the embodiments of the present application; Figure 4 A flowchart of a cloud data center network security detection method provided by the embodiments of the present application; Figure 5A flowchart of a method for detecting network security of an Internet of Things smart home provided by an embodiment of the present application is shown in the figure. Figure 6 A flowchart of a method for detecting network security of a mobile operator network provided by an embodiment of the present application is shown in the figure. Figure 7 A component structure diagram of a computer network security detection system provided by an embodiment of the present application is shown in the figure. DETAILED DESCRIPTION
[0009] To make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described below in connection with the drawings in the embodiments of the present application. Obviously, the described embodiments are some of the embodiments of the present application, but not all the embodiments of the present application. The following embodiments are used to explain the present application, but not to limit the scope of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of the present application.
[0010] In the following description, “some embodiments” are described, which describe a subset of all possible embodiments, but it can be understood that “some embodiments” can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.
[0011] It should be noted that the terms “first\second\third” involved in the embodiments of the present application are only to distinguish similar objects, and do not represent a specific order of the objects. It can be understood that “first\second\third” can be interchanged with a specific order or sequence as allowed, so that the embodiments of the present application described here can be implemented in an order other than that illustrated or described here.
[0012] Those skilled in the art can understand that, unless otherwise defined, all terms (including technical terms and scientific terms) used herein have the same meaning as generally understood by those of ordinary skill in the art to which the embodiments of the present application belong. It should also be understood that terms such as those defined in a general dictionary should be understood to have meanings consistent with those in the context of the prior art, and should not be interpreted in an idealized or overly formal sense unless specifically defined as such.
[0013] The embodiment of the present application provides a computer network security detection method, which is applied to an electronic device. The electronic device includes but is not limited to a mobile phone, a notebook computer, a tablet computer, a palm internet device, a multimedia device, a streaming media device, a mobile internet device, a wearable device or other types of electronic devices. The function realized by the method can be realized by calling program code in the processor of the electronic device, and of course the program code can be saved in a computer storage medium. Therefore, the electronic device at least includes a processor and a storage medium. The processor can be used for processing the computer network security detection process, and the storage can be used for storing the data required in the computer network security detection process and the generated data.
[0014] Figure 1 A flowchart of a computer network security detection method provided by the embodiment of the present application is provided, the method is applied to a computer network security detection system, the system includes a data source layer, a data collection layer, a data processing layer, a priority classification layer and a detection execution layer, as shown in Figure 1 The method at least includes the following steps: Step S110, the data collection layer acquires network data of the data source layer and encrypts, and sends the encrypted network data to the data processing layer; As shown in Figure 2 The data source layer is a physical source of original data, and can provide a detection object for the system. The data source layer can include network equipment (such as a router, a switch and the like), a user terminal (such as a PC (personal computer), a mobile phone and the like) and an Internet of Things device (such as a smart sensor, a camera and the like). The network equipment generates network flow data (also referred to as flow data), the user terminal generates user behavior data (also referred to as behavior data, such as access records, operation logs and the like), and the Internet of Things device generates device running log data (also referred to as log data). These data are collectively referred to as network data.
[0015] As shown in Figure 2 The data collection layer is deployed at a network key node (such as a boundary router, a terminal equipment interface). The data collection layer includes a quantum data collection module. The quantum data collection module captures original data (i.e. network data) from the data source layer, encrypts the network data, and outputs the encrypted network data. The encrypted network data can also be referred to as quantum encrypted data. The encrypted network data can be transmitted to the data processing layer.
[0016] Step S120, the data processing layer extracts features from the encrypted network data by using a quantum hash coding algorithm, obtains a target quantum feature vector, and sends the target quantum feature vector to the priority classification layer; As shown inFigure 2 As shown, the data processing layer includes a quantum hash preprocessing module, which can perform a quantum hash encoding (QHE) algorithm on the encrypted network data to extract features from the encrypted network data and generate a target quantum feature vector, laying a foundation for subsequent processing.
[0017] In step S130, the priority classification layer generates a priority classification number based on the target quantum feature vector, event type, historical response record, and system load through a reinforcement learning model, and sends the priority classification number and the target quantum feature vector to the detection execution layer. The priority classification number includes a priority level and a sub-classification number. The priority classification number can include a main classification and a sub-classification. The main classification P1 / P2 / P3 corresponds to high, medium, and low priority levels, which is a preliminary division of the emergency level of network security events. The sub-classification C1-C5 further classifies attack types in a fine-grained manner. For example, in the high priority level, P1-C1 may represent a specific type of high threat attack, P1-C2 represents another type, and so on. This hierarchical classification method can more accurately manage the priority of different types of security events.
[0018] As shown in Figure 2 The priority classification layer can include a reinforcement learning priority classification module for priority determination and detection. The reinforcement learning model can be deployed in the reinforcement learning priority classification module. The reinforcement learning model can be a Double DQN (Double Deep Q-Network) model. The target quantum feature vector, event type (one-hot encoding), historical response record (the last 10 results), and system load (CPU / memory utilization) are input into the Double DQN model. The model outputs a priority classification number (e.g., P1-C3, where P1 is high priority and C3 is a ransomware sub-class). The priority classification layer transmits the priority classification number and the target quantum feature vector to the detection execution layer.
[0019] In step S140, the detection execution layer calls a corresponding detection model based on the priority classification number to detect the target quantum feature vector, obtaining a detection result and a confidence level. The detection result represents the security properties of the network data.
[0020] As shown in Figure 2As shown, the detection execution layer includes an intelligent detection engine module, which can dynamically call a detection model according to a priority classification number, output a detection result and a confidence level; the detection model can be a QNN (Quantum Neural Network) or a CNN (Convolutional Neural Network), etc., and the detection result can include normal events, suspicious events and attack events, as well as corresponding confidence levels.
[0021] In the above embodiments, the data processing layer generates a target quantum feature vector using a quantum hash coding algorithm, which can resist attacks by Grover's algorithm in quantum computing and avoid feature forgery compared with traditional hash algorithms (such as SHA-256), thereby ensuring the anti-collision property and uniqueness of feature extraction from the bottom; the data collection layer first encrypts the network data before transmission, and in combination with the security of the quantum feature vector, forms a dual security mechanism of "encrypted transmission + quantum-level feature protection", solves the problem of insecure transmission of the traditional encryption system in the face of Shor's algorithm, and prevents attackers from eavesdropping or tampering with data; the priority classification layer not only bases on the target quantum feature vector, but also fuses multi-dimensional information such as event type, historical response record and system load, dynamically generates a priority classification number through reinforcement learning, can adapt to the dynamic evolution of attacks in real time, and reduces the situation that high-risk events are misjudged as low priority; in combination with the historical response record, a "feedback-optimization" closed loop is formed to continuously improve the classification accuracy and solve the contradiction between static mechanism and dynamic attack; the detection execution layer calls the corresponding detection model according to the priority classification number, realizes a differentiated detection strategy, and can solve the problem of insufficient adaptability of static rules to new attacks.
[0022] The quantum hash coding technology is used to realize secure extraction of data features, which improves the anti-attack property and uniqueness of the features; the reinforcement learning algorithm is used to dynamically generate a priority classification number, a hierarchical detection mechanism is constructed, the efficient allocation of resources is realized, the response efficiency and classification efficiency are improved, and the intelligence and dynamic adaptability of network security detection are improved.
[0023] In some embodiments, the system further includes a response execution layer, and the method further includes: Step S150, the detection execution layer sends the detection result and the confidence level to the response execution layer; Step S160, the response execution layer executes a corresponding response strategy based on the detection result and the confidence level, and obtains a corresponding response result and a detection log.
[0024] In some embodiments, the system further includes a response execution layer, and the method further includes: Figure 2As shown, the response execution layer includes an adaptive response module, and the intelligent detection engine module of the detection execution layer can also deliver the detection result and the confidence to the adaptive response module of the response execution layer to perform operations, the adaptive response module executes a hierarchical response according to the detection result to obtain a response result and a detection log; the response result can include isolation success / failure, response time consumption, etc.
[0025] In the case of the detection result being an attack event, the corresponding response strategy can be to immediately switch the network connection, isolate the device, and send a red alert (such as an email or a short message); in the case of the detection result being a suspicious event, the corresponding response strategy can be to trigger real-time traffic monitoring, record logs, and send a yellow alert; in the case of the detection result being a normal event, the corresponding response strategy can be to only archive the logs and not trigger active responses.
[0026] In the above embodiments, a “detection-response” closed loop is formed, the detection result is converted into specific defense actions (such as interception, isolation, and warning), and differential responses are performed according to the detection result, which improves the real combat defense capability of the system; at the same time, the generation of the response result and the detection log provides detailed original data for subsequent system optimization, and also realizes the traceability of the attack disposal process, which is convenient for post-analysis and summary of experience.
[0027] In some embodiments, the system further includes a feedback optimization layer, and the method further includes: Step S170, the response execution layer generates response feedback data based on the response result and the detection log, and sends the response feedback data to the feedback optimization layer; Step S180, the feedback optimization layer iteratively trains the reinforcement learning model in the priority classification layer based on the response feedback data to dynamically adjust the generation strategy of the priority classification number.
[0028] As shown in Figure 2 As shown, the adaptive response module of the response execution layer collects the response result and the detection log to generate response feedback data, and the response feedback data includes priority misjudgment, excessively long detection time consumption, etc. The feedback optimization layer includes a knowledge base updating module and a model training module, the knowledge base updating module is used to analyze the detection log, extract new attack features (such as 0day vulnerability fingerprints), and update to the system feature library for subsequent detection matching; the model training module is used to retrain the reinforcement learning model in the reinforcement learning priority classification module by using the response feedback data (such as the records of the previous three similar attack failures) to adjust the generation strategy of the limited classification number, forming a “detection-response-optimization” closed loop.
[0029] As shown in Figure 3As shown, the data link includes quantum data acquisition → quantum hash feature extraction → dynamic priority classification → hierarchical detection → adaptive response, forming a closed-loop feedback. The control link includes the adaptive response module feeding the processing results (such as attack success / failure, response time consumption) to the reinforcement learning priority classification module to optimize the generation strategy of the priority classification number.
[0030] In the above embodiments, the response result and the detection log are fed back to the reinforcement learning priority classification module to optimize the priority generation strategy, forming a closed-loop system; the feedback optimization layer iteratively trains the reinforcement learning model of the priority classification layer using the response feedback data to dynamically adjust the generation strategy of the priority classification number. This design builds a full-link closed loop of “detection - response - feedback - optimization”, enabling the system to have the ability of self-learning and self-optimization. Compared with the defect that the traditional static classification mechanism cannot adjust according to the actual application effect, this layer can make the generation of the priority classification number more accurate and better adapt to the changing network attack scene, reducing the occurrence of misjudgment and omission.
[0031] In some embodiments, the above step S180 “the feedback optimization layer iteratively trains the reinforcement learning model in the priority classification layer based on the response feedback data to dynamically adjust the generation strategy of the priority classification number” includes the following steps: Step S801, the feedback optimization layer acquires the current state based on the state space definition of the reinforcement learning model in the priority classification layer, the current state including the target quantum feature vector, the event type, the historical response record and the system load; In the system, the reinforcement learning model is used to dynamically generate the priority classification number of network security events, and the core framework includes: state space (S), action space (A) and reward function R. The state space is constructed , wherein: H represents the 256-dimensional target quantum feature vector from the quantum hash encoding algorithm, which carries the unique quantum features of the network data block and is one of the important bases for judging the priority of the event. T represents the event type (normal / suspicious / attack), which is represented by one-hot encoding, which can accurately distinguish normal traffic, known attacks, new attacks and other categories; for example, a virus attack can be coded as 1000, and a vulnerability exploit can be coded as 0100, etc. This coding method can clearly distinguish different types of network security events. Y represents the historical response record recording the processing results of the last 10 times, including success, failure or delay, etc. These historical data can help the model learn the best decision strategy under different conditions. L represents a system load indicator, covering CPU utilization, memory occupation, network bandwidth and other information. By considering the system load, the model can allocate detection resources more reasonably, avoiding low detection efficiency caused by insufficient or excessive allocation of resources.
[0032] According to the characteristics of the reinforcement learning model, the system uses a deep reinforcement learning priority classification generation algorithm (DRL-PCG, Deep Reinforcement Learning for Priority Classification Generation) as the core reinforcement learning algorithm.
[0033] In step S802, the reinforcement learning model selects a priority classification number from a preset priority classification number set in the action space as an action based on the current state. The preset priority classification number set a in the action space can be represented as a . Wherein: The main classifications P1 / P2 / P3 correspond to high, medium and low priority levels, which is a preliminary division of the emergency level of network security events.
[0034] The sub-classifications C1-C5 further classify attack types in a fine-grained manner. For example, in the high priority level, P1-C1 may represent a certain type of high threat attack, P1-C2 represents another type, and so on. This hierarchical classification method can more accurately manage the priority of different types of security events.
[0035] In step S803, the reinforcement learning model updates the event type, the historical response record and the system load according to the response result generated by executing the action and the detection log, generates a new state containing the target quantum feature vector, the updated event type, the updated historical response record and the updated system load in combination with the target quantum feature vector, and constructs a reward function based on the response feedback data. The reward corresponding to the action is calculated through the reward function. The current state can be represented as s, the action can be represented as a, and the new state can be represented as The reward function can be represented as .
[0036] In step S804, the reinforcement learning model dynamically iteratively updates parameters based on the current state, the action, the reward and the new state to adjust the generation strategy of the priority classification number.
[0037] Through continuous interaction (i.e., the "detection-classification-feedback" cycle), the optimal strategy that can maximize the long-term cumulative reward (i.e., the decision logic for generating reasonable priority classification numbers) is learned.
[0038] Here, the reinforcement learning priority classification module can be based on a deep Q network (Double DQN) architecture, and the following mechanisms can be used to achieve efficient training and optimization of the model: Policy execution: the agent selects an action a based on the current state s according to the ε-greedy policy, and obtains a reward after execution and updates the system state to ; Experience replay: store training samples in the experience replay pool, and use TD-error (Temporal Difference-error) priority sampling to extract a mini-batch for parameter updating to improve training efficiency; (Parameter updating: use a deep neural network to fit the action-value function, minimize the MSE (Mean Squared Error) loss function, and use the Adam optimizer (Adaptive Moment Estimation) to iteratively update the network parameters to approximate the optimal policy.) Stable training: use the target network synchronization mechanism to periodically copy the main network parameters to the target network to avoid value function overestimation during training; Convergence optimization: reduce the exploration rate during the model convergence phase, fine-tune the weight coefficients, and achieve multi-objective balanced optimization to ensure that the model continues to adapt to system dynamics. By optimizing the classification strategy through experience replay (Replay Buffer) and target network synchronization mechanism, the priority identification accuracy of 0day attacks can be improved by 40% by automatically updating the priority rules every hour. In some embodiments, after 1000 detections, the evaluation network parameters can be synchronized to the target network to stabilize the training process; every 100 detections triggers experience replay, which can randomly sample data from the replay pool to optimize the model and avoid overfitting.
[0039] In the above embodiments, the iterative logic of the reinforcement learning model is quantified, and the closed-loop update of "current state-action-reward-new state" enables the model to accurately learn "which priority classification strategy is more effective", improving the dynamic adaptability of priority classification; combined with the response result and detection log update state (such as event type, history record, system load), the model training is more suitable for real-time scenarios, enhancing the pertinence and flexibility of the classification strategy.
[0040] In some embodiments, the step S803 of "constructing a reward function based on the response feedback data" includes the following steps: Step S8031, based on the response feedback data, determine the detection accuracy, response efficiency and resource utilization; The reward function can be designed based on a linearly weighted piecewise function, which can be determined by Balancing the three core indicators of detection accuracy, response efficiency, and resource utilization, represents the reward corresponding to the detection accuracy, represents the adjusted first reference weight, represents the reward corresponding to the response efficiency, represents the adjusted second reference weight, represents the reward corresponding to the resource utilization, represents the adjusted third reference weight. The design of the reward function aims to guide the model to learn the optimal priority classification strategy. By giving different rewards or penalties for different results, the model can gradually understand which action to take in which state to achieve the maximum cumulative reward, thereby continuously optimizing its decision-making process.
[0041] Step S8032, the reference weights of the detection accuracy, the response efficiency, and the resource utilization are the first reference weight, the second reference weight, and the third reference weight, respectively, when the system running state is in the normal load period; when the system running state is in the high load warning period, the first reference weight is reduced and the second reference weight is increased; when the system running state is in the new attack outbreak period, the first reference weight and the second reference weight are reduced, and the third reference weight is increased; The weights can be dynamically adjusted according to the system running state , , ( ), for example: in the normal load period, = 0.5, , (prioritize detection accuracy); in the high load warning period, = 0.3, 5, (improve high-priority response speed); in the new attack outbreak period, = 0.4, , (optimize computing power to respond to unknown threats).
[0042] The system load, new attack proportion, false positive rate, and other indicators can be counted every hour, and the current state can be automatically matched and the weight coefficient can be adjusted by a fuzzy logic controller.
[0043] Step S8033, based on the detection accuracy, the adjusted first reference weight, the response efficiency, the adjusted second reference weight, the resource utilization, and the adjusted third reference weight, a reward function is constructed by weighted summation.
[0044] In some embodiments, rewards are given for detection accuracy ( Positive incentives include a +100 reward for correctly identifying high-priority attacks (such as P1 ransomware) to encourage accurate detection of high-risk threats; and a +20 reward for accurately judging low-priority normal events (such as P3 regular traffic) to avoid oversensitivity leading to resource waste. Negative penalties include a -200 penalty for missing high-priority attacks (such as unidentified P1 attacks), which is much higher than the positive incentives, forcing the model to pay attention to high-risk threats; and a -50 penalty for misjudging normal events (such as P3 being marked as an attack) to avoid false alarms interfering with the system. Dynamic adjustments include doubling the reward for the first correct classification of a new type of attack (such as a +200 reward for the first identification of a 0-day vulnerability attack); and a 10% reduction in reward after three consecutive correct classifications of a known attack event (such as a reward of +90 for the fourth correct identification of a known virus, which is reduced from +100). In this embodiment, the correctness of high-priority events (such as P1 attacks) has the greatest impact, and therefore the reward and penalty are the strongest; the impact of low-priority events (such as normal P3 events) is smaller, and the reward and penalty are weaker.
[0045] In some embodiments, for response efficiency rewards ( The high-efficiency response rewards include a +50 reward for high-priority events with a response time <50ms, requiring extreme speed, such as instantaneous blocking of APT (Advanced Persistent Threat) attacks; a +30 reward for medium-priority events with a response time <100ms, balancing speed and resources; latency penalties include a -150 penalty for high-priority events with a response time >200ms, severely punishing delays and preventing the spread of high-risk attacks; latency caused by improper computing power allocation (such as insufficient GPU (Graphics Processing Unit) allocation to P1) is penalized with -0.5·t (t is the latency time); dynamic thresholds include a 50% relaxation of the response time threshold under high load (such as CPU utilization > 80%) (e.g., relaxing 50ms for P1 to 75ms), avoiding excessive pursuit of speed leading to system crashes; and a halving of related rewards when low-priority computing power usage exceeds the limit, preventing resource waste; in this embodiment, response speed is related to event priority (high priority requires faster response), and the threshold changes dynamically with system load (appropriately relaxed under high load).
[0046] In some embodiments, resource utilization rewards ( ), the reasonable allocation of rewards includes high-priority event GPU allocation ≥ 50% reward + 30 (ensure that the deep detection has sufficient computing power, such as QNN model running), low-priority event computing power consumption < 5% reward + 20 (encourage fast processing with lightweight algorithms); the low-efficiency penalty includes high-priority event computing power allocation < 30% penalty - 100 (force to guarantee the computing power of high-risk events to avoid missing detection due to insufficient resources), and the same priority computing power fluctuation exceeds 20% penalty - 40 (require stable resource allocation to avoid system shock); dynamic adjustment includes automatically adjusting the computing power threshold according to historical load data (such as increasing the GPU quota of P1 from 50% to 60% during peak hours), temporarily increasing the computing power quota of new attack corresponding sub-classification (such as C4 new vulnerability) by 10%, and increasing the reward coefficient to 1.2 times (such as changing + 30 to + 36), supporting in-depth analysis of unknown threats. In the embodiment of the application, the reward and punishment guiding model allocates resources on demand, allocates more to high-priority events, and occupies less to low-priority events, avoiding resource overload or waste.
[0047] In the above embodiment, the scenario adaptation of the reward mechanism is realized, the three are balanced under normal load, the response speed is preferentially guaranteed under high load, the resource utilization is preferentially optimized when a new attack breaks out, and resource overload is avoided, so that the optimization target of the reinforcement learning model is dynamically matched with the actual demand of the system, the problem of "inability to cope with complex scene changes" of the fixed weight reward function is solved, and the robustness of the system under different threat scenes is improved.
[0048] In some embodiments, the step S120 of "the data processing layer extracts features from the encrypted network data by a quantum hash coding algorithm to obtain a target quantum feature vector" includes the following steps: Step S201, the data processing layer extracts features from the encrypted network data by quantum state preparation, quantum entanglement transformation and probability measurement of the quantum hash coding algorithm to obtain an initial quantum feature vector containing phase verification information; The quantum hash preprocessing module of the data processing layer can execute the quantum hash coding algorithm on the encrypted data, generate an initial quantum feature vector containing phase verification information (such as a 256-dimensional binary string) through quantum state preparation, entanglement transformation and probability measurement.
[0049] Step S202, the data processing layer standardizes the initial quantum feature vector to obtain a target quantum feature vector.
[0050] The initial quantum feature vector can be standardized to unify the data format for subsequent analysis.
[0051] In the above embodiment, the initial quantum feature vector is obtained through quantum state preparation, quantum entanglement transformation and probability measurement of the quantum hash coding algorithm, so that the generated feature vector has stronger resistance to quantum attack, solves the problem that the traditional hash feature is easy to be cracked by quantum computing, and ensures the uniqueness and collision resistance of the feature vector. After the initial vector containing phase verification information is standardized, the feature format is unified, consistent input is provided for subsequent priority classification and detection model calling, and the detection stability is improved.
[0052] In some embodiments, the step S201 "the data processing layer performs feature extraction on the encrypted network data through quantum state preparation, quantum entanglement transformation and probability measurement of the quantum hash coding algorithm to obtain an initial quantum feature vector containing phase verification information" includes the following steps: Step S2011, the data processing layer converts the encrypted network data into a quantum bit sequence byte by byte, and applies a Hadamard gate to each quantum bit in the quantum bit sequence to convert the corresponding quantum bit from a ground state to a superposition state; Wherein, the quantum hash coding algorithm creates multiple branch hash values for each network data block by virtue of the superposition principle of quantum state. In the quantum world, each bit of the data block corresponds to multiple evolution paths of quantum states. Through quantum measurement means, a feature vector with a unique probability distribution can be obtained, which is used as the quantum hash feature of the data block.
[0053] The encrypted network data is in the form of a byte stream and can be represented as . Wherein, represents the i-th byte in the data block. The quantum hash coding algorithm outputs a quantum hash feature vector (i.e. the initial quantum feature vector) in the form of a binary string, which will be used as the key data identifier for subsequent processing.
[0054] Step S2011 is a quantum state preparation step. In this step, the input data block can be converted into a quantum bit sequence byte by byte. Each byte is accurately mapped to 8 quantum bits (each byte of eight binary numbers corresponds to eight quantum bits, binary 0 is mapped to quantum ground state , and 1 is mapped to ), and the initial state is set to ground state . For example, byte 0x41 (decimal 65) is represented in binary as 01000001, which is converted into 8 quantum bits with an initial state of .
[0055] A Hadamard gate is applied to each quantum bit to convert it from a ground state to a superposition state , and satisfies This step makes the quantum bit have the quantum superposition property through quantum gate operation, increases the dimension of information carrying, is a quantum state, representing the state of a quantum bit, and are the ground states of the quantum bit, corresponding to 0 and 1 of the classical bit, is the amplitude of the quantum bit in the state, is the amplitude of the quantum bit in the state, is the probability of the quantum bit in the state, is the probability of the quantum bit in the state.
[0056] Step S2012, the data processing layer introduces auxiliary quantum bits in the sequence of quantum bits, and constructs an entangled state through a CNOT gate; Wherein, step S2012 and step S2013 are quantum entanglement transformation steps, auxiliary quantum bits are introduced, and an entangled state is constructed through a CNOT gate (controlled NOT gate). For example, after introducing 1 auxiliary quantum bit, an entangled system of n+1 quantum bits is formed from n data quantum bits. This entangled state causes non-classical correlation between different quantum bits.
[0057] Step S2013, the data processing layer applies quantum Fourier transform to the entangled state to obtain a high-dimensional entangled state, so as to realize superposition state mapping of high-dimensional Hilbert space; The quantum Fourier transform (QFT, Quantum Fourier Transform) is applied to the entangled state to generate a superposition state in dimensional Hilbert space, which greatly expands the characteristic dimension, which provides a basis for generating a more unique hash value subsequently.
[0058] Step S2014, the data processing layer performs quantum measurement on the high-dimensional entangled state, obtains the probability distribution of each quantum bit collapsing into a classical bit, selects a preset number of quantum bit combinations with the highest probability from the probability distribution, and takes the coefficient ratio as additional phase verification information, to obtain an initial quantum feature vector containing phase verification information.
[0059] Wherein, step S2014 is a probability measurement step, and quantum measurement is performed on the high-dimensional entangled state to obtain the probability distribution of each quantum bit collapsing into a classical bit. The measurement process follows the probability interpretation of quantum mechanics, and different quantum bits collapse into or The probability is determined by the previous quantum state.
[0060] The probability distribution is selected from the probability of the highest m-bit combination (usually set to 256 bits), and a unique quantum hash value H is generated; at the same time, the phase information (such as the α / β ratio) is attached as an anti-collision check code, which can further enhance the uniqueness and anti-collision ability of the hash value, making the hash value more secure when facing complex attacks.
[0061] In the above embodiment, through the mapping of quantum bit superposition state and high-dimensional entangled state, the feature extraction dimension is higher, the information density is larger, and the subtle attack features that traditional hash cannot identify can be captured; the phase check information and probability screening are introduced, which further improves the uniqueness and tamper resistance of the features, and ensures the reliability of the target quantum feature vector, providing accurate "digital fingerprint" for subsequent detection. Through quantum hash coding technology, a hash value with quantum state superposition characteristics is generated, which can resist quantum computing attacks and data tampering, and improve the anti-attack performance of feature extraction. Through quantum state superposition and entanglement, the hash collision probability is reduced to , which is much higher than the of traditional algorithms.
[0062] In some embodiments, the step S140 of "the detection execution layer calls the corresponding detection model based on the priority classification number to detect the target quantum feature vector and obtains the detection result and confidence" includes the following steps: Step S401, in the case that the priority classification number represents high priority, using quantum neural network to mine deep features of the target quantum feature vector through quantum dot product operation, and obtaining the detection result and confidence; Step S402, in the case that the priority classification number represents medium priority, using a light convolutional neural network to perform suspicious traffic screening on the target quantum feature vector, and obtaining the detection result and confidence; Step S403, in the case that the priority classification number represents low priority, performing known attack feature matching on the target quantum feature vector through a rule matching engine, and obtaining the detection result and confidence.
[0063] Among them, the intelligent detection engine module of the detection execution layer can dynamically call the detection model according to the priority classification number. When the priority classification number is high priority (P1), the quantum neural network (QNN) is enabled to mine deep features through quantum dot product operation, and the detection speed can be improved by 3 times compared with traditional CNN; when the priority classification number is medium priority (P2), the light convolutional neural network (Light CNN) is used to quickly screen suspicious traffic; when the priority classification number is medium priority (P3), the rule engine is used to match known attack features, which can realize fast filtering.
[0064] In the above embodiment, a hierarchical detection mechanism is constructed: based on the priority classification number, quantum neural network deep detection is adopted for high-risk events, and lightweight algorithm is adopted for fast processing of low-risk events, balancing detection accuracy and efficiency. Different priority data are detected by quantum neural network, lightweight convolutional neural network or rule engine, realizing differentiated allocation of detection resources, combining the quantum computing advantage of QNN and the efficiency advantage of Light CNN / rule engine, while ensuring the detection accuracy of high-risk events, improving the detection speed and resource utilization of the overall system. Compared with the unified detection strategy, the detection speed of high-priority events is improved by 3 times, and the overall resource utilization of the system is optimized by 50%.
[0065] In some embodiments, the step S110 of "the data collection layer acquires network data of the data source layer and encrypts" includes the following steps: Step S101, the data collection layer acquires network data of the data source layer; Step S102, the data collection layer generates a shared key through quantum key distribution technology, encrypts the network data using the shared key, and obtains encrypted network data.
[0066] Among them, the data can be encrypted in real time through quantum key distribution (QKD, Quantum Key Distribution) technology, and the byte stream in the format of quantum state coding is obtained, which ensures that the transmission process is not tampered with or eavesdropped.
[0067] In the above embodiment, quantum key distribution is based on the principles of quantum mechanics, and the key cannot be cloned and eavesdropping will be detected, effectively solving the problem that the traditional key distribution method (such as RSA) is easy to be cracked by Shor algorithm. Encryption transmission is realized from the data collection source, and forms an "end-to-end" quantum-level security protection with subsequent quantum feature extraction, completely building a security defense line for data transmission.
[0068] Figure 3 A working flow diagram of a computer network security detection system is provided for the embodiments of the present application, as shown in Figure 4 The data collection stage, the system collects data from two dimensions of network traffic and device logs, captures network data packets through quantum sensors, and collects the running logs of servers, routers and other devices. All collected data is immediately encrypted using quantum keys to ensure the security of data transmission.
[0069] In the feature extraction stage, the encrypted data enters the quantum hash preprocessing module. Through quantum state preparation, entanglement transformation, and probability measurement, the original data is converted into a fixed-length quantum hash value, and a representative quantum feature vector is extracted, providing basic data for subsequent priority classification.
[0070] In the priority classification stage, the reinforcement learning priority classification module receives the quantum feature vector, combines historical response records and current system load information, and dynamically generates a priority classification number using a deep Q network model to classify security events into high (P1), medium (P2), and low (P3) priority levels.
[0071] In the hierarchical detection stage, different detection algorithms are used for security analysis based on the priority classification results. High-priority events are detected using a quantum neural network, medium-priority events are quickly screened using a lightweight convolutional neural network, and low-priority events are matched and detected using a rule engine, optimizing the allocation of detection resources.
[0072] In the response processing stage, the system executes corresponding response strategies based on the detection results. For confirmed attack behavior, the network connection is immediately cut off, the infected device is isolated, and a high-level alarm (red alert level three) is sent. For suspicious behavior, real-time traffic monitoring is performed and a medium-level alarm (yellow alert level two) is sent. For normal events, only event logs are recorded. The results of all response operations will be used as feedback information to update the priority strategy of the reinforcement learning model.
[0073] In the closed-loop optimization stage, the system feeds back the effects of each response and detection logs to the reinforcement learning module, continuously learns and adjusts the priority generation strategy, forms a complete closed-loop optimization mechanism, and continuously improves the system's detection and response capabilities to new security threats.
[0074] The embodiment of the application provides a quantum hash encoding process (taking a TCP packet as an example), which includes the following steps: Data block: a 1500-byte TCP packet is divided into 15 100-byte data blocks . Quantum state preparation: apply a Hadamard gate to 800 bits (100 bytes x 8) of each data block to generate 800 superposition state quantum bits . Entanglement transformation: introduce an auxiliary quantum bit, entangle it with each data quantum bit through a CNOT gate, and form an entangled state of 801 quantum bits.
[0075] Measurement generation: 1024 measurements are made on the entangled state, the collapse probability of each bit is counted, and the first 256 high-probability bits are selected to generate the quantum hash value H = "10101101...0101", and the phase information a = 0.6, b = 0.8 is added.
[0076] The embodiment of the application provides a reinforcement learning priority generation process (for new ransomware attacks), which comprises the following steps: State input: quantum features H: including encrypted traffic features (such as quantum hash values of AES encryption fingerprints); event type T: one-hot encoding as a ransomware attack (0010); historical response: response delay (failure record) caused by priority misjudgment in the previous 3 similar attacks. System load: CPU utilization 60%, memory occupation 75%. Action selection: evaluate network Calculate the Q value of each priority, and select the action with the highest Q value (high priority, ransomware subclass 3), wherein s represents the current state, a represents the action, represent parameters of the network Q, such as weights and biases of the network Q.
[0077] Response feedback: the detection engine identifies the attack features through the quantum neural network, the response module immediately isolates the infected host, the response time is 80 ms, and the reward +100 is obtained.
[0078] Experience is stored in the replay pool, and the evaluation network parameter update is triggered. The embodiment of the application provides a hierarchical detection engine workflow, which comprises the following steps: Priority classification number analysis: receiving a P1-C3 classified data packet triggers a high-priority detection process. Resource allocation: scheduling 60% of the GPU computing power to the quantum neural network (QNN), and inputting a 256-dimensional quantum feature vector. QNN detection steps: quantum hidden layer: 128 quantum neurons, performing quantum dot product operation (U is a quantum weight matrix); classical output layer: outputting attack type probability through a Softmax function, and determining the attack when the ransomware probability is greater than 95%, represents an input quantum state, represents a transformed quantum state, that is, the input quantum state is transformed into a transformed quantum state by a quantum weight matrix U . Response execution: sending a blocking instruction to a firewall, and sending a level 3 red alert to an administrator, with an attack feature hash value for tracing.
[0079] In the embodiments of the present application, the quantum hash coding technology generates multi-branch hash values using the superposition principle of quantum states to improve the uniqueness and attack resistance of features. The dynamic priority classification mechanism optimizes the priority allocation strategy in real time through reinforcement learning to solve the lag of static rules. The hierarchical detection engine architecture dynamically calls QNN / lightweight CNN detection models according to the priority classification number to achieve efficient resource allocation.
[0080] The present application provides a cloud data center network security detection embodiment. The implementation scenario is a data center operated by a large cloud service provider, which carries the core business data of thousands of enterprises, has a complex network structure, and has huge traffic. As shown in Figure 5 The implementation process includes the following steps: Data collection; Quantum data collection modules are deployed on key nodes such as border routers, core switches, and virtualization platforms in the cloud data center. The module continuously collects network traffic data, device state data, virtual machine log data, server log data, and virtual machine running state data. Quantum hash preprocessing; The collected data is transmitted to the quantum hash preprocessing module after being encrypted by the quantum key. The module uses quantum hash coding algorithms to convert the data into quantum hash values and extract quantum feature vectors. For example, for log data generated by a virtual machine, it is converted into a unique quantum feature vector through quantum hash coding for subsequent analysis. Reinforcement learning priority classification; The reinforcement learning priority classification module receives the quantum feature vector, combines information such as the importance of the cloud data center's business, historical attack records, and current system load, and uses a deep Q network model to generate a priority classification number. If abnormal traffic of a virtual machine is detected and the virtual machine carries the core business of an important enterprise, the system may determine it as high priority (P1). Hierarchical detection; The intelligent detection engine module performs detection according to the priority classification number. For high-priority data, a quantum neural network is enabled for deep detection; for low-priority data, a lightweight convolutional neural network and a rule engine are used for detection, respectively. For example, after detecting abnormal traffic with high priority, the quantum neural network will perform deep mining on the features of the traffic to determine whether there is a malicious attack. Response processing; The adaptive response module performs corresponding operations according to the detection results. If it is confirmed that there is an attack, the infected virtual machine is immediately isolated, its network connection is cut off, and a high-level alarm (such as a three-level red alarm) is sent to the security operation team of the cloud service provider. At the same time, response feedback information is sent to the reinforcement learning priority classification module to optimize the priority judgment strategy in the future.
[0081] The application provides an Internet of Things smart home network security detection embodiment. The implementation scenario is a smart home network composed of various smart devices, including smart door locks, cameras, thermostats, smart speakers, etc. The number of devices is large and the communication protocols are diverse. As shown in Figure 6 , the implementation process includes the following steps: Data collection; Deploy a quantum data collection module at the smart home gateway to collect communication data between each smart device and the gateway, including device status information (such as door opening records, video streams, temperature data, etc.), control instruction data (such as voice instructions), etc. Quantum hash preprocessing; The collected data is processed by quantum hash coding to convert it into a quantum feature vector. For example, the door opening record data of a smart door lock is converted into a unique quantum feature vector after quantum hash coding. Reinforcement learning priority classification; The reinforcement learning priority classification module generates a priority classification number based on the type, usage frequency, and historical security records of the smart device, combined with the quantum feature vector. If the video stream of a smart camera is abnormal and the camera is in an important position for home security monitoring, the system will determine it as a high priority. Hierarchical detection; The smart detection engine module detects data of different priorities based on the priority classification number. High-priority data is detected using a quantum neural network, and low-priority data is detected using a lightweight convolutional neural network and a rule engine, respectively. For example, for high-priority smart camera video stream abnormal data, the quantum neural network analyzes whether the video content contains illegal intrusion or other abnormal situations. Response processing; The adaptive response module executes response operations based on the detection results. If a smart door lock is detected to be illegally attempted to be unlocked, an alarm message is immediately sent to the user's mobile phone, and the door lock is automatically locked; if an abnormal instruction is detected by a smart speaker, its work is suspended and relevant information is recorded. At the same time, the response feedback is sent to the reinforcement learning priority classification module to optimize the subsequent detection strategy.
[0082] The application provides a mobile operator network security detection embodiment. The implementation scenario is the core network of a mobile operator, connecting a large number of base stations, user terminals, processing massive voice and data traffic, and facing various network attack threats. As shown in Figure 7 , the implementation process includes the following steps: Data collection; Quantum data collection modules are deployed at positions such as core network elements and base station controllers of mobile operators to collect signaling data, user traffic data and the like in the network. Quantum hash preprocessing; The collected data is quantum hash coded to extract a quantum feature vector. For example, for user online traffic data, a specific quantum feature vector is formed through quantum hash coding. Reinforcement learning priority classification; The reinforcement learning priority classification module analyzes the quantum feature vector in combination with information such as the service type (such as ordinary data service and high-definition video service) of the user, the user level (such as VIP user and ordinary user), and historical attack data to generate a priority classification number. If the online traffic of a VIP user abnormally increases, the system may determine it as high priority. Hierarchical detection; The intelligent detection engine module performs hierarchical detection on the data according to the priority classification number. High-priority data is detected in depth by a quantum neural network, and low-priority data is detected by a lightweight convolutional neural network and a rule engine, respectively. For example, for high-priority abnormal traffic data of a VIP user, the quantum neural network deeply analyzes the source, destination, and data content of the traffic to determine whether there is an attack behavior.
[0083] Response processing; The adaptive response module executes a response operation according to the detection result. If a malicious attack against a VIP user is detected, measures such as traffic cleaning and switching of network path are taken to ensure normal operation of the user service, and an alarm is sent to the security and operation personnel of the operator. At the same time, response feedback information is sent to the reinforcement learning priority classification module to adjust the priority judgment strategy.
[0084] The embodiments of the present application provide comprehensive and accurate security detection: combining network device inspection and analysis, network traffic tracking analysis, and user network behavior detection and evaluation, the network security situation can be comprehensively analyzed from multiple dimensions. By comprehensively judging various factors, the network security risk is accurately evaluated, avoiding the limitations of single detection methods, and greatly improving the accuracy and reliability of detection. For example, in the above embodiments of enterprises and colleges and universities, network device vulnerabilities, abnormal traffic, and user violations can be accurately found, and potential security threats can be identified in a timely manner. The embodiments of the present application provide timely and effective early warning response: the system can monitor the network state in real time, and once a security problem is found, it can quickly generate an alarm and notify the relevant personnel. This timely early warning mechanism enables network administrators to take measures to respond to security threats in the first time, effectively reducing the loss caused by security incidents. For example, in the enterprise case, when signs of DDoS (Distributed Denial of Service) attacks and abnormal employee behavior are detected, the system timely issues an early warning, enabling administrators to quickly block attacks and regulate employee behavior, avoiding serious consequences. The embodiments of the present application can improve network security and stability: by regularly scanning for vulnerabilities and checking the configuration of network devices, potential security vulnerabilities can be found and repaired in a timely manner, and network device configuration can be optimized, reducing the risk of network device attacks and improving network security. At the same time, real-time monitoring of network traffic and timely processing of abnormal traffic, as well as regulation of user network behavior, ensure the stable operation of the network and avoid network congestion, interruption and other problems caused by abnormal traffic or malicious behavior. In the case of a university campus network, by solving network device problems and regulating student network behavior, the security and stability of the campus network are effectively improved. The embodiments of the present application can reduce the cost of network security management: compared with traditional network security management methods, the system realizes automatic detection and evaluation, reducing the workload of manual inspection and analysis and reducing labor costs. Moreover, since security problems can be found and solved in a timely manner, high losses caused by business interruptions, data loss and other problems are avoided, and the overall cost of network security management is reduced.
[0085] The embodiments of the present application can promote network use compliance: monitoring and evaluating user network behavior helps to ensure that users comply with network use regulations and prevent security risks caused by user violations. Whether it is an employee or a student, under the supervision of the system, they can use the network more normatively, creating a safe and compliant network environment.
[0086] Based on the foregoing embodiments, the embodiments of the present application further provide a computer network security detection system, which comprises modules and units included in the modules, and can be implemented by a processor in an electronic device. Of course, the system can also be implemented by a specific logic circuit. In the implementation process, the processor can be a central processing unit (CPU), a micro processing unit (MPU), a digital signal processor (DSP), or a field programmable gate array (FPGA).
[0087] Figure 7 A schematic diagram of the composition structure of a computer network security detection system provided by the embodiments of the present application is shown in FIG. 7. As shown in FIG. 7, the system 700 comprises a data source layer 710, a data collection layer 720, a data processing layer 730, a priority classification layer 740, and a detection execution layer 750, wherein: The data collection layer 720 is configured to acquire network data of the data source layer 710 and encrypt the network data, and send the encrypted network data to the data processing layer 730. The data processing layer 730 is configured to perform feature extraction on the encrypted network data by using a quantum hash coding algorithm, to obtain a target quantum feature vector, and send the target quantum feature vector to the priority classification layer 740. The priority classification layer 740 is configured to generate a priority classification number based on the target quantum feature vector, an event type, a historical response record, and a system load by using a reinforcement learning model, and send the priority classification number and the target quantum feature vector to the detection execution layer 750. The priority classification number comprises a priority level and a sub-classification number. The detection execution layer 750 is configured to perform detection on the target quantum feature vector based on the priority classification number by using a corresponding detection model, to obtain a detection result and a confidence level, and use the detection result to represent a security attribute of the network data.
[0088] In some possible embodiments, the system further comprises a response execution layer, The detection execution layer 750 is further configured to send the detection result and the confidence level to the response execution layer. The response execution layer is configured to execute a corresponding response strategy based on the detection result and the confidence level, to obtain a corresponding response result and a detection log.
[0089] In some possible embodiments, the system further comprises a feedback optimization layer, The response execution layer is further configured to generate response feedback data based on the response result and the detection log, and send the response feedback data to the feedback optimization layer; and the feedback optimization layer is configured to perform iterative training on the reinforcement learning model in the priority classification layer based on the response feedback data, so as to dynamically adjust the generation strategy of the priority classification number.
[0090] In some possible embodiments, the feedback optimization layer is further configured to obtain a current state based on a state space definition of the reinforcement learning model in the priority classification layer 740, the current state including the target quantum feature vector, the event type, the historical response record and the system load; the reinforcement learning model is configured to select a priority classification number from a preset priority classification number set in an action space as an action based on the current state; the reinforcement learning model is further configured to update the event type, the historical response record and the system load according to a response result generated by executing the action and a detection log, generate a new state including the target quantum feature vector, the updated event type, the updated historical response record and the updated system load in combination with the target quantum feature vector, construct a reward function based on the response feedback data, and calculate a reward corresponding to the action through the reward function; and the reinforcement learning model is further configured to dynamically and iteratively update parameters based on the current state, the action, the reward and the new state, so as to adjust the generation strategy of the priority classification number.
[0091] In some possible embodiments, the reinforcement learning model is further configured to determine detection accuracy, response efficiency and resource utilization based on the response feedback data; and benchmark weights of the detection accuracy, the response efficiency and the resource utilization are respectively a first benchmark weight, a second benchmark weight and a third benchmark weight when a system running state is a normal load period; the first benchmark weight is reduced and the second benchmark weight is increased when the system running state is a high load early warning period; the first benchmark weight and the second benchmark weight are reduced and the third benchmark weight is increased when the system running state is a new attack outbreak period; and a reward function is constructed through weighted summation based on the detection accuracy, the adjusted first benchmark weight, the response efficiency, the adjusted second benchmark weight, the resource utilization and the adjusted third benchmark weight.
[0092] In some possible embodiments, the data processing layer 730 is further configured to perform feature extraction on the encrypted network data through quantum state preparation, quantum entanglement transformation and probability measurement of a quantum hash coding algorithm, to obtain an initial quantum feature vector including phase verification information; and the data processing layer 730 is further configured to perform standardization processing on the initial quantum feature vector, to obtain a target quantum feature vector.
[0093] In some possible embodiments, the data processing layer 730 is further configured to convert the encrypted network data into a sequence of qubits byte by byte, apply a Hadamard gate to each qubit in the sequence of qubits to convert the corresponding qubit from a ground state to a superposition state, introduce an ancillary qubit into the sequence of qubits, and construct an entangled state through a CNOT gate; the data processing layer 730 is further configured to apply a quantum Fourier transform to the entangled state to obtain a high-dimensional entangled state, so as to realize superposition state mapping of a high-dimensional Hilbert space; the data processing layer 730 is further configured to perform quantum measurement on the high-dimensional entangled state, obtain a probability distribution of each qubit collapsing into a classical bit, filter a preset number of qubit combinations with the highest probability from the probability distribution, and take a coefficient ratio as additional phase verification information, to obtain an initial quantum feature vector containing phase verification information.
[0094] In some possible embodiments, the data acquisition layer 720 is configured to acquire network data of the data source layer 210, and generate a shared key through a quantum key distribution technology, and encrypt the network data by using the shared key to obtain encrypted network data.
[0095] It should be noted that the above description of the system embodiments is similar to the description of the above method embodiments, and has similar beneficial effects to the method embodiments. For technical details not disclosed in the system embodiments of the present application, please refer to the description of the method embodiments of the present application for understanding.
[0096] It should be understood that, throughout the specification, "one embodiment" or "an embodiment" means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the application. Therefore, "in one embodiment" or "in an embodiment" appearing in various places throughout the specification does not necessarily refer to the same embodiment. In addition, these particular features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. It should be understood that, in various embodiments of the present application, the size of the serial number of each process does not mean the execution order, and the execution order of each process should be determined according to its function and inherent logic, and should not constitute any limitation on the implementation process of the embodiments of the present application. The serial number of the above embodiments of the present application is only for description, not representing the advantages and disadvantages of the embodiments.
[0097] It should be noted that, in the present document, the terms "comprises", "comprising", or any other variations thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can also include other elements not expressly listed or inherent to such process, method, article, or apparatus. Without further limitation, an element preceded by "comprises... a" does not, without more constraints, foreclose the existence of additional identical elements in the process, method, article, or apparatus that comprises the element.
[0098] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other manners. The above described device embodiments are merely schematic. For example, the division of the units is only a logical function division. There can be another division manner for the actual implementation, for example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed coupling, or direct coupling or communication connection between the components can be indirect coupling or communication connection through some interfaces, devices, or units, and can be electrical, mechanical, or in other forms.
[0099] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units; they can be located in one place, or distributed on multiple network units; and some or all of the units can be selected according to actual needs to achieve the purpose of the embodiments of the present application. In addition, each functional unit in the embodiments of the present application can be integrated into one processing unit, or each unit can be a separate unit, or two or more units can be integrated into one unit; the integrated unit can be realized in the form of hardware, or in the form of hardware plus software functional units.
[0100] Alternatively, the integrated units of the present application, if implemented in the form of software functional modules and sold or used as independent products, can also be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the embodiments of the present application can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling an apparatus to perform all or part of the methods described in the embodiments of the present application. The aforementioned storage medium includes: mobile storage devices, ROM, magnetic disks or optical disks, and various other media that can store program codes.
[0101] The methods disclosed in the several method embodiments provided by the present application can be combined arbitrarily without conflict to obtain new method embodiments. The features disclosed in the several method or device embodiments provided by the present application can be combined arbitrarily without conflict to obtain new method embodiments or device embodiments.
[0102] The above merely provides a method of the present application, but the protection scope of the present application is not limited thereto, any person skilled in the art can easily think of changes or replacements within the technical scope disclosed by the present application, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A computer network security detection method, characterized by, The application is applied to a computer network security detection system, and the system comprises a data source layer, a data collection layer, a data processing layer, a priority classification layer and a detection execution layer, and the method comprises the following steps: The data collection layer acquires network data of the data source layer and encrypts the network data, and sends the encrypted network data to the data processing layer; the data processing layer extracts features from the encrypted network data by using a quantum hash coding algorithm to obtain a target quantum feature vector, and sends the target quantum feature vector to the priority classification layer; The priority classification layer generates a priority classification number based on the target quantum feature vector, an event type, historical response records and system load by using a reinforcement learning model, and sends the priority classification number and the target quantum feature vector to the detection execution layer, wherein the priority classification number comprises a priority level and a sub-classification number; The detection execution layer calls a corresponding detection model based on the priority classification number to detect the target quantum feature vector, and obtains a detection result and a confidence level, wherein the detection result is used to represent a security attribute of the network data.
2. The method of claim 1, wherein, The system further comprises a response execution layer, and the method further comprises the following steps: The detection execution layer sends the detection result and the confidence level to the response execution layer; The response execution layer executes a corresponding response strategy based on the detection result and the confidence level to obtain a corresponding response result and a detection log.
3. The method of claim 2, wherein, The system further comprises a feedback optimization layer, and the method further comprises the following steps: The response execution layer generates response feedback data based on the response result and the detection log, and sends the response feedback data to the feedback optimization layer; The feedback optimization layer iteratively trains the reinforcement learning model in the priority classification layer based on the response feedback data to dynamically adjust a generation strategy of the priority classification number.
4. The method of claim 3, wherein, The feedback optimization layer iteratively trains the reinforcement learning model in the priority classification layer based on the response feedback data to dynamically adjust a generation strategy of the priority classification number, which comprises the following steps: The feedback optimization layer acquires a current state based on a state space definition of the reinforcement learning model in the priority classification layer, wherein the current state comprises the target quantum feature vector, the event type, the historical response records and the system load; the reinforcement learning model selects a priority classification number from a preset priority classification number set in an action space as an action based on the current state; The reinforcement learning model updates the event type, the historical response records and the system load according to a response result and a detection log generated by executing the action, generates a new state comprising the target quantum feature vector, updated event type, updated historical response records and updated system load in combination with the target quantum feature vector, constructs a reward function based on the response feedback data, and calculates a reward corresponding to the action by using the reward function; The reinforcement learning model dynamically and iteratively updates parameters based on the current state, the action, the reward and the new state to adjust the generation strategy of the priority classification number.
5. The method of claim 4, wherein, The constructing a reward function based on the response feedback data comprises: Based on the response feedback data, the detection accuracy, response efficiency and resource utilization are determined; The baseline weights of the detection accuracy, the response efficiency and the resource utilization are respectively a first baseline weight, a second baseline weight and a third baseline weight when the system running state is a normal load period; the first baseline weight is reduced and the second baseline weight is increased when the system running state is a high load early warning period; the first baseline weight and the second baseline weight are reduced and the third baseline weight is increased when the system running state is a new attack outbreak period; Based on the detection accuracy, the adjusted first baseline weight, the response efficiency, the adjusted second baseline weight, the resource utilization and the adjusted third baseline weight, a reward function is constructed by weighted summation.
6. The method of claim 1, wherein, The data processing layer extracts features from the encrypted network data through a quantum hash coding algorithm to obtain a target quantum feature vector, comprising: The data processing layer extracts features from the encrypted network data through quantum state preparation, quantum entanglement transformation and probability measurement of the quantum hash coding algorithm to obtain an initial quantum feature vector containing phase verification information; The data processing layer normalizes the initial quantum feature vector to obtain a target quantum feature vector.
7. The method of claim 6, wherein, The data processing layer extracts features from the encrypted network data through quantum state preparation, quantum entanglement transformation and probability measurement of the quantum hash coding algorithm to obtain an initial quantum feature vector containing phase verification information, comprising: The data processing layer converts the encrypted network data byte by byte into a quantum bit sequence, and applies a Hadamard gate to each quantum bit in the quantum bit sequence to convert the corresponding quantum bit from a ground state to a superposition state; The data processing layer introduces auxiliary quantum bits in the quantum bit sequence and constructs an entangled state through a CNOT gate; The data processing layer applies quantum Fourier transform to the entangled state to obtain a high-dimensional entangled state to realize superposition state mapping of a high-dimensional Hilbert space; The data processing layer performs quantum measurement on the high-dimensional entangled state to obtain the probability distribution of each quantum bit collapsing into a classical bit, selects a preset number of quantum bit combinations with the highest probability from the probability distribution, and takes the coefficient ratio as additional phase verification information to obtain an initial quantum feature vector containing phase verification information.
8. The method of claim 1, wherein, The detection execution layer calls a corresponding detection model based on the priority classification number to detect the target quantum feature vector to obtain a detection result and a confidence level, comprising: In the case that the priority classification number represents high priority, a quantum neural network is used to mine deep features of the target quantum feature vector through quantum dot product operation to obtain a detection result and a confidence level; In the case that the priority classification number represents medium priority, a lightweight convolutional neural network is used to perform suspicious traffic screening on the target quantum feature vector to obtain a detection result and a confidence level; In the case that the priority classification number represents a low priority, known attack feature matching on the target quantum feature vector is performed by the rule matching engine to obtain a detection result and a confidence level.
9. The method of claim 1, wherein, The data acquisition layer acquires and encrypts the network data of the data source layer, including: The data acquisition layer acquires the network data of the data source layer; The data acquisition layer generates a shared key through quantum key distribution technology, encrypts the network data using the shared key, and obtains encrypted network data.
10. A computer network security detection system, characterized by, Including: A data source layer, a data acquisition layer, a data processing layer, a priority classification layer, and a detection execution layer, wherein: The data acquisition layer is configured to acquire and encrypt the network data of the data source layer, and send the encrypted network data to the data processing layer; The data processing layer is configured to perform feature extraction on the encrypted network data through a quantum hash coding algorithm to obtain a target quantum feature vector, and send the target quantum feature vector to the priority classification layer; The priority classification layer is configured to generate a priority classification number based on the target quantum feature vector, an event type, a historical response record, and a system load through a reinforcement learning model, and send the priority classification number and the target quantum feature vector to the detection execution layer, wherein the priority classification number includes a priority level and a sub-classification number; The detection execution layer is configured to call a corresponding detection model based on the priority classification number to detect the target quantum feature vector, and obtain a detection result and a confidence level, wherein the detection result is used to represent the security attribute of the network data.