Intelligent Vulnerability Scanning Policy Generation Method and Device Based on Reinforcement Learning

Through the intelligent vulnerability scanning strategy generation method based on reinforcement learning, the scanning plug-in selection strategy is dynamically generated, which solves the problem that manually configuring scanning strategies in the existing technology is difficult to cope with complex network environments, and achieves efficient and accurate network security scanning.

CN119892505BActive Publication Date: 2025-06-20HANGZHOU HIKVISION DIGITAL TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510361591.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-06-20
Estimated Expiration
2045-03-25

AI Technical Summary

Technical Problem

In the existing network security defense system, manually configured scanning strategies are difficult to cope with the increasingly complex network attack and defense confrontation environment, resulting in insufficient scanning efficiency and accuracy.

Method used

The intelligent vulnerability scanning strategy generation method based on reinforcement learning is adopted, and the use priority of scanning plug-ins is optimized by determining the device portrait of the target device, generating device feature vectors, dynamically generating scanning plug-in selection strategies using historical scanning behaviors and reinforcement learning models.

Benefits of technology

The adaptive dynamic generation scanning plug-in strategy is implemented, which improves the dynamic adaptability of scanning, reduces false alarms and missed alarms, and improves detection efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119892505B_ABST
    Figure CN119892505B_ABST
Patent Text Reader

Abstract

The present application provides a method and apparatus for generating intelligent vulnerability scanning strategies based on reinforcement learning. In this embodiment, based on the device feature vector of the target device perceived in real time and the historical scanning results of the target device through reinforcement learning, a current scanning plug-in selection strategy is dynamically generated to ensure that the dynamically generated scanning plug-in selection strategy has strong dynamic adaptability, thereby reducing false positives and false negatives. Further, with the help of the reward mechanism of reinforcement learning, according to the historical performance of different scanning plug-ins and the current scanning results of the target device (and even plus user feedback information), the Actor model in the reinforcement learning model is optimized and adjusted in real time to continuously optimize the subsequent scanning plug-in selection strategy and improve the detection efficiency and accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to network security, and particularly to an intelligent vulnerability scanning policy generation method and device based on reinforcement learning. Background Art

[0002] In the network security defense system, scanning devices is often the core means to actively discover system vulnerabilities. However, the scanning highly depends on the rationality of the scanning policy. At present, the scanning policy is often configured manually. And the manually configured scanning policy is difficult to cope with the increasingly complex network attack and defense environment. Summary of the Invention

[0003] This application provides an intelligent vulnerability scanning policy generation method and device based on reinforcement learning to realize intelligent generation of scanning policies based on reinforcement learning.

[0004] An embodiment of this application provides an intelligent vulnerability scanning policy generation method based on reinforcement learning. The method is applied to a scanning device and includes:

[0005] Determine the encoding method matched by each feature data in the device profile of the target device; the device profile is generated based on the device information of the target device obtained by actively scanning the target device; the device profile contains feature data corresponding to multiple different feature types, and the encoding methods matched by the feature data of at least two different feature types are different;

[0006] Encode each feature data in the device profile according to the encoding method matched by each feature data in the device profile, and generate a device feature vector of the target device according to the encoding results of each feature data;

[0007] Obtain a reference scanning result based on the historical scanning behavior of the target device that has been executed; the reference scanning result at least includes: the scanning situation of each scanning plugin when scanning the target device, and the scanning situation at least includes: missed reports of abnormal information, false reports, unrecognized abnormal information, correctly recognized abnormal information;

[0008] Use the policy model Actor model in the current reinforcement learning model to evaluate the potential benefits of each scanning plugin based on the device feature vector and the reference scanning result to obtain a scanning plugin selection policy; and for each scanning plugin that needs to be used when currently scanning the target device indicated by the scanning plugin selection policy, determine the scanning priority of the scanning plugin when currently scanning the target device according to the resources occupied by the scanning plugin during execution, the risk situation of the abnormal information that can be scanned out by the scanning plugin, and the most recent time when the scanning plugin was used;

[0009] Schedule each scanning plugin to scan the target device according to the scanning priorities of the scanning plugins and in accordance with the priority scheduling algorithm, to obtain the current scanning results of each scanning plugin; determine the reward values of each scanning plugin based on the current scanning results of each scanning plugin and the reward function of each scanning plugin, and use the value Critic model in the current reinforcement learning model to evaluate the behavior performance of the scanning plugin selection strategy based on the reward values of each scanning plugin, so as to optimize the Actor model according to the evaluation results of the Critic model; the optimized Actor model outputs a better scanning plugin selection strategy compared to before optimization.

[0010] An embodiment of the present application further provides a scanning device, and the device includes:

[0011] A data perception module, configured to determine the encoding methods matched by the feature data in the device profile of the target device; the device profile is generated based on the device information of the target device obtained by actively scanning the target device; the device profile contains feature data corresponding to multiple different feature types, and the encoding methods matched by the feature data of at least two different feature types are different; and, configured to encode the feature data in the device profile according to the encoding methods matched by the feature data in the device profile, and generate a device feature vector of the target device according to the encoding results of the feature data.

[0012] A policy generation module, configured to obtain a reference scanning result based on the historical scanning behaviors performed on the target device; the reference scanning result at least includes: the scanning situations of each scanning plugin when scanning the target device, and the scanning situations at least include: missed reports of abnormal information, false reports, unrecognized abnormal information, correctly recognized abnormal information; and, use the policy model Actor model in the current reinforcement learning model to evaluate the potential benefits of each scanning plugin based on the device feature vector and the reference scanning result to obtain a scanning plugin selection strategy.

[0013] A policy execution module, configured to, for each scanning plugin indicated by the scanning plugin selection strategy and required to be used when currently scanning the target device, determine the scanning priority of the scanning plugin when currently scanning the target device according to the resources occupied by the scanning plugin during execution, the risk situation of the abnormal information that can be scanned out by the scanning plugin, and the most recent time when the scanning plugin is used; and

[0014] Schedule each scanning plugin to scan the target device according to the scanning priorities of the scanning plugins and in accordance with the priority scheduling algorithm, to obtain the current scanning results of each scanning plugin.

[0015] A policy feedback module, configured to determine the reward values of each scanning plugin based on the current scanning results of each scanning plugin and the reward function of each scanning plugin, and feedback the reward values of each scanning plugin and the current scanning results of each scanning plugin to the policy generation module;

[0016] The policy generation module is further configured to, according to the feedback from the policy feedback module and using the value Critic model in the current reinforcement learning model, evaluate the behavior performance of the scanning plugin selection policy based on the reward values of each scanning plugin, so as to optimize the Actor model according to the evaluation results of the Critic model; compared with the Actor model before optimization, the optimized Actor model outputs a better scanning plugin selection policy.

[0017] As can be seen from the above technical solutions, in this application, in this embodiment, through reinforcement learning, based on the device feature vector of the target device perceived in real time and the historical scanning results of the target device, a current scanning plugin selection policy is dynamically generated to ensure that the dynamically generated scanning plugin selection policy has strong dynamic adaptability, thereby reducing false positives and false negatives.

[0018] Furthermore, in this embodiment, by means of the reward mechanism of reinforcement learning, according to the historical performance of different scanning plugins and the current scanning results of the target device (and even plus user feedback information), the Actor model in the reinforcement learning model is optimized and adjusted in real time to continuously optimize the subsequent scanning plugin selection policy and improve the detection efficiency and accuracy.

[0019] Furthermore, in this embodiment, the target device is profiled to describe the target device with as comprehensive information as possible, and according to the coding method matched by each feature data in the device profile, each feature data in the device profile is encoded to generate the device feature vector of the target device, which ensures the planning and management of data, avoids adding useless data and other noises such as forcing each feature data to be encoded into an equal-length vector in a unified coding format by a conventional convolutional neural network, and can further ensure that when generating a scanning policy for the target device based on the device feature vector, the accuracy of the policy is improved. Brief Description of the Drawings

[0020] The accompanying drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure.

[0021] Figure 1 It is a flowchart of the method provided by the embodiment of the present application;

[0022] Figure 2 It is a schematic diagram of the device feature vector provided by the embodiment of the present application;

[0023] Figure 3 Provided by the embodiment of the present applicationFigure 1 Schematic diagram of the shown process;

[0024] Figure 4 Flowchart for implementing step 104 provided by an embodiment of the present application;

[0025] Figure 5 Device structure diagram provided by an embodiment of the present application;

[0026] Figure 6 Electronic device structure diagram provided by an embodiment of the present application. Specific implementation manners

[0027] In order to enable those skilled in the art to better understand the technical solutions provided by the embodiments of the present application and make the above-mentioned objects, features, and advantages of the embodiments of the present application more obvious and understandable, the technical solutions in the embodiments of the present application will be further described in detail below with reference to the accompanying drawings.

[0028] The method provided by the embodiment of the present application can break through the dependence on manual experience and realize the adaptive dynamic generation of scanning plug-in strategies. Moreover, the embodiment of the present application will also construct a closed-loop feedback system based on the dynamically generated scanning plug-in strategies to optimize the subsequently generated scanning plug-in strategies (indicating which scanning plug-ins can be used to scan the device to be scanned), reduce the frequent occurrence of false alarms and the situation of multiple unrecognized abnormal information, and improve the detection efficiency.

[0029] The method provided by the embodiment of the present application will be described below:

[0030] Refer to Figure 1 , Figure 1 which is the flowchart of the method provided by the embodiment of the present application. This process can be applied to a scanning device. The present embodiment does not specifically limit the deployment location of the scanning device in the network. As an embodiment, the scanning device can be deployed in existing network devices in the network such as switches, gateways, etc., or can be independent of the existing network devices in the network, and the present embodiment does not specifically limit.

[0031] When it is necessary to perform a scan on a target device, the scanning device can be enabled to perform a scanning task on the target device. For details, refer to Figure 1 the shown process.

[0032] As shown in Figure 1 , this process may include:

[0033] Step 101, determining the encoding method that matches each feature data in the device profile of the target device; the device profile is generated based on the device information of the target device obtained by actively scanning the target device.

[0034] In this embodiment, device information of a target device is collected through active scanning to generate a device portrait of the target device. The device portrait of the target device is used to provide a description of the target device that is as complete as possible.

[0035] Specifically, the device portrait may include feature data corresponding to multiple different feature types. Here, the multiple different feature types include, for example, port type, service version type, device description type, operating system type, etc.

[0036] For the port type, it is used to indicate the service opening status of ports on the target device. That is, the feature data under the port type at least includes: the service opening status of at least one port. For example, the service opening status of any port may include: whether the port is open, whether the port is a high-risk port that meets the high-risk definition such as port 80, port 3389, etc., and the services supported by the port such as TCP service, UDP service, etc.

[0037] For the service version type, it is used to indicate the service version information of the target device. That is, the feature data under the service version type at least includes: the service version information of the target device.

[0038] For the device description type, it is used to indicate the description information of the target device. That is, the feature data under the device description type at least includes: the description information of the target device such as device manufacturer, device type, device model, device firmware version information, etc.

[0039] For the operating system type, it is used to indicate the operating system of the target device. That is, the feature data under it at least includes: the operating system of the target device such as Windows, Linux, etc.

[0040] Those skilled in the art know that for a convolutional neural network, the input data is required to be a regular vector matrix. In this embodiment, most of the feature data are discrete categorical variables and of variable length. According to the requirements of the convolutional neural network, it is necessary to force each feature data to be encoded into an equal-length vector in a unified encoding format, which will lead to the introduction of useless data and other noises. In this embodiment, the different characteristics of each feature data are retained to match the encoding method for each feature data according to the features of each feature data. This makes it such that in this embodiment, the encoding methods matched by the feature data of at least two different feature types are different. For example, the encoding method matched by the feature data under the service version type, the encoding method matched by the feature data under the device description type, and the encoding method matched by the feature data under the operating system type are different from the encoding method matched by the feature data under the port type.

[0041] Specifically, for each feature type, the encoding method for the feature data under that feature type can be determined based on the longest length of the feature data under that feature type. For example, for the feature data under the port type, which is generally relatively long, the high-dimensional sparse matrix encoding method can be first used to encode the service opening status of each port under the port type. Among them, when encoding, if there is a high-risk port that meets the high-risk definition, a corresponding weight coefficient is added to the service opening status of that high-risk port; then, dimensionality reduction is performed on the high-dimensional sparse matrix encoding.

[0042] For the encoding method for the feature data matching under the service version type, the bert encoding method can be used for encoding; for the encoding method for the feature data matching under the device description type, the transE algorithm can be used to generate word embedding vectors; for the encoding method for the feature data matching under the operating system type, one-dimensional vectors can be used to encode the feature data under the operating system type.

[0043] Step 102: According to the encoding methods for the respective feature data in the device profile, encode the respective feature data in the device profile, and generate a device feature vector of the target device based on the encoding results of the respective feature data.

[0044] After encoding the respective feature data in the device profile according to the encoding methods for the respective feature data in the device profile, the encoded data can be concatenated, and finally a device feature vector of the target device is obtained. Figure 2 The specific processes of steps 101 to 102 are illustrated by way of example.

[0045] Step 103: Obtain a reference scan result based on the historical scan behaviors performed on the target device; the reference scan result at least includes: the scan situations when each scan plugin scans the target device, and the scan situations at least include: missed reports of abnormal information, false alarms, unrecognized abnormal information, correctly recognized abnormal information.

[0046] It should be noted that in this embodiment, the historical scan results of the target device are not blindly used as the reference scan result. Instead, by means of the historical scan results, an information is obtained: pay attention to the most recent scan result of the target device, but also refer to other non-recent scan results of the target device. This information is recorded as the reference scan result.

[0047] As for how to obtain the reference scan result, in this embodiment, it is obtained by means of the prediction mechanism of the LSTM model that combines long-term preference and short-term preference. Specifically, the historical scan results of the target device are input into the LSTM model, so that the LSTM model analyzes the historical scan results of the target device and generates the reference scan result; the reference scan result focuses on the most recent scan result of the target device and also refers to other non-recent scan results of the target device.

[0048] The reference scan results here at least include: the scan situations when each scan plugin scans the target device, and the scan situations at least include: missed reports of abnormal information, false reports, unrecognized abnormal information, and correctly recognized abnormal information.

[0049] Step 104: Use the policy model Actor model in the current reinforcement learning model to evaluate the potential benefits of each scan plugin based on the device feature vector and the reference scan results to obtain a scan plugin selection policy; and for each scan plugin that needs to be used when currently scanning the target device indicated by the scan plugin selection policy, determine the scan priority of this scan plugin when currently scanning the target device according to the resources occupied during its execution, the risk situation of the abnormal information that can be scanned out by this scan plugin, and the most recent time when this scan plugin was used.

[0050] In this embodiment, a state space can be constructed based on the device feature vector and the reference scan results. The state space here is defined as: {device feature vector, reference scan results}. Then, the state space is input into the current reinforcement learning model, so that the policy model (denoted as the Actor model) in the current reinforcement learning model can evaluate the potential benefits of each scan plugin based on the input state space to obtain a scan plugin selection policy. It can be seen that in this embodiment, when currently scanning the target device, based on the current features of the dynamically perceived target device, that is, the above-mentioned device feature vector, and the reference scan results obtained based on the historical scan behavior of the target device (which can include user feedback such as missed reports, false reports, unrecognized abnormal information, correctly recognized abnormal information, etc.), a scan plugin selection policy for currently scanning the target device is dynamically generated (which indicates which scan plugins are determined to be used for scanning the target device when currently scanning the target device).

[0051] As an embodiment, the scan plugin selection policy can be represented by an action space. The action space can be defined as: {set of whether to call the scan plugin}. For example, if there are three scan plugins A, B, and C, the action space can be {1, 0, 1}, indicating that scan plugins A and C can be used for currently scanning the target device, and scan plugin B is not used for currently scanning the target device.

[0052] Optionally, in this embodiment, for each scanning plugin indicated by the scanning plugin selection policy and required to be used when currently scanning a target device, the scanning priority of the scanning plugin when currently scanning the target device is determined according to the resources occupied by the scanning plugin during execution, the risk situation of the abnormal information that can be scanned by the scanning plugin, and the most recent time when the scanning plugin was used. The purpose of determining the scanning priority of each scanning plugin when currently scanning the target device is that subsequently, according to the priority scheduling algorithm, the scanning plugin with a higher scanning priority can be preferentially called for scanning.

[0053] One way to determine the scanning priority will be described by way of example below and will not be elaborated here for the time being.

[0054] Step 105: According to the scanning priorities of the scanning plugins and in accordance with the priority scheduling algorithm, schedule the scanning plugins to scan the target device to obtain the current scanning results of the scanning plugins; determine the reward values of the scanning plugins based on the current scanning results of the scanning plugins and the reward functions of the scanning plugins, and use the value Critic model in the current reinforcement learning model to evaluate the behavior performance of the scanning plugin selection policy based on the reward values of the scanning plugins, so as to optimize the Actor model according to the evaluation results of the Critic model; compared with before optimization, the optimized Actor model outputs a better scanning plugin selection policy.

[0055] It can be seen that in this embodiment, even if the scanning plugin selection policy is obtained, it does not blindly use each scanning plugin indicated by the scanning plugin selection policy and required to be used when currently scanning the target device for scanning. Instead, by means of the dynamically determined scanning priorities of the scanning plugins, in accordance with the priority scheduling algorithm, the scanning plugin with the highest scanning priority is preferentially scheduled for scanning, which realizes that the scanning plugin with the highest benefit is preferentially scheduled.

[0056] It can also be seen that in this embodiment, after scheduling the scanning plugins to scan the target device, the reward values of the scanning plugins are determined based on the current scanning results of the scanning plugins and the reward functions of the scanning plugins, and the Actor model is adjusted and optimized in real time, that is, it realizes the real-time update of the Actor model by using the current scanning results feedback of each scanning plugin during the application process. Different from the existing convolutional neural network, once the training is completed, the fixed model parameters of the convolutional neural network are used during subsequent multiple uses, and the model parameters cannot be dynamically feedback adjusted, which is likely to result in the repeated invocation of the same scanning plugin in multiple rounds of scanning, leading to resource waste, or a certain actually existing scanning plugin not being invoked in multiple rounds of scanning, resulting in false alarms.

[0057] As an embodiment, the reward function of any scanning plugin can be expressed by the following formula:

[0058] ;

[0059] Among them, represents the reward value, , , are weight coefficients.

[0060] As an example, represents the scanning reward item, which is determined according to the risk situation of the abnormal information that can be scanned by the scanning plug-in and the success ratio of the scanning plug-in being successfully called to scan the abnormal information. Specifically, is represented by the following formula: ; among them, represents the risk score of the i-th scanning plug-in, which is specifically determined according to the risk situation of the abnormal information that can be scanned by the i-th scanning plug-in. For example, the higher the risk level of the abnormal information that can be scanned by the i-th scanning plug-in, the is higher, and vice versa. represents the success ratio of the i-th scanning plug-in being successfully called to scan the abnormal information within a period of time.

[0061] As an example, represents the penalty item for false alarms of the scanning plug-in and the situation where the scanning plug-in is called but no abnormal information is detected, which is used to indicate that the call to the scanning plug-in is reduced when the scanning plug-in performs poorly in the current state. Specifically, is represented by the following formula: ; among them, when the scanning plug-in does not have false alarms and the scanning plug-in is called but no abnormal information is detected, is the first value such as 0, otherwise it is the second value such as 1.

[0062] As an example, represents the exploration reward item; is determined based on the number of times the scanning plug-in is called and the frequency suppression factor to prevent the scanning plug-in from being over-called; the number of times the new scanning plug-in is called is recorded as the initial value such as 0. Specifically, is represented by the following formula: ; among them, , is an expression, which is the default value such as 1 when the number of times the scanning plug-in is called is less than the set threshold, and can be the set value such as 0 in other cases. When is the default value such as 1, it means is valid; represents the frequency suppression factor, represents the number of times the scanning plug-in is called within a period of time, represents the total number of scans during this period.

[0063] After determining the reward values of each scanning plugin based on the current scanning results of each scanning plugin and the reward function of each scanning plugin, the reward values of each scanning plugin and the above-mentioned device feature vector are input into the current reinforcement learning model, so that the value Critic model in the current reinforcement learning model evaluates the behavior performance of the scanning plugin selection strategy, and the above-mentioned Actor model is optimized according to the evaluation result of the Critic model; compared with before optimization, the optimized Actor model outputs a better scanning plugin selection strategy. That is, by means of the evaluation result of the Critic model, the Actor model is continuously optimized to understand the logical relationship between different state spaces when using the reinforcement learning model subsequently, so that a better scanning plugin selection strategy will be output when scanning the target device in the same state space subsequently, achieving the purpose of continuously optimizing the scanning plugin selection strategy.

[0064] So far, the Figure 1 shown process is completed. To make Figure 1 the shown process more intuitive and clear, Figure 3 the specific implementation is illustrated by the keywords involved in each step.

[0065] Through Figure 1 the shown process, it can be seen that in this embodiment, through reinforcement learning, based on the device feature vector of the target device sensed in real time and the historical scanning results of the target device, the current scanning plugin selection strategy is dynamically generated to ensure that the dynamically generated scanning plugin selection strategy has strong dynamic adaptability, thereby reducing false positives and false negatives.

[0066] Furthermore, in this embodiment, by means of the reward mechanism of reinforcement learning, according to the historical performance of different scanning plugins and the current scanning results of the target device (and even plus user feedback information), the Actor model in the reinforcement learning model is optimized and adjusted in real time to continuously optimize the subsequent scanning plugin selection strategy and improve the detection efficiency and accuracy.

[0067] Furthermore, in this embodiment, the target device is profiled to describe the target device with as comprehensive information as possible, and according to the coding method matched by each feature data in the device profile, each feature data in the device profile is encoded to generate the device feature vector of the target device, which ensures the planning and management of data, avoids the conventional convolutional neural network forcing each feature data to be encoded into an equal-length vector according to a unified coding format and increasing useless data and other noises, and can further ensure the accuracy of the strategy when generating a scanning strategy for the target device based on the device feature vector subsequently.

[0068] The following describes how to determine the scanning priority of the scanning plugin in step 104 above:

[0069] See Figure 4 , Figure 4This is the flowchart for implementing step 104 provided by the embodiments of this application. As Figure 4 shown, this process may include the following steps:

[0070] Step 401: For each scanning plugin, determine the risk score of the scanning plugin according to the risk situation of the abnormal information that can be scanned by the scanning plugin .

[0071] For example, the higher the risk level of the abnormal information that can be scanned by the i-th scanning plugin, the higher the risk score of the i-th scanning plugin , and vice versa.

[0072] Step 402: Determine the time decay factor corresponding to the scanning plugin according to the current time and the most recent time when the scanning plugin was used .

[0073] As an embodiment, the time decay factor can be implemented by an exponential function. For example, the time decay factor is determined according to the following formula: ;

[0074] where represents the current time, and represents the most recent time when the scanning plugin was used.

[0075] Step 403: Obtain the similarity Sim between the feature vector corresponding to the scanning plugin and the device feature vector.

[0076] In this embodiment, the feature vector corresponding to the scanning plugin is also encoded according to the encoding method of each feature data matching in the plugin profile of the scanning plugin. The feature vector corresponding to the scanning plugin is used to describe the scanning plugin. The plugin profile contains feature data corresponding to multiple different feature types, and each feature type in the plugin profile is similar to each feature type in the above device profile.

[0077] Step 404: Determine the comprehensive score of the scanning plugin according to the weight assigned to the risk score , the weight assigned to the time decay factor , and the weight assigned to the similarity Sim.

[0078] For example, is represented by the following formula:

[0079] = ; where , , are the risk scores Allocated weight, time decay factor The allocated weight, and the weight allocated to the similarity Sim.

[0080] Step 405, according to the score corresponding to the resources occupied by the scanning plug-in during execution and the comprehensive score , determine the scanning priority of the scanning plug-in when currently scanning the target device.

[0081] For example, the more resources the scanning plug-in occupies during execution, the smaller the corresponding score Conversely, the fewer resources the scanning plug-in occupies during execution, the larger the corresponding score will be.

[0082] As an embodiment, this embodiment can use and the quotient of as the scanning priority of the scanning plug-in. Optionally, the smaller the quotient, the higher the scanning priority.

[0083] So far, the process shown Figure 4 is completed.

[0084] Through Figure 4 the process shown, an example is given of how to determine the scanning priority of the scanning plug-in.

[0085] As an embodiment, this embodiment adopts -greedy strategy. When the Actor model decides the scanning plug-in selection strategy, a very small positive number is set to select an unknown scanning plug-in with a probability of (<1), and the remaining 1 - probability is used to select the scanning plug-in with the largest value among the existing scanning plug-ins. As the reinforcement learning model is used, the number of unknown scanning plug-ins decided by the Actor model becomes less and less, gradually reaching the above . Under this premise, the reinforcement learning model changes from exploration to exploitation. That is, the current reinforcement learning model at this time is the optimal reinforcement learning model. Under this premise, before this embodiment determines the reward value of each scanning plug-in based on the current scanning results of each scanning plug-in and the reward function of each scanning plug-in, it further identifies whether the current reinforcement learning model is the optimal reinforcement learning model. If so, the current process ends. If not, continue to execute to determine the reward value of each scanning plug-in based on the current scanning results of each scanning plug-in and the reward function of each scanning plug-in.

[0086] The method provided by the embodiments of the present application has been described above. Next, the device provided by the embodiments of the present application will be described.

[0087] See Figure 5 , Figure 5The structural diagram of the device provided by the embodiment of the present application. As Figure 5 shown, this process mainly includes: a data perception module, a policy generation module, a policy execution module, and a policy feedback module.

[0088] As an embodiment, the data perception module is used to collect the device profile of the target device, determine the encoding method matched by each feature data in the device profile of the target device; the device profile contains feature data corresponding to multiple different feature types, and the encoding methods matched by the feature data of at least two different feature types are different; encode each feature data in the device profile according to the encoding method matched by each feature data in the device profile, and generate a device feature vector of the target device according to the encoding results of each feature data; and collect the historical scan results of the target device.

[0089] The policy generation module is used to obtain a reference scan result based on the historical scan result of the target device, and use the policy model Actor model in the current reinforcement learning model to evaluate the potential benefits of each scan plug-in based on the device feature vector and the reference scan result to obtain a scan plug-in selection policy.

[0090] The policy execution module is used for each scan plug-in that needs to be used when scanning the target device currently indicated by the scan plug-in selection policy generated by the policy generation module. According to the resources occupied by the scan plug-in during execution, the risk situation of the abnormal information that the scan plug-in can scan out, the similarity between the feature vector corresponding to the scan plug-in and the device feature vector, and the most recent time when the scan plug-in was used, determine the scan priority of the scan plug-in when scanning the target device currently; according to the scan priorities of each scan plug-in and according to the priority scheduling algorithm, schedule each scan plug-in to scan the target device to obtain the current scan results of each scan plug-in.

[0091] The policy feedback module is used to feedback the current scan results, such as whether the called scan plug-in recognizes abnormal information, whether false alarms or missed alarms occur, to the policy generation module, and determine the reward value of each scan plug-in based on the current scan results of each scan plug-in and the reward function of each scan plug-in and feedback it to the policy generation module.

[0092] The policy generation module, according to the feedback of the policy feedback module, uses the value Critic model in the current reinforcement learning model to evaluate the behavior performance of the scan plug-in selection policy based on the reward values of each scan plug-in, so as to optimize the Actor model according to the evaluation results of the Critic model; the optimized Actor model outputs a better scan plug-in selection policy compared with before optimization.

[0093] As an example, the device portrait involves at least two of the port type, service version type, device description type, and operating system type;

[0094] The feature data under the port type at least includes: the service opening status of at least one port; the coding method for the feature data under the port type is: first, use the high-dimensional sparse matrix coding method to code the service opening status of each port under the port type. Among them, when coding, if there is a high-risk port that meets the high-risk definition, then a corresponding weight coefficient is added to the service opening status of this high-risk port; then, the high-dimensional sparse matrix coding is reduced in dimension;

[0095] The feature data under the service version type at least includes: the service version information of the target device; the feature data under the device description type at least includes: the description information of the target device; the feature data under the operating system type at least includes: the operating system of the target device;

[0096] Among them, the coding method for the feature data under the service version type, the coding method for the feature data under the device description type, and the coding method for the feature data under the operating system type are different from the coding method for the feature data under the port type.

[0097] As an example, the coding method for the feature data under the service version type is: use the bert coding method for coding;

[0098] The description information at least includes the device manufacturer, device type, device model, and device firmware version information; the coding method for the feature data under the device description type is: use the transE algorithm to generate word embedding vectors;

[0099] The coding method for the feature data under the operating system type is: use a one-dimensional vector to code the feature data under the operating system type.

[0100] As an example, obtaining the reference scan result based on the historical scan behavior executed on the target device includes:

[0101] Input the historical scan result of the target device into the LSTM model, so that the LSTM model analyzes the historical scan result of the target device and generates a reference scan result; the reference scan result focuses on the most recent scan result of the target device and also refers to other non-recent scan results of the target device.

[0102] As an example, determining the scanning priority of the scanning plugin when currently scanning the target device based on the resources occupied by the scanning plugin during execution, the risk situation of the abnormal information that can be scanned by the scanning plugin, and the most recent time when the scanning plugin was used includes:

[0103] Determine the risk score of the scanning plugin according to the risk situation of the abnormal information that can be scanned by the scanning plugin ;

[0104] Determine the time decay factor corresponding to the scanning plugin according to the current time and the most recent time when the scanning plugin was used ;

[0105] Obtain the similarity Sim between the feature vector corresponding to the scanning plugin and the device feature vector;

[0106] Based on the risk score The assigned weight, the time decay factor The assigned weight, and the weight assigned to the similarity Sim, determine the comprehensive score of the scanning plugin ;

[0107] According to the score corresponding to the resource occupation situation of the scanning plugin during execution And the comprehensive score , determine the scanning priority of the scanning plugin when currently scanning the target device.

[0108] As an example, the time decay factor Is determined according to the following formula:

[0109] ;

[0110] Wherein, Represents the current time, Represents the most recent time when the scanning plugin was used.

[0111] As an example, the reward function of any scanning plugin is represented by the following formula: ;

[0112] Wherein, Represents the scanning reward item, which is determined according to the risk situation of the abnormal information that can be scanned by the scanning plugin and the success ratio of the scanning plugin being called successfully to scan the abnormal information;

[0113] Represents the penalty item for false positives of the scanning plugin and the scanning plugin being called but no abnormal information being detected, which is used to indicate reducing the call of the scanning plugin when the scanning plugin performs poorly in the current state;

[0114] Indicates the exploration reward item; Determined based on the number of times the scanning plugin is called and the frequency suppression factor to prevent the scanning plugin from being over-called; the number of times the new scanning plugin is called is recorded as the initial value;

[0115] 、 、 Are weight coefficients.

[0116] As an example, Is represented by the following formula:

[0117] ;

[0118] Where, Represents the risk score of the i-th scanning plugin, and is specifically determined according to the risk situation of the abnormal information that can be scanned by the i-th scanning plugin; Represents the success ratio of the i-th scanning plugin being called to successfully scan abnormal information within a period of time;

[0119] Is represented by the following formula: ; Where, when the scanning plugin does not give a false alarm and the scanning plugin is called without detecting abnormal information, Is the first value, otherwise it is the second value;

[0120] Is represented by the following formula: ;

[0121] Where, , Is the default value when the number of times the scanning plugin is called is less than the set threshold, Represents the frequency suppression factor, Represents the number of times the scanning plugin is called within a period of time, Represents the total number of scans during this period.

[0122] As an example, before determining the reward value of each scanning plugin based on the current scanning results of each scanning plugin and the reward function of each scanning plugin, the method further includes:

[0123] Identifying whether the current reinforcement learning model meets The optimal reinforcement learning model required by the greedy policy, the Greedy policy is used to describe the unknown probability Decreases over time, and when the probability of the unknown scanning plugin explored by the reinforcement learning model is less than or equal to , Determine that the current reinforcement learning model meets The optimal reinforcement learning model required by the greedy strategy; if so, end the current process, and if not, continue to determine the reward values of each scanning plugin based on the current scanning results of each scanning plugin and the reward function of each scanning plugin.

[0124] So far, the Figure 5 structural description of the device shown is completed.

[0125] The embodiments of the present application also provide Figure 5 the hardware structure of the device shown. Refer to Figure 6 , Figure 6 which is the structural diagram of the electronic device provided by the embodiments of the present application. As Figure 6 shown, the hardware structure may include: a processor and a machine-readable storage medium, and the machine-readable storage medium stores machine-executable instructions that can be executed by the processor; the processor is used to execute the machine-executable instructions to implement the method disclosed in the above examples of the present application.

[0126] Based on the same application concept as the above method, the embodiments of the present application also provide a machine-readable storage medium, and a number of computer instructions are stored on the machine-readable storage medium. When the computer instructions are executed by a processor, the method disclosed in the above examples of the present application can be implemented.

[0127] Exemplarily, the above machine-readable storage medium can be any electronic, magnetic, optical or other physical storage device, and can contain or store information, such as executable instructions, data, etc. For example, the machine-readable storage medium can be: RAM (Random Access Memory), volatile memory, non-volatile memory, flash memory, storage drive (such as a hard disk drive), solid-state drive, any type of storage disk (such as an optical disc, DVD, etc.), or a similar storage medium, or a combination thereof.

[0128] The above are only the embodiments of the present application and are not used to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.

Claims

1. A method for generating intelligent vulnerability scanning strategies based on reinforcement learning, characterized in that: The method is applied to a scanning device, and the method comprises: Determine an encoding method for matching each feature data in a device portrait of a target device; the device portrait is generated based on device information of the target device obtained by actively scanning the target device; the device portrait includes feature data corresponding to a plurality of different feature types, and feature data of at least two different feature types are matched with different encoding methods; According to the encoding method matched with each feature data in the device portrait, each feature data in the device portrait is encoded, and according to the encoding result of each feature data, a device feature vector of the target device is generated; Based on the historical scanning behavior executed on the target device, a reference scanning result is obtained; the reference scanning result at least includes: the scanning situation of each scanning plug-in when scanning the target device, and the scanning situation at least includes: missed abnormal information, false positive, unrecognized abnormal information, and correctly recognized abnormal information; Using the strategy model Actor model in the current reinforcement learning model, based on the device feature vector and the reference scanning result, the potential benefits of each scanning plug-in are evaluated to obtain a scanning plug-in selection strategy; and for each scanning plug-in indicated by the scanning plug-in selection strategy to be used when scanning the target device, according to the resources occupied by the scanning plug-in during execution, the risk of abnormal information that can be scanned by the scanning plug-in, and the most recent time when the scanning plug-in was used, the scanning priority of the scanning plug-in when scanning the target device is currently determined; According to the scanning priority of each scanning plug-in and in accordance with the priority scheduling algorithm, each scanning plug-in is scheduled to scan the target device to obtain the current scanning result of each scanning plug-in; based on the current scanning result of each scanning plug-in and the reward function of each scanning plug-in, the reward value of each scanning plug-in is determined, and the value Critic model in the current reinforcement learning model evaluates the behavioral performance of the scanning plug-in selection strategy based on the reward value of each scanning plug-in, so as to optimize the Actor model according to the evaluation result of the Critic model; the optimized Actor model outputs a better scanning plug-in selection strategy than before optimization.

2. The method according to claim 1, characterized in that The device portrait involves at least two of a port type, a service version type, a device description type, and an operating system type; The characteristic data under the port type at least includes: the service opening status of at least one port; the encoding method of matching the characteristic data under the port type is: firstly use a high-dimensional sparse matrix encoding method to encode the service opening status of each port under the port type, wherein, when encoding, if there is a high-risk port that meets the high-risk definition, the corresponding weight coefficient is added to the service opening status of the high-risk port; then, the high-dimensional sparse matrix encoding is reduced in dimension; The characteristic data under the service version type at least includes: the service version information of the target device; the characteristic data under the device description type at least includes: the description information of the target device; the characteristic data under the operating system type at least includes: the operating system of the target device; Among them, the encoding method of feature data matching under the service version type, the encoding method of feature data matching under the device description type, and the encoding method of feature data matching under the operating system type are different from the encoding method of feature data matching under the port type.

3. The method according to claim 2, characterized in that The encoding method of the feature data matching under the service version type is: encoding using the BERT encoding method; The description information at least includes the device manufacturer, device type, device model, and device firmware version information; the encoding method of the feature data matching under the device description type is: using the transE algorithm to generate a word embedding vector; The encoding method for matching the characteristic data under the operating system type is: using a one-dimensional vector to encode the characteristic data under the operating system type.

4. The method according to claim 1, characterized in that: The obtaining of a reference scanning result based on a historical scanning behavior performed on the target device includes: The historical scan results of the target device are input into the LSTM model, so that the LSTM model analyzes the historical scan results of the target device to generate a reference scan result; the reference scan result focuses on the most recent scan result of the target device and also refers to other non-most recent scan results of the target device.

5. The method according to claim 1, characterized in that The determining, according to the resources occupied by the scanning plug-in when executing, the risk of abnormal information that can be scanned by the scanning plug-in, and the most recent time when the scanning plug-in was used, the scanning priority of the scanning plug-in when currently scanning the target device comprises: Determine the risk score of the scanning plug-in based on the risk of abnormal information that the scanning plug-in can scan ; Determine the time decay factor corresponding to the scan plugin based on the current time and the most recent time the scan plugin was used ; Obtaining the similarity Sim between the feature vector corresponding to the scanning plug-in and the device feature vector; Based on risk score Assigned weights, time decay factors The weight assigned to the scan plugin and the weight assigned to the similarity Sim determine the overall score of the scan plugin. ; The score corresponding to the resources occupied by the scanning plug-in during execution and the composite score , determining the scanning priority of the scanning plug-in when currently scanning the target device.

6. The method according to claim 5, characterized in that The time decay factor Determine as follows: ; in, Indicates the current time. Indicates the most recent time the scan plugin was used.

7. The method according to claim 1, characterized in that The reward function of any scanning plugin is expressed as follows: ; in, Indicates the scanning reward item, which is determined based on the risk of abnormal information that can be scanned by the scanning plug-in and the success rate of the scanning plug-in being called to successfully scan abnormal information; Indicates the penalty item for scanning plug-in false alarms and scanning plug-in being called but no abnormal information is detected. It is used to instruct the scanning plug-in to reduce the call of the scanning plug-in when it performs poorly in the current state; Indicates exploration reward items; The number of times the scanning plug-in is called is determined based on the number of times the scanning plug-in is called and the frequency suppression factor to prevent the scanning plug-in from being called excessively; the number of times the new scanning plug-in is called is recorded as the initial value; , , is the weight coefficient.

8. The method according to claim 7, characterized in that It is expressed by the following formula: ; in, represents the risk score of the i-th scanning plug-in, which is determined according to the risk of abnormal information that can be scanned by the i-th scanning plug-in; Indicates the success rate of the i-th scanning plug-in being called to successfully scan abnormal information within a period of time; It is expressed by the following formula: ; Among them, when the scanning plug-in does not have a false alarm and the scanning plug-in is called without detecting abnormal information, is the first value, otherwise it is the second value; It is expressed by the following formula: ; in, , when the number of times the scanning plug-in is called is less than the set threshold efficient, is the default value; represents the frequency suppression factor, Indicates the number of times the scanning plug-in is called within a period of time. Indicates the total number of scans during this period of time.

9. The method according to claim 1, characterized in that: Before determining the reward value of each scanning plug-in based on the current scanning result of each scanning plug-in and the reward function of each scanning plug-in, the method further includes: Identify whether the current reinforcement learning model meets -The optimal reinforcement learning model required by the greedy strategy, the ϵ-greedy strategy is used to describe the unknown probability Decreases over time, and the probability of an unknown scanning plug-in explored by the reinforcement learning model is less than or equal to When the current reinforcement learning model is determined to satisfy -The optimal reinforcement learning model required by the greedy strategy; If yes, then the current process ends; if no, then the reward value of each scanning plug-in is determined based on the current scanning result of each scanning plug-in and the reward function of each scanning plug-in.

10. A scanning device, characterized in that: The device includes: A data perception module, used to determine the encoding method matching each feature data in the device portrait of the target device; the device portrait is generated based on the device information of the target device obtained by actively scanning the target device; the device portrait contains feature data corresponding to multiple different feature types, and the encoding methods matching feature data of at least two different feature types are different; and, used to encode each feature data in the device portrait according to the encoding method matching each feature data in the device portrait, and generate a device feature vector of the target device according to the encoding results of each feature data; A strategy generation module is used to obtain a reference scanning result based on the historical scanning behavior executed on the target device; the reference scanning result at least includes: the scanning situation of each scanning plug-in when scanning the target device, and the scanning situation at least includes: missed abnormal information, false positive, unrecognized abnormal information, and correctly recognized abnormal information; and, using the strategy model Actor model in the current reinforcement learning model, based on the device feature vector and the reference scanning result, evaluate the potential benefits of each scanning plug-in to obtain a scanning plug-in selection strategy; a policy execution module, configured to determine, for each scanning plug-in that is required to be used when scanning the target device currently as indicated by the scanning plug-in selection policy, a scanning priority of the scanning plug-in when scanning the target device currently according to resources occupied by the scanning plug-in when executing, risk of abnormal information that can be scanned by the scanning plug-in, and a recent time when the scanning plug-in was used; and According to the scanning priority of each scanning plug-in and the priority scheduling algorithm, each scanning plug-in is scheduled to scan the target device to obtain the current scanning result of each scanning plug-in; A strategy feedback module, used to determine the reward value of each scanning plug-in based on the current scanning result of each scanning plug-in and the reward function of each scanning plug-in, and feed back the reward value of each scanning plug-in and the current scanning result of each scanning plug-in to the strategy generation module; The strategy generation module is further used to evaluate the behavioral performance of the scanning plug-in selection strategy based on the reward value of each scanning plug-in according to the feedback of the strategy feedback module and using the value Critic model in the current reinforcement learning model, so as to optimize the Actor model according to the evaluation result of the Critic model; the optimized Actor model outputs a better scanning plug-in selection strategy than before optimization.

Citation Information

Patent Citations

  • Off-line reinforcement learning method, device and equipment for target control

    CN114186474A

  • Distributed network detection method of ant colony algorithm based on adaptive pheromone update

    CN117221182A