Vulnerability identification method and device, computer equipment, readable storage medium and program product

By using RASP components to obtain behavior samples while the application is running and using reinforcement learning and integrated learning to train the vulnerability identification model, the problem that traditional methods are difficult to process and update data in real time is solved, and accurate detection and real-time defense of vulnerability behavior is achieved.

CN120197182APending Publication Date: 2025-06-24CHINA SOUTHERN POWER GRID COMPANY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510451659.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

Traditional methods are difficult to process and update data at the runtime of applications in real time, resulting in inefficient vulnerability identification and defense.

Method used

By applying the self-protecting RASP components at runtime, we can obtain normal samples and vulnerability attack samples of multiple behaviors during the application operation, forming multiple vulnerability data sets. Then, the single vulnerability recognition model is trained using reinforcement learning, and the comprehensive vulnerability recognition model is trained through feature cross-learning and integrated learning to achieve real-time identification and defense against vulnerability attacks.

Benefits of technology

It realizes more accurate detection and real-time defense of vulnerability behavior, and improves the dynamic analysis and defense capabilities of application security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120197182A_ABST
    Figure CN120197182A_ABST
Patent Text Reader

Abstract

The invention relates to a vulnerability identification method and device, computer equipment, a computer readable storage medium and a computer program product. The method comprises the following steps: acquiring a preset number of categories of normal behavior samples and corresponding categories of vulnerability attack samples in the running process of an application program through a runtime application self-protection RASP component, and forming a preset number of vulnerability data sets; each vulnerability data set is composed of a normal behavior sample and a vulnerability attack sample of a corresponding category; for each vulnerability data set, training the to-be-trained single vulnerability recognition model of the corresponding category in a reinforcement learning mode to obtain a preset number of single vulnerability recognition models; based on the preset number of vulnerability data sets and the preset number of single vulnerability recognition models, training the vulnerability recognition models through feature cross learning and ensemble learning; the vulnerability identification model is used for identifying vulnerability attacks. By adopting the method, a more accurate vulnerability identification model can be obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of application security, and particularly to a vulnerability identification method, apparatus, computer device, computer-readable storage medium, and computer program product. Background Art

[0002] The RASP behavior self-learning vulnerability identification patent technology aims to solve technical problems in the field of application security, particularly those related to vulnerability identification and defense. The goal of this technology is to dynamically analyze and learn the behavior of an application during runtime in order to detect and defend against potential vulnerabilities in a timely manner. It uses a self-learning method, applying machine learning techniques, to extract and analyze features from the execution process of the application, establish models, and through ensemble learning between models to identify known or unknown vulnerabilities and attacks.

[0003] Traditional methods rely on collecting and processing a large amount of application runtime data, and it is difficult to process and update models in real time for high-frequency data updates. Summary of the Invention

[0004] Based on this, it is necessary to provide a vulnerability identification method, apparatus, computer device, computer-readable storage medium, and computer program product that can process and update in real time for the above technical problems.

[0005] In a first aspect, the present application provides a vulnerability identification method, including:

[0006] By running the Runtime Application Self-Protection (RASP) component, obtain normal behavior samples and vulnerability attack samples corresponding to various behaviors during the running of the application, and form multiple vulnerability data sets; each vulnerability data set is composed of normal behavior samples and vulnerability attack samples corresponding to one behavior;

[0007] For each vulnerability data set, determine the target behavior corresponding to the current vulnerability data set, and obtain the single-vulnerability identification model to be trained corresponding to the target behavior; based on the current vulnerability data set, train the single-vulnerability identification model to be trained by means of reinforcement learning to obtain the trained single-vulnerability identification model corresponding to the target behavior;

[0008] Based on the multiple vulnerability data sets and the trained single-vulnerability identification model corresponding to each behavior, train the vulnerability identification model by means of feature cross-learning and ensemble learning; the vulnerability identification model is used to identify vulnerability attacks.

[0009] In one of the embodiments, the training the single-vulnerability identification model to be trained by means of reinforcement learning based on the current vulnerability data set to obtain the trained single-vulnerability identification model corresponding to the target behavior includes:

[0010] Input the current training samples in the current vulnerability dataset into the single vulnerability recognition model to be trained, and obtain vulnerability operation behaviors and vulnerability operation rewards; put the current training samples, the vulnerability operation behaviors, the vulnerability operation rewards, and the next training sample into a buffer as an experience sample; whenever the number of experience samples put into the buffer exceeds a preset value, optimize the single vulnerability recognition model to be trained based on the experience samples in the buffer to obtain a trained single vulnerability recognition model.

[0011] In one embodiment, optimizing the single vulnerability recognition model to be trained based on the experience samples in the buffer includes:

[0012] Randomly sample a first preset number of experience samples from the buffer; based on the first preset number of experience samples, obtain a first preset number of vulnerability operation rewards, and determine a first preset number of target vulnerability values; the vulnerability operation rewards and the target vulnerability values correspond one by one; based on the vulnerability operation rewards and the target vulnerability values, determine a loss function, and optimize the single vulnerability recognition model to be trained based on the loss function.

[0013] In one embodiment, training the vulnerability recognition model by means of feature cross-learning and ensemble learning includes:

[0014] Input the training samples of each vulnerability dataset into the corresponding single vulnerability recognition model for feature extraction to obtain a vulnerability feature set corresponding to each vulnerability dataset; perform combination and cross-analysis on the vulnerability features in different vulnerability feature sets to obtain a cross-feature dataset; based on the cross-feature dataset, iteratively train the vulnerability recognition model by means of ensemble learning.

[0015] In one embodiment, the iterative training of the vulnerability recognition model by means of ensemble learning, where the process of one iterative training includes:

[0016] For the training process of the i-th iteration, based on the cross-feature dataset and the vulnerability recognition model of the (i - 1)-th iteration, obtain the negative gradient; i is not less than 2; based on the negative gradient and the cross-feature dataset, obtain the decision tree of the i-th iteration; based on the vulnerability recognition model of the i-th iteration and the decision tree of the i-th iteration, obtain the vulnerability recognition model of the i-th iteration.

[0017] In one embodiment, the multiple behaviors include at least one of user input behavior, API interface call behavior, file operation behavior, network communication behavior, abnormal behavior, permission and access control behavior, and authorization and authentication behavior.

[0018] In a second aspect, the present application also provides a vulnerability recognition device, including:

[0019] An acquisition module, configured to obtain normal behavior samples and vulnerability attack samples corresponding to various behaviors during the running of an application by running a runtime application self-protection (RASP) component, and form a plurality of vulnerability data sets; each vulnerability data set consists of normal behavior samples and vulnerability attack samples corresponding to one behavior.

[0020] A strengthening module, configured to, for each vulnerability data set, determine a target behavior corresponding to the current vulnerability data set, and obtain a to-be-trained single-vulnerability recognition model corresponding to the target behavior; based on the current vulnerability data set, train the to-be-trained single-vulnerability recognition model by using a reinforcement learning method to obtain a trained single-vulnerability recognition model corresponding to the target behavior.

[0021] An integration module, configured to train a vulnerability recognition model by using a feature cross-learning and an ensemble learning method based on the plurality of vulnerability data sets and the trained single-vulnerability recognition models corresponding to each behavior; the vulnerability recognition model is used to identify vulnerability attacks.

[0022] In a third aspect, the present application further provides a computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:

[0023] By running a runtime application self-protection (RASP) component, obtain normal behavior samples and vulnerability attack samples corresponding to various behaviors during the running of an application, and form a plurality of vulnerability data sets; each vulnerability data set consists of normal behavior samples and vulnerability attack samples corresponding to one behavior.

[0024] For each vulnerability data set, determine a target behavior corresponding to the current vulnerability data set, and obtain a to-be-trained single-vulnerability recognition model corresponding to the target behavior; based on the current vulnerability data set, train the to-be-trained single-vulnerability recognition model by using a reinforcement learning method to obtain a trained single-vulnerability recognition model corresponding to the target behavior.

[0025] Based on the plurality of vulnerability data sets and the trained single-vulnerability recognition models corresponding to each behavior, train a vulnerability recognition model by using a feature cross-learning and an ensemble learning method; the vulnerability recognition model is used to identify vulnerability attacks.

[0026] In a fourth aspect, the present application further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the following steps are implemented:

[0027] By running the runtime application self - protection RASP component, normal behavior samples and vulnerability attack samples corresponding to various behaviors during the running of the application are obtained, forming multiple vulnerability data sets; each vulnerability data set consists of normal behavior samples and vulnerability attack samples corresponding to one behavior.

[0028] For each vulnerability data set, determine the target behavior corresponding to the current vulnerability data set, and obtain the single - vulnerability recognition model to be trained corresponding to the target behavior; based on the current vulnerability data set, train the single - vulnerability recognition model to be trained by means of reinforcement learning to obtain the trained single - vulnerability recognition model corresponding to the target behavior.

[0029] Based on the multiple vulnerability data sets and the trained single - vulnerability recognition models corresponding to each behavior, train a vulnerability recognition model by means of feature cross - learning and ensemble learning; the vulnerability recognition model is used to identify vulnerability attacks.

[0030] In a fifth aspect, the present application also provides a computer program product, including a computer program, which when executed by a processor implements the following steps:

[0031] By running the runtime application self - protection RASP component, normal behavior samples and vulnerability attack samples corresponding to various behaviors during the running of the application are obtained, forming multiple vulnerability data sets; each vulnerability data set consists of normal behavior samples and vulnerability attack samples corresponding to one behavior.

[0032] For each vulnerability data set, determine the target behavior corresponding to the current vulnerability data set, and obtain the single - vulnerability recognition model to be trained corresponding to the target behavior; based on the current vulnerability data set, train the single - vulnerability recognition model to be trained by means of reinforcement learning to obtain the trained single - vulnerability recognition model corresponding to the target behavior.

[0033] Based on the multiple vulnerability data sets and the trained single - vulnerability recognition models corresponding to each behavior, train a vulnerability recognition model by means of feature cross - learning and ensemble learning; the vulnerability recognition model is used to identify vulnerability attacks.

[0034] The above-mentioned vulnerability identification method, device, computer device, computer-readable storage medium and computer program product obtain normal behavior samples and vulnerability attack samples corresponding to various behaviors during the operation of the application through the runtime application self-protection (RASP) component, and form multiple vulnerability data sets; each vulnerability data set is composed of normal behavior samples and vulnerability attack samples corresponding to one behavior; for each vulnerability data set, determine the target behavior corresponding to the current vulnerability data set, and obtain the single-vulnerability identification model to be trained corresponding to the target behavior; based on the current vulnerability data set, train the single-vulnerability identification model to be trained by means of reinforcement learning to obtain the trained single-vulnerability identification model corresponding to the target behavior; based on the multiple vulnerability data sets and the trained single-vulnerability identification models corresponding to each behavior, train the vulnerability identification model by means of feature cross-learning and ensemble learning; the vulnerability identification model is used to identify vulnerability attacks. It can detect vulnerability behaviors more accurately. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required for describing the embodiments of the present application or related technologies. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other related drawings can be obtained without creative efforts based on these drawings.

[0036] Figure 1 It is a schematic flowchart of the vulnerability identification method in an embodiment;

[0037] Figure 2 It is a detailed flowchart of the vulnerability identification method in an embodiment;

[0038] Figure 3 It is a structural block diagram of the vulnerability identification device in an embodiment;

[0039] Figure 4 It is an internal structure diagram of the computer device in an embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0040] In order to make the purpose, technical solutions and advantages of the present application clearer, the following further describes the present application in detail with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application, and are not used to limit the present application.

[0041] In one embodiment, as Figure 1As shown, a vulnerability identification method is provided. In this embodiment, the method is exemplified by being applied to a terminal. It can be understood that the method can also be applied to a server, or to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. In this embodiment, the method includes the following steps:

[0042] Step 102: By running the Runtime Application Self-Protection (RASP) component, obtain normal behavior samples and vulnerability attack samples corresponding to various behaviors during the running of the application, and form multiple vulnerability data sets; each vulnerability data set is composed of normal behavior samples and vulnerability attack samples corresponding to one behavior.

[0043] Among them, the RASP component is a tool that can dynamically monitor and defend against security threats during the running of the application; it is integrated into the application internally, monitors the running state of the application in real time, and can actively discover and prevent various attacks on the application.

[0044] Optionally, the above-mentioned various behaviors may include user input behaviors and API interface call behaviors, and normal behavior samples and vulnerability attack samples of user input behaviors during the running of the application, as well as normal behavior samples and vulnerability attack samples of API interface call behaviors can be obtained.

[0045] Among them, the normal behavior sample of the user input behavior is the case where the keywords, data length, and input format entered by the user are all normal, and the vulnerability attack sample of the user input behavior is the case where the keywords, data length, and input format entered by the user are all abnormal.

[0046] Step 104: For each vulnerability data set, determine the target behavior corresponding to the current vulnerability data set, and obtain the single-vulnerability identification model to be trained corresponding to the target behavior; based on the current vulnerability data set, train the single-vulnerability identification model to be trained in a reinforcement learning manner to obtain the trained single-vulnerability identification model corresponding to the target behavior.

[0047] Among them, for each vulnerability data set, a single-vulnerability identification model to be trained is correspondingly trained, and each single-vulnerability identification model is specifically used to identify one type of vulnerability attack behavior.

[0048] Optionally, the single-vulnerability identification model can be a Transformer model.

[0049] Among them, the Transformer model is a deep learning architecture, and the core idea is to use the Attention mechanism to learn the long-term dependencies in the text, that is, to enable the model to dynamically focus on different parts of the input sequence when processing the information at each position.

[0050] Step 106: Based on the multiple vulnerability data sets and the trained single-vulnerability recognition models corresponding to each type of behavior, train a vulnerability recognition model through feature cross-learning and ensemble learning; the vulnerability recognition model is used to identify vulnerability attacks.

[0051] Among them, feature cross-learning refers to the process of combining multiple features in the original data to generate new features. In this way, the interaction and dependence relationships between different features can be captured, thereby improving the model's expressive ability and prediction performance. Ensemble learning is a machine learning technique that creates a more powerful and robust learning model by combining multiple base learners to achieve better performance than a single learner.

[0052] The above-mentioned vulnerability recognition method, device, computer device, computer-readable storage medium, and computer program product obtain normal behavior samples and vulnerability attack samples corresponding to various behaviors during the operation of the application through the runtime application self-protection (RASP) component, and form multiple vulnerability data sets; each vulnerability data set is composed of normal behavior samples and vulnerability attack samples corresponding to one type of behavior; for each vulnerability data set, determine the target behavior corresponding to the current vulnerability data set, and obtain the single-vulnerability recognition model to be trained corresponding to the target behavior; based on the current vulnerability data set, train the single-vulnerability recognition model to be trained by means of reinforcement learning to obtain the trained single-vulnerability recognition model corresponding to the target behavior; based on the multiple vulnerability data sets and the trained single-vulnerability recognition models corresponding to each type of behavior, train a vulnerability recognition model through feature cross-learning and ensemble learning; the vulnerability recognition model is used to identify vulnerability attacks. It can detect vulnerability behaviors more accurately.

[0053] In an exemplary embodiment, the training of the single-vulnerability recognition model to be trained by means of reinforcement learning based on the current vulnerability data set to obtain the trained single-vulnerability recognition model corresponding to the target behavior includes:

[0054] Input the current training sample in the current vulnerability data set into the single-vulnerability recognition model to be trained to obtain a vulnerability operation behavior and a vulnerability operation reward; put the current training sample, the vulnerability operation behavior, the vulnerability operation reward, and the next training sample into a buffer as an experience sample; whenever the number of experience samples put into the buffer exceeds a preset value, optimize the single-vulnerability recognition model to be trained based on the experience samples in the buffer to obtain a trained single-vulnerability recognition model.

[0055] Exemplarily, for a single vulnerability dataset, the current training sample in the current vulnerability dataset is input into the single-vulnerability recognition model to be trained corresponding to the current vulnerability dataset, and vulnerability operation behaviors and vulnerability operation rewards are obtained; the vulnerability operation behaviors include: recording logs but not taking measures, marking as normal traffic, marking as suspicious traffic (warning), directly blocking requests, and triggering an automatic defense strategy (such as modifying firewall rules). The vulnerability operation rewards include: a reward of +1 for correctly identifying normal behavior, a reward of +10 for successfully detecting and blocking an attack, a penalty of -5 for misjudging normal behavior as an attack, a penalty of -10 for failing to block an attack, and a reward of 0 when no abnormality occurs. The current training sample, the vulnerability operation behavior, the vulnerability operation reward, and the next training sample are used as an experience sample and put into the buffer; whenever the number of experience samples put into the buffer exceeds a preset value, the single-vulnerability recognition model to be trained is optimized based on the experience samples in the buffer. After multiple optimizations, the trained single-vulnerability recognition model is obtained.

[0056] In this embodiment, by setting the vulnerability operation behavior and the vulnerability operation reward, the model can gradually and correctly identify the vulnerability attack behavior and take corresponding measures.

[0057] In an exemplary embodiment, the optimizing the single-vulnerability recognition model to be trained based on the experience samples in the buffer includes:

[0058] Randomly sampling a first preset number of experience samples from the buffer; based on the first preset number of experience samples, obtaining the first preset number of vulnerability operation rewards and determining the first preset number of target vulnerability values; the vulnerability operation rewards and the target vulnerability values are in one-to-one correspondence; based on the vulnerability operation rewards and the target vulnerability values, determining a loss function, and optimizing the single-vulnerability recognition model based on the loss function.

[0059] Exemplarily, randomly sampling a first preset number of experience samples from the buffer, obtaining the first preset number of vulnerability operation rewards from the first preset number of experience samples, determining the first preset number of target vulnerability values based on the first preset number of experience samples; the vulnerability operation rewards and the target vulnerability values are in one-to-one correspondence; calculating the model loss according to the first preset number of vulnerability operation rewards and the first preset number of target vulnerability values, and optimizing the single-vulnerability recognition model to be trained through the model loss; if the average loss function of the current first preset number of experience samples is not less than the average loss function of the first preset number of experience samples during the previous training, then the trained single-vulnerability recognition model is obtained.

[0060] In this embodiment, by iteratively training the single-vulnerability recognition model, a model for identifying vulnerabilities can be accurately obtained.

[0061] In an exemplary embodiment, training the vulnerability identification model by means of feature cross-learning and ensemble learning includes:

[0062] Input the training samples of each vulnerability dataset into the corresponding single-vulnerability identification model for feature extraction to obtain a vulnerability feature set corresponding to each vulnerability dataset; perform combination and cross-analysis on the vulnerability features in different vulnerability feature sets to obtain a cross-feature dataset; based on the cross-feature dataset, iteratively train the vulnerability identification model by means of ensemble learning.

[0063] Exemplarily, input the training samples of each vulnerability dataset into the corresponding single-vulnerability identification model for feature extraction to obtain a vulnerability feature set corresponding to each vulnerability dataset; perform combination and cross-analysis on the vulnerability features in different vulnerability feature sets to obtain a cross-feature dataset; based on the cross-feature dataset, use the gradient boosting tree algorithm to iteratively train to obtain a trained vulnerability identification model.

[0064] In this embodiment, by means of feature cross-learning and ensemble learning, multiple single-vulnerability identification models are integrated, which can identify vulnerability attack behaviors more accurately and comprehensively.

[0065] In an exemplary embodiment, the method of iteratively training the vulnerability identification model by means of ensemble learning, where the process of one iteration of training includes:

[0066] For the training process of the i-th iteration, based on the cross-feature dataset and the vulnerability identification model of the (i - 1)-th iteration, obtain the negative gradient; where i is not less than 2; based on the negative gradient and the cross-feature dataset, obtain the decision tree of the i-th iteration; based on the vulnerability identification model of the i-th iteration and the decision tree of the i-th iteration, obtain the vulnerability identification model of the i-th iteration.

[0067] Exemplarily, for the training process of the i-th iteration, based on the cross-feature dataset and the vulnerability identification model of the (i - 1)-th iteration, obtain the negative gradient; where i is not less than 2; based on the cross-feature dataset, establish the decision tree of the i-th iteration by fitting the negative gradient as the target; based on the vulnerability identification model of the i-th iteration and the decision tree of the i-th iteration, obtain the vulnerability identification model of the i-th iteration.

[0068] In this embodiment, the vulnerability identification model is obtained through the gradient boosting tree algorithm, which can gradually optimize the vulnerability identification model, thereby identifying vulnerability attack behaviors more accurately and comprehensively.

[0069] In an exemplary embodiment, the multiple behaviors include at least one of user input behavior, API interface call behavior, file operation behavior, network communication behavior, abnormal behavior, permission and access control behavior, and authorization and authentication behavior.

[0070] Among them, multiple behaviors are not limited to the behaviors of the above categories, and can also be other types of behaviors.

[0071] In this embodiment, through different types of behaviors, rich normal behavior samples and vulnerability attack samples can be obtained, so as to obtain more abundant training data.

[0072] In an exemplary embodiment, such as Figure 2As shown in the figure, a vulnerability identification method is provided. By running the Runtime Application Self-Protection (RASP) component, normal behavior samples of a preset number of categories and vulnerability attack samples of corresponding categories during the operation of the application are obtained, including at least one of: normal user input behavior samples and user input behavior vulnerability attack samples, normal API interface call behavior samples and API interface call vulnerability attack samples, normal file operation behavior samples and file operation behavior vulnerability attack samples, normal network communication behavior samples and network communication behavior vulnerability attack samples, normal abnormal behavior samples and abnormal behavior vulnerability attack samples, normal permission and access control behavior samples and permission and access control behavior vulnerability attack samples, normal authorization and authentication behavior samples and authorization and authentication behavior vulnerability attack samples; and a preset number of vulnerability data sets are formed; each vulnerability data set is composed of normal behavior samples and vulnerability attack samples of corresponding categories. For each vulnerability data set, the current training sample in the vulnerability data set is input into the single vulnerability identification model to be trained corresponding to the vulnerability data set, and vulnerability operation behaviors and vulnerability operation rewards are obtained; the vulnerability operation behaviors include: recording logs but not taking measures, marking as normal traffic, marking as suspicious traffic (warning), directly blocking requests, and triggering an automatic defense strategy (such as modifying firewall rules). The vulnerability operation rewards include: +1 for correctly identifying normal behavior, +10 for successfully detecting and blocking an attack, -5 for misjudging normal behavior as an attack, -10 for failing to block an attack, and 0 for no abnormality. The current training sample, vulnerability operation behavior, vulnerability operation reward, and the next training sample are taken as an experience sample and put into a buffer; whenever the number of experience samples put into the buffer exceeds a preset value, a first preset number of experience samples are randomly sampled from the buffer, the first preset number of vulnerability operation rewards are obtained from the first preset number of experience samples, and based on the first preset number of experience samples, a first preset number of target vulnerability values are determined; the vulnerability operation rewards and the target vulnerability values correspond one by one; the model loss is calculated according to the first preset number of vulnerability operation rewards and the first preset number of target vulnerability values, and the single vulnerability identification model to be trained is optimized through the model loss; if the average loss function of the current first preset number of experience samples is not less than the average loss function of the first preset number of experience samples during the previous training, the trained single vulnerability identification model is obtained; thus, a preset number of single vulnerability identification models are obtained.Based on the preset number of vulnerability data sets and the preset number of single-vulnerability recognition models, input the training samples of each vulnerability data set into the corresponding single-vulnerability recognition model for feature extraction to obtain the vulnerability feature set corresponding to each vulnerability data set; combine and cross-analyze the vulnerability features in different vulnerability feature sets to obtain a cross-feature data set; based on the cross-feature data set, use the gradient boosting tree algorithm for iterative training, where the training process of one iteration includes: for the training process of the i-th iteration, based on the cross-feature data set and the vulnerability recognition model of the (i - 1)-th iteration, obtain the negative gradient; where i is not less than 2; based on the cross-feature data set, establish the decision tree of the i-th iteration by fitting the negative gradient as the target; based on the vulnerability recognition model of the i-th iteration and the decision tree of the i-th iteration, obtain the vulnerability recognition model of the i-th iteration. After multiple iterations, obtain the trained vulnerability recognition model; the vulnerability recognition model is used to identify vulnerability attacks.

[0073] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps does not have a strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or steps or stages in other steps.

[0074] In an exemplary embodiment, as Figure 3 shown, a vulnerability recognition device is provided, including: an acquisition module 301, a reinforcement module 302, and an integration module 303, where:

[0075] The acquisition module is used to obtain normal behavior samples and vulnerability attack samples corresponding to various behaviors during the running of the application program through the runtime application self-protection (RASP) component, and form multiple vulnerability data sets; each vulnerability data set is composed of normal behavior samples and vulnerability attack samples corresponding to one behavior;

[0076] The reinforcement module is used to, for each vulnerability data set, determine the target behavior corresponding to the current vulnerability data set, and obtain the to-be-trained single-vulnerability recognition model corresponding to the target behavior; based on the current vulnerability data set, train the to-be-trained single-vulnerability recognition model in a reinforcement learning manner to obtain the trained single-vulnerability recognition model corresponding to the target behavior;

[0077] An integration module is used to train a vulnerability recognition model by means of feature cross - learning and ensemble learning based on the multiple vulnerability data sets and the single - vulnerability recognition models that have been trained for each type of behavior; the vulnerability recognition model is used to identify vulnerability attacks.

[0078] In an exemplary embodiment, the reinforcement module is further configured to:

[0079] Input the current training samples in the current vulnerability data set into the single - vulnerability recognition model to be trained, and obtain vulnerability operation behaviors and vulnerability operation rewards; put the current training samples, the vulnerability operation behaviors, the vulnerability operation rewards, and the next training sample into a buffer as an experience sample; whenever the number of experience samples put into the buffer exceeds a preset value, optimize the single - vulnerability recognition model to be trained based on the experience samples in the buffer to obtain a trained single - vulnerability recognition model.

[0080] In an exemplary embodiment, the reinforcement module is further configured to:

[0081] Randomly sample a first preset number of experience samples from the buffer; based on the first preset number of experience samples, obtain a first preset number of vulnerability operation rewards and determine a first preset number of target vulnerability values; the vulnerability operation rewards and the target vulnerability values are in one - to - one correspondence; based on the vulnerability operation rewards and the target vulnerability values, determine a loss function and optimize the single - vulnerability recognition model to be trained based on the loss function.

[0082] In an exemplary embodiment, the integration module is further configured to:

[0083] Input the training samples of each vulnerability data set into the corresponding single - vulnerability recognition model for feature extraction to obtain a vulnerability feature set corresponding to each vulnerability data set; perform combination and cross - analysis on the vulnerability features in different vulnerability feature sets to obtain a cross - feature data set; based on the cross - feature data set, iteratively train the vulnerability recognition model by means of ensemble learning.

[0084] In an exemplary embodiment, the integration module is further configured to:

[0085] For the training process of the i - th iteration, based on the cross - feature data set and the vulnerability recognition model of the (i - 1) - th iteration, obtain a negative gradient; where i is not less than 2; based on the negative gradient and the cross - feature data set, obtain the decision tree of the i - th iteration; based on the vulnerability recognition model of the i - th iteration and the decision tree of the i - th iteration, obtain the vulnerability recognition model of the i - th iteration.

[0086] In an exemplary embodiment, the multiple behaviors include at least one of user input behavior, API interface call behavior, file operation behavior, network communication behavior, abnormal behavior, permission and access control behavior, and authorization and authentication behavior.

[0087] Each module in the above vulnerability identification device can be implemented in whole or in part by software, hardware, and their combination. Each of the above modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to each of the above modules.

[0088] In an exemplary embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 4 shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store normal behavior samples and vulnerability attack samples. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals through a network connection. When the computer program is executed by the processor, it implements a vulnerability identification method.

[0089] Those skilled in the art can understand that Figure 4 the structure shown in

[0090] is only a block diagram of a part of the structure related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0091] By running the Runtime Application Self-Protection (RASP) component, obtain normal behavior samples and vulnerability attack samples corresponding to multiple behaviors during the running of the application program, and form multiple vulnerability data sets; each vulnerability data set is composed of normal behavior samples and vulnerability attack samples corresponding to one behavior;

[0092] For each vulnerability dataset, determine the target behavior corresponding to the current vulnerability dataset, and obtain the single-vulnerability recognition model to be trained corresponding to the target behavior; based on the current vulnerability dataset, train the to-be-trained single-vulnerability recognition model by means of reinforcement learning to obtain the trained single-vulnerability recognition model corresponding to the target behavior;

[0093] Based on the multiple vulnerability datasets and the trained single-vulnerability recognition models corresponding to each behavior, train a vulnerability recognition model by means of feature cross-learning and ensemble learning; the vulnerability recognition model is used to identify vulnerability attacks.

[0094] In one embodiment, when the processor executes the computer program, the following steps are further implemented:

[0095] Input the current training sample in the current vulnerability dataset into the to-be-trained single-vulnerability recognition model to obtain a vulnerability operation behavior and a vulnerability operation reward; put the current training sample, the vulnerability operation behavior, the vulnerability operation reward, and the next training sample into a buffer as an experience sample; whenever the number of experience samples put into the buffer exceeds a preset value, optimize the to-be-trained single-vulnerability recognition model based on the experience samples in the buffer to obtain the trained single-vulnerability recognition model.

[0096] In one embodiment, when the processor executes the computer program, the following steps are further implemented:

[0097] Randomly sample a first preset number of experience samples from the buffer; based on the first preset number of experience samples, obtain a first preset number of vulnerability operation rewards, and determine a first preset number of target vulnerability values; the vulnerability operation rewards and the target vulnerability values are in one-to-one correspondence; based on the vulnerability operation rewards and the target vulnerability values, determine a loss function, and optimize the to-be-trained single-vulnerability recognition model based on the loss function.

[0098] In one embodiment, when the processor executes the computer program, the following steps are further implemented:

[0099] Input the training samples of each vulnerability dataset into the corresponding single-vulnerability recognition model for feature extraction to obtain a vulnerability feature set corresponding to each vulnerability dataset; perform combination and cross-analysis on the vulnerability features in different vulnerability feature sets to obtain a cross-feature dataset; based on the cross-feature dataset, iteratively train the vulnerability recognition model by means of ensemble learning.

[0100] In one embodiment, when the processor executes the computer program, the following steps are further implemented:

[0101] For the training process of the i-th iteration, based on the cross-feature dataset and the vulnerability identification model of the (i - 1)-th iteration, obtain the negative gradient; where i is not less than 2; based on the negative gradient and the cross-feature dataset, obtain the decision tree of the i-th iteration; based on the vulnerability identification model of the i-th iteration and the decision tree of the i-th iteration, obtain the vulnerability identification model of the i-th iteration.

[0102] In one embodiment, the multiple behaviors include at least one of: user input behavior, API interface call behavior, file operation behavior, network communication behavior, abnormal behavior, permission and access control behavior, authorization and authentication behavior.

[0103] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0104] By running the runtime application self-protection (RASP) component, obtain normal behavior samples and vulnerability attack samples corresponding to multiple behaviors during the running of the application program, and form multiple vulnerability datasets; each vulnerability dataset consists of normal behavior samples and vulnerability attack samples corresponding to one behavior;

[0105] For each vulnerability dataset, determine the target behavior corresponding to the current vulnerability dataset, and obtain the single-vulnerability identification model to be trained corresponding to the target behavior; based on the current vulnerability dataset, train the single-vulnerability identification model to be trained in a reinforcement learning manner to obtain the trained single-vulnerability identification model corresponding to the target behavior;

[0106] Based on the multiple vulnerability datasets and the trained single-vulnerability identification model corresponding to each behavior, train the vulnerability identification model through feature cross-learning and ensemble learning; the vulnerability identification model is used to identify vulnerability attacks.

[0107] In one embodiment, when the computer program is executed by the processor, the following steps are further implemented:

[0108] Input the current training sample in the current vulnerability dataset into the single-vulnerability identification model to be trained, and obtain the vulnerability operation behavior and the vulnerability operation reward; take the current training sample, the vulnerability operation behavior, the vulnerability operation reward, and the next training sample as an experience sample and put it into the buffer; whenever the number of experience samples put into the buffer exceeds a preset value, optimize the single-vulnerability identification model to be trained based on the experience samples in the buffer to obtain the trained single-vulnerability identification model.

[0109] In one embodiment, when the computer program is executed by the processor, the following steps are further implemented:

[0110] Randomly sample the first preset number of empirical samples from the buffer; based on the first preset number of empirical samples, obtain the first preset number of vulnerability operation rewards, and determine the first preset number of target vulnerability values; the vulnerability operation rewards and the target vulnerability values correspond one by one; based on the vulnerability operation rewards and the target vulnerability values, determine the loss function, and optimize the to-be-trained single-vulnerability recognition model based on the loss function.

[0111] In one embodiment, when the computer program is executed by the processor, the following steps are further implemented:

[0112] Input the training samples of each vulnerability dataset into the corresponding single-vulnerability recognition model for feature extraction to obtain the vulnerability feature set corresponding to each vulnerability dataset; perform combination and cross-analysis on the vulnerability features in different vulnerability feature sets to obtain the cross-feature dataset; based on the cross-feature dataset, iteratively train the vulnerability recognition model by means of ensemble learning.

[0113] In one embodiment, when the computer program is executed by the processor, the following steps are further implemented:

[0114] For the training process of the i-th iteration, based on the cross-feature dataset and the vulnerability recognition model of the (i - 1)-th iteration, obtain the negative gradient; where i is not less than 2; based on the negative gradient and the cross-feature dataset, obtain the decision tree of the i-th iteration; based on the vulnerability recognition model of the i-th iteration and the decision tree of the i-th iteration, obtain the vulnerability recognition model of the i-th iteration.

[0115] In one of the embodiments, the multiple behaviors include at least one of user input behavior, API interface call behavior, file operation behavior, network communication behavior, abnormal behavior, permission and access control behavior, and authorization and authentication behavior.

[0116] In one embodiment, a computer program product is provided, including a computer program, which when executed by the processor implements the following steps:

[0117] By running the runtime application self-protection (RASP) component, obtain the normal behavior samples and vulnerability attack samples corresponding to multiple behaviors during the running of the application program, and form multiple vulnerability datasets; each vulnerability dataset is composed of the normal behavior samples and vulnerability attack samples corresponding to one behavior;

[0118] For each vulnerability dataset, determine the target behavior corresponding to the current vulnerability dataset, and obtain the to-be-trained single-vulnerability recognition model corresponding to the target behavior; based on the current vulnerability dataset, train the to-be-trained single-vulnerability recognition model by means of reinforcement learning to obtain the trained single-vulnerability recognition model corresponding to the target behavior;

[0119] Based on the multiple vulnerability data sets and the single-vulnerability recognition models trained for each behavior, a vulnerability recognition model is trained through feature cross-learning and ensemble learning; the vulnerability recognition model is used to identify vulnerability attacks.

[0120] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:

[0121] Input the current training samples in the current vulnerability data set into the single-vulnerability recognition model to be trained, and obtain vulnerability operation behaviors and vulnerability operation rewards; put the current training samples, the vulnerability operation behaviors, the vulnerability operation rewards, and the next training sample into a buffer as an experience sample; whenever the number of experience samples put into the buffer exceeds a preset value, optimize the single-vulnerability recognition model to be trained based on the experience samples in the buffer to obtain a trained single-vulnerability recognition model.

[0122] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:

[0123] Randomly sample a first preset number of experience samples from the buffer; based on the first preset number of experience samples, obtain the first preset number of vulnerability operation rewards, and determine the first preset number of target vulnerability values; the vulnerability operation rewards and the target vulnerability values correspond one by one; based on the vulnerability operation rewards and the target vulnerability values, determine a loss function, and optimize the single-vulnerability recognition model to be trained based on the loss function.

[0124] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:

[0125] Input the training samples of each vulnerability data set into the corresponding single-vulnerability recognition model for feature extraction to obtain a vulnerability feature set corresponding to each vulnerability data set; perform combination and cross-analysis on the vulnerability features in different vulnerability feature sets to obtain a cross-feature data set; based on the cross-feature data set, iteratively train the vulnerability recognition model by means of ensemble learning.

[0126] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:

[0127] For the training process of the i-th iteration, based on the cross-feature data set and the vulnerability recognition model of the (i - 1)-th iteration, obtain the negative gradient; i is not less than 2; based on the negative gradient and the cross-feature data set, obtain the decision tree of the i-th iteration; based on the vulnerability recognition model of the i-th iteration and the decision tree of the i-th iteration, obtain the vulnerability recognition model of the i-th iteration.

[0128] In one of the embodiments, the multiple types of behaviors include at least one of the following: user input behavior, API interface call behavior, file operation behavior, network communication behavior, abnormal behavior, permission and access control behavior, and authorization and authentication behavior.

[0129] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, artificial intelligence (AI) processors, etc., without limitation.

[0130] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in the present application.

[0131] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.

Claims

1. A vulnerability identification method, characterized in that: The method comprises: By using the runtime self-protection RASP component, normal behavior samples and vulnerability attack samples corresponding to various behaviors during the application running process are obtained to form multiple vulnerability data sets; each vulnerability data set consists of normal behavior samples and vulnerability attack samples corresponding to one behavior; For each vulnerability data set, determine the target behavior corresponding to the current vulnerability data set, and obtain the single vulnerability identification model to be trained corresponding to the target behavior; based on the current vulnerability data set, train the single vulnerability identification model to be trained by reinforcement learning to obtain the trained single vulnerability identification model corresponding to the target behavior; Based on the multiple vulnerability data sets and the trained single vulnerability identification model corresponding to each behavior, the vulnerability identification model is trained by feature cross-learning and ensemble learning; the vulnerability identification model is used to identify vulnerability attacks.

2. The method according to claim 1, characterized in that The method of training the single vulnerability identification model to be trained by reinforcement learning based on the current vulnerability data set to obtain a trained single vulnerability identification model corresponding to the target behavior includes: Inputting the current training sample in the current vulnerability data set into the single vulnerability identification model to be trained to obtain vulnerability operation behavior and vulnerability operation reward; Put the current training sample, the vulnerability operation behavior, the vulnerability operation reward and the next training sample into a buffer as an experience sample; Whenever the experience samples put into the buffer exceed the preset value, the to-be-trained single vulnerability recognition model is optimized based on the experience samples in the buffer to obtain a trained single vulnerability recognition model.

3. The method according to claim 2, characterized in that The optimizing the single vulnerability identification model to be trained based on the experience samples in the buffer includes: randomly sampling a first preset number of experience samples from the buffer; Based on the first preset number of experience samples, obtaining a first preset number of vulnerability operation rewards, and determining a first preset number of target vulnerability values; the vulnerability operation rewards and the target vulnerability values ​​correspond one to one; Based on the vulnerability operation reward and the target vulnerability value, a loss function is determined, and the single vulnerability identification model to be trained is optimized based on the loss function.

4. The method according to claim 1, characterized in that: The vulnerability identification model is trained by feature cross-learning and ensemble learning, including: Input the training samples of each vulnerability data set into the corresponding single vulnerability identification model to extract features and obtain the vulnerability feature set corresponding to each vulnerability data set; Combine and cross-analyze the vulnerability features in different vulnerability feature sets to obtain a cross-feature dataset; Based on the cross-feature dataset, an ensemble learning approach is used to iteratively train the vulnerability identification model.

5. The method according to claim 4, characterized in that The vulnerability identification model is iteratively trained by using an ensemble learning method, wherein the process of one iterative training includes: For the training process of the i-th iteration, a negative gradient is obtained based on the cross-feature data set and the vulnerability identification model of the i-1-th iteration; the i is not less than 2; Based on the negative gradient and the cross-feature data set, obtaining an i-th decision tree; Based on the vulnerability identification model of the i-th iteration and the i-th decision tree, a vulnerability identification model of the i-th iteration is obtained.

6. The method according to claim 1, characterized in that The multiple behaviors include: at least one of: user input behavior, API interface call behavior, file operation behavior, network communication behavior, abnormal behavior, permission and access control behavior, and authorization and authentication behavior.

7. A vulnerability identification device, characterized in that: The device comprises: An acquisition module is used to acquire normal behavior samples and vulnerability attack samples corresponding to various behaviors during the running of the application program through the runtime application self-protection RASP component, so as to form multiple vulnerability data sets; each vulnerability data set consists of normal behavior samples and vulnerability attack samples corresponding to one behavior; A reinforcement module is used to determine, for each vulnerability data set, a target behavior corresponding to the current vulnerability data set, and obtain a single vulnerability identification model to be trained corresponding to the target behavior; based on the current vulnerability data set, the single vulnerability identification model to be trained is trained by reinforcement learning to obtain a trained single vulnerability identification model corresponding to the target behavior; The integration module is used to train the vulnerability identification model through feature cross-learning and integrated learning based on the multiple vulnerability data sets and the trained single vulnerability identification model corresponding to each behavior; the vulnerability identification model is used to identify vulnerability attacks.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.