Vulnerability detection method and device based on feature engineering

Through feature engineering and Stacking integrated model, the problem of insufficient universality and adaptability of vulnerability detection in the existing technology is solved, and efficient and accurate vulnerability detection is achieved, which is suitable for multi-type information systems.

CN120337223APending Publication Date: 2025-07-18ELECTRIC POWER RESEARCH INSTITUTE OF STATE GRID NINGXIA ELECTRIC POWER COMPANY
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510270915.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-07-18

Smart Images

  • Figure CN120337223A_ABST
    Figure CN120337223A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of information security, in particular to a feature engineering-based vulnerability detection method and device. The method comprises the following steps: acquiring multi-source heterogeneous data of an information system; on the basis of the multi-source heterogeneous data, session-level statistical features are generated through a time sequence sliding window; performing dimension reduction on the session-level statistical features by utilizing mutual information and XGBoost feature importance, and replacing the session-level statistical features with target codes; inputting the session-level statistical features which are replaced by the target codes into a pre-constructed Stacking integration model for training; and outputting a vulnerability detection result through the trained Stacking integration model. According to the method, the session-level statistical features are constructed, and dynamic feature dimension reduction is carried out by utilizing mutual information and XGBoost feature importance double evaluation, so that the feature selection precision is improved. A Stacking integration model is introduced, and an adversarial training mechanism is combined, so that the robustness and generalization ability of the model are enhanced, and the accuracy of a detection result is improved; in addition, the comprehensiveness of vulnerability detection is improved by collecting multi-source heterogeneous data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of information security technology, and particularly relates to a vulnerability detection method and device based on feature engineering. Background Art

[0002] Existing information system vulnerability detection solutions mainly rely on expert knowledge and rule matching. By collecting vulnerability information daily and manually adding it to the expert knowledge base and rule base. By comparing the vulnerabilities of the system to be scanned with the known vulnerability information, analyzing the matching degree and similarity, and then generating a solution, the vulnerability detection effect completely depends on the quantity of the vulnerability knowledge base and the quality of the known vulnerability information.

[0003] In the prior art, since a large amount of vulnerability information needs to be manually collected, sorted out, and normalized, this not only increases the workload, but also leads to the problems of poor generality and adaptability. Especially in the face of a rapidly changing network environment and emerging new vulnerability types, the detection method based on fixed rules is difficult to update in a timely manner. This one-sided perspective makes vulnerability detection often only limited to the matching of known patterns, and cannot comprehensively cover potential security threats, and affects the timeliness and accuracy of vulnerability detection.

[0004] Therefore, there is an urgent need for a vulnerability detection solution based on feature engineering to ensure the comprehensiveness of vulnerability detection and improve the timeliness and accuracy in the vulnerability scanning process. Summary of the Invention

[0005] In view of this, it is necessary to provide a vulnerability detection method and device for information systems based on feature engineering to solve the problems of insufficient comprehensiveness and accuracy in the prior art in vulnerability detection.

[0006] A vulnerability detection method for information systems based on feature engineering includes: Obtaining multi-source heterogeneous data of the information system; Generating session-level statistical features through a time-series sliding window based on the multi-source heterogeneous data; Reducing the dimension of the session-level statistical features by using mutual information and XGBoost feature importance, and converting to obtain session-level statistical features based on target encoding; Inputting the session-level statistical features converted to target encoding into a pre-constructed Stacking ensemble model for training; Outputting a vulnerability detection result through the trained Stacking ensemble model.

[0007] Preferably, the reducing the dimension of the session-level statistical features by using mutual information and XGBoost feature importance, and converting to obtain session-level statistical features based on target encoding includes: Measure the correlation between the session-level statistical features and a preset target variable through mutual information, and filter out the first feature set; Use the XGBoost algorithm to evaluate the importance of the features in the first feature set, and perform weighted fusion according to the evaluation results to obtain a comprehensive score; Filter out a preset proportion of the second feature set according to the comprehensive score to achieve dimensionality reduction; Convert the selected second feature set according to the selected coding strategy.

[0008] Preferably, after filtering out a preset proportion of the second feature set according to the comprehensive score, it further includes: Use a reinforcement learning strategy to adjust the weights to optimize feature selection, and filter out the third feature set.

[0009] Preferably, the Stacking ensemble model includes a base model with LightGBM+BiLSTM in the first layer and a meta-model enhanced by the Attention mechanism in the second layer.

[0010] Preferably, it further includes: when training the Stacking ensemble model, generating adversarial samples through FGSM to train the Stacking ensemble model.

[0011] Preferably, the multi-source heterogeneous data includes network traffic, system logs, API call sequences, and memory status data.

[0012] Preferably, after outputting the vulnerability detection result, it further includes: sending known normal samples and abnormal samples to the information system under test respectively and recording the multi-source heterogeneous data generated by the information system; Analyze the multi-source heterogeneous data generated by the information system to obtain vulnerability features; Feedback the vulnerability features to the Stacking ensemble model to dynamically update the Stacking ensemble model.

[0013] Preferably, when a vulnerability is confirmed, generate corresponding proof-of-concept code.

[0014] An information system vulnerability detection device based on feature engineering, including: A data acquisition module for acquiring multi-source heterogeneous data of the information system; A statistical feature generation module for generating session-level statistical features based on the multi-source heterogeneous data through a time-series sliding window; A feature dimensionality reduction module for reducing the dimensionality of the session-level statistical features by using mutual information and XGBoost feature importance, and converting to obtain session-level statistical features based on target coding; A model training module for inputting the session-level statistical features after being converted into the target encoding into a pre-constructed Stacking ensemble model for training; A prediction module for outputting vulnerability detection results through the trained Stacking ensemble model.

[0015] Preferably, it further includes: a differential testing and feedback module for, after outputting the vulnerability detection results, separately sending known normal samples and abnormal samples to the information system under test and recording the multi-source heterogeneous data generated by the information system under test, analyzing the multi-source heterogeneous data to obtain vulnerability features, and feeding back the vulnerability features to the Stacking ensemble model to dynamically update the Stacking ensemble model.

[0016] Compared with the prior art, the beneficial effects of the present application are as follows: The technical solution provided by the present application constructs session-level statistical features and uses double evaluations of mutual information and XGBoost feature importance to perform dynamic feature dimensionality reduction on the session-level statistical features, improving the accuracy of feature selection. A more complex Stacking ensemble model is introduced and combined with an adversarial training mechanism to enhance the robustness and generalization ability of the model and improve the accuracy of the detection results; in addition, the comprehensiveness of vulnerability detection is increased by collecting multi-source heterogeneous data. Description of the Drawings

[0017] Figure 1 is a schematic flowchart of a method for detecting information system vulnerabilities based on feature engineering provided by an embodiment of the present application.

[0018] Figure 2 is a schematic flowchart of another method for detecting information system vulnerabilities based on feature engineering provided by an embodiment of the present application.

[0019] Figure 3 is a schematic structural diagram of an information system vulnerability detection device 300 based on feature engineering provided by an embodiment of the present application. Detailed Embodiments

[0020] To more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required to be used in the embodiments. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0021] Please refer to Figure 1 - Figure 2 , Figure 1 is a schematic flowchart of a method for detecting information system vulnerabilities based on feature engineering provided by an embodiment of the present application. Figure 2It is a schematic flowchart of information system vulnerability detection based on feature engineering provided by an embodiment of the present application; the information system vulnerability detection method based on feature engineering includes: S101: Obtain multi-source heterogeneous data of the information system.

[0022] The multi-source heterogeneous data may include network traffic, system logs, API call sequences, memory state data, etc., covering four aspects: TCP / IP five-tuple, syslog structured parsing, Hook technology capture, and RASP runtime monitoring. Basic fields such as HTTP request headers, POST parameters, response status codes, and SQL query execution times are collected for the Web system.

[0023] By processing the multi-source heterogeneous data in this application, the comprehensiveness of vulnerability scanning is effectively improved. It avoids the problem that the prior art usually only analyzes surface features such as protocol fields, ignores deep correlation information such as system running status and context semantics, and is difficult to quickly and comprehensively cover all rule points, resulting in a single feature dimension and poor detection effect.

[0024] S102: Process the multi-source heterogeneous data through a time-series sliding window to generate session-level statistical features.

[0025] The time-series sliding window technology is used to generate session-level statistical features, providing a basic input for subsequent feature selection and model training. Among them, the time-series sliding window technology extracts data segments by defining a fixed-size window in the time dimension and gradually moving this window along the time axis. The data within each window can be used to calculate statistical features, which directly reflect certain behavioral characteristics of the system and help identify potential security threats. For example, the proportion or frequency of abnormal requests among all requests from the same IP address can be calculated, including but not limited to multi-source heterogeneous data such as network traffic (such as TCP / IP five-tuple), system logs (such as HTTP request headers, POST parameters, response status codes, etc.), API call sequences, and memory status. Such as the proportion or number of abnormal requests initiated from the same IP address within 5 minutes (judging whether it is abnormal according to predefined criteria), and other statistical metrics that may reflect the system behavior pattern or potential threat, such as average response time, SQL query execution time, etc.

[0026] In this embodiment, using the time-series sliding window technology to generate session-level statistical features can provide feature representations at different time scales through different window size and step settings, improve the representativeness of features, and a feature repository can also be set up to store the selected or determined features, providing high-quality input data for subsequent deep learning models.

[0027] S103: Dimension reduction is performed on the session-level statistical features using mutual information and XGBoost feature importance, and session-level statistical features based on target encoding are obtained through conversion.

[0028] The correlation between the session-level statistical features and a preset target variable is measured using mutual information, and a first feature set is selected. The importance of the features in the first feature set is evaluated using the XGBoost algorithm, and weighted fusion is performed according to the evaluation results to obtain a comprehensive score. A second feature set with a preset ratio is selected based on the comprehensive score to achieve dimension reduction. The selected feature set is converted according to the selected encoding strategy.

[0029] Specifically, first, the association degree between each feature and the target variable is evaluated using mutual information, and a feature set with a relatively high correlation is selected. For each feature, the mutual information value between it and the target variable (such as whether there is a vulnerability) is calculated. Mutual information is a measure of the mutual dependence between two random variables, and it can reveal the association strength between the feature and the target variable even in non-linear cases. Among them, the target variable is a pre-defined label used to identify whether a data point involves a security threat, and this label comes from a professionally annotated data set. These data sets contain historical real cases and can come from internal accumulation, public data, and simulation experiments, etc. Each instance in them has been labeled as whether it contains a vulnerability. Usually, the target variable is usually a binary label (such as 0 or 1), indicating whether a specific data point (such as a network request, a system log entry, etc.) is a potential security threat or vulnerability. Specifically, 1 may represent the existence of a security threat or vulnerability, while 0 represents the non-existence.

[0030] All features are sorted according to the calculated mutual information values, and those features with relatively high mutual information values are selected for the next round of evaluation. This step helps to remove those features that are almost irrelevant or redundant to the target variable and reduce the amount of data for subsequent processing.

[0031] Then, the XGBoost model is applied to calculate the importance scores of these features, and the features are sorted and selected based on this score. This can not only effectively reduce the feature dimension but also ensure that the selected features have high value for the prediction model. This process may include: Training the XGBoost model: Use the feature set selected through mutual information evaluation as input to train an XGBoost classifier. XGBoost is an efficient gradient boosting decision tree algorithm, which is good at handling high-dimensional data and automatically discovering complex patterns. Based on the trained XGBoost model, calculate the importance scores of each feature. XGBoost provides various ways to quantify the importance of features, such as metrics like Gain, Cover, etc. Here, Gain is usually selected as the main basis because it directly reflects the contribution of features to the improvement of model performance. Then combine the results of mutual information evaluation with the XGBoost feature importance scores, and calculate the comprehensive score of each feature through dynamic weighting. The weights can be set according to historical data or experience, aiming to reasonably reflect the evaluation results from two different perspectives in the final score.

[0032] During the final feature selection, combine the evaluation results of mutual information and XGBoost feature importance to determine the final feature set to be retained: Re-sort all features according to the comprehensive score, and set a cumulative score ratio (such as 90%), which means only retain those features whose cumulative contribution of the comprehensive score reaches more than 90% of the total score. Select the top several features according to the set ratio to form the final feature subset. This process not only considers the correlation between features and the target variable but also takes into account their performance in the actual prediction task, ensuring that the selected features are both explanatory and can effectively improve model performance. On this basis, reinforcement learning strategies can also be applied to dynamically adjust the weights according to the historical feature selection effect. For example, give a higher retention probability to features that performed well before, and continuously optimize the feature selection process.

[0033] This embodiment uses mutual information and XGBoost feature importance to reduce the dimension of the session-level statistical features, which can effectively extract the most valuable information from the original data. This process ensures that only the most representative and predictive features will be selected and used subsequently, thereby improving the accuracy and efficiency of the entire vulnerability detection system. It can also significantly reduce the risk of overfitting and enhance the robustness and generalization ability of the system.

[0034] Furthermore, the problem of the curse of dimensionality caused by high-dimensional features can also be addressed through encoding replacement methods, enabling the model to complete training within a reasonable time without overfitting. For example, One-Hot Encoding or Label Encoding can be used. For instance, for categorical features, One-Hot Encoding can be adopted to convert each category into an independent binary feature; for numerical features, standardization or other forms of encoding may be required. This can effectively map high-dimensional features into a lower-dimensional space while retaining as much information as possible from the original features. It not only reduces the complexity of the feature space and computational cost but also avoids overfitting problems caused by excessive dimensions.

[0035] Furthermore, reinforcement learning strategies can also be applied to dynamically adjust weights based on the performance of historical features. For example, features that have performed well previously can be given a higher probability of retention, continuously optimizing the feature selection process. For instance, applying a reinforcement learning strategy to the second feature set and performing selection to obtain the third feature set not only effectively improves the accuracy of feature selection but also further reduces the amount of data to be processed.

[0036] S104: Input the session-level statistical features converted to the target encoding into a pre-constructed Stacking ensemble model for training.

[0037] Stacking is an ensemble learning technique that combines the prediction results of multiple base models (also known as base learners or first-layer models) and uses another model (meta-learner or second-layer model) to learn how to best combine these predictions to generate the final prediction. This method can effectively leverage the advantages of different models to improve the overall performance and stability of the model.

[0038] This application also provides a Stacking integrated model, including a base model using LightGBM + BiLSTM in the first layer and a meta-model strengthened by the Attention mechanism in the second layer. LightGBM is an efficient framework based on the Gradient Boosting Decision Tree (GBDT) algorithm, which is particularly suitable for processing large-scale datasets. It makes predictions by constructing a series of decision trees, and each tree tries to correct the mistakes of the previous tree. BiLSTM (Bidirectional Long Short-Term Memory Network) is suitable for processing sequence data. Compared with the traditional unidirectional LSTM, BiLSTM can read sequence information both forward and backward simultaneously, thus capturing more context information, which is very useful for understanding time series features. The second-layer meta-model uses a model with the Attention mechanism as the meta-learner. The Attention mechanism allows the model to pay more attention to the key parts of the input data when making decisions, which is particularly important for identifying complex patterns and can help improve the resistance to obfuscation attacks.

[0039] The Stacking integrated model proposed in this application can not only make full use of the advantages of different types of models, but also effectively cope with complex and ever-changing information security challenges, providing more accurate and reliable vulnerability detection services.

[0040] Furthermore, the training of the Stacking integrated model can include: First, the screened feature data (feature vectors) are respectively input into the two base models of LightGBM and BiLSTM for training, and each base model will output a prediction result. Further, during the training process, an adversarial training mechanism is introduced to enable the model to learn the patterns of adversarial perturbations, thereby significantly enhancing the defense ability against unknown vulnerabilities, especially obfuscation attacks. Adversarial samples (such as generated by the FGSM method) are introduced and mixed with the original samples for training to enhance the robustness of the model against obfuscation attacks. Finally, the prediction results of all base models are combined as new input features for the meta-model (the model with the Attention mechanism) to further learn how to best integrate these prediction results.

[0041] Furthermore, the method of adjusting the loss function and training strategy in real time can also be adopted during the training process to ensure that the model has good generalization ability while maintaining high accuracy.

[0042] S105: Output the vulnerability detection result through the trained Stacking integrated model.

[0043] Based on the trained Stacking ensemble model, when new data arrives, the multi-source heterogeneous data is first converted into feature vectors. Then, these feature vectors are also fed into the base models to obtain their respective prediction results. Next, these prediction results are passed to the meta-model, which comprehensively considers the outputs of each base model and gives the final vulnerability detection result.

[0044] Furthermore, after outputting the vulnerability detection result, it also includes analysis and verification using differential testing technology. Differential testing technology is mainly applied to compare the behavior differences between the system under test and the virtual sandbox environment to identify potential security threats. Specifically: Send known normal samples and abnormal samples to the information system under test respectively and record the multi-source heterogeneous data generated by the information system; Analyze the multi-source heterogeneous data generated by the information system to obtain vulnerability features; Feed the vulnerability features back to the Stacking ensemble model to dynamically update the Stacking ensemble model.

[0045] To reduce the false positive rate, this application not only relies on the direct comparison of behavior patterns, but also considers additional context information, such as the time, frequency of event occurrence, and the relationship with other events. In addition, a normal behavior baseline can be established based on historical data, and reasonable thresholds can be set to distinguish normal variations from abnormal behaviors. When an anomaly is detected, penetration test cases will be automatically generated for verification, which helps to confirm whether it is a real vulnerability rather than a false positive.

[0046] Furthermore, multiple detection methods can be designed. For input validation type vulnerabilities, adopt the methods of test case generation and differential analysis optimization; For memory security type vulnerabilities, adopt dynamic instrumentation technology and differential trigger conditions; For logical vulnerabilities, adopt multi-role comparison testing and state machine differential technology. Finally, combine the false positive / missed positive results to adjust the test case generation strategy (such as increasing the weight of specific attack patterns). When a vulnerability is confirmed to exist, generate the corresponding proof-of-concept code (POC) to achieve automatic generation of POC.

[0047] The above embodiments of this application use statistics, analysis, and construction of vulnerability features for accurate identification, reduce the cost of vulnerability detection, and save the time of matching vulnerability rules one by one. At the same time, a deep detection model is used for feature adaptation, which can improve the robustness of the model against obfuscation attacks and greatly enhance the detection effect.

[0048] Based on the same inventive concept, this application also provides an information system vulnerability detection device 300 based on feature engineering, as Figure 3 shown in the structural schematic diagram of the information system vulnerability detection device 300 based on feature engineering, including: A data acquisition module 301, configured to acquire multi-source heterogeneous data of the information system; The statistical feature generation module 302 is configured to generate session-level statistical features based on the multi-source heterogeneous data through a time-series sliding window; The feature dimensionality reduction module 303 is configured to reduce the dimensionality of the session-level statistical features by using mutual information and XGBoost feature importance, and convert them into session-level statistical features based on target encoding; The model training module 304 is configured to input the session-level statistical features converted to target encoding into a pre-constructed Stacking ensemble model for training; The prediction module 305 is configured to output a vulnerability detection result through the trained Stacking ensemble model.

[0049] Furthermore, it further includes: a differential testing and feedback module 306, configured to, after outputting the vulnerability detection result, send known normal samples and abnormal samples to the information system under test respectively, record the multi-source heterogeneous data generated by the information system under test, analyze the data to obtain vulnerability features, and feedback the vulnerability features to the Stacking ensemble model to dynamically update the Stacking ensemble model.

[0050] The above data acquisition module 301, statistical feature generation module 302, feature dimensionality reduction module 303, model training module 304, prediction module 305, and differential testing and feedback module 306 are used to execute the methods of the above steps S101-S105 and any optional embodiments thereof, which will not be elaborated in this embodiment.

[0051] In summary, the information system vulnerability detection solution based on feature engineering provided by this application, through multi-source heterogeneous data collection, intelligent feature engineering, complex Stacking ensemble models, adversarial training mechanisms, differential testing techniques, and reinforcement learning strategies, etc., significantly improves the accuracy of vulnerability detection, the robustness and generalization ability of the model, and effectively reduces the false positive rate, and is applicable to the security assessment of various types of information systems such as Web applications, Internet of Things devices, and industrial control systems.

Claims

1. An information system vulnerability detection method based on feature engineering, characterized in that, Including: Obtain multi-source heterogeneous data of the information system; Process the multi-source heterogeneous data through a time-series sliding window to generate session-level statistical features; Reduce the dimension of the session-level statistical features using mutual information and XGBoost feature importance, and transform to obtain session-level statistical features based on target encoding; Input the session-level statistical features transformed into target encoding into a pre-constructed Stacking ensemble model for training; Output the vulnerability detection result through the trained Stacking ensemble model.

2. The information system vulnerability detection method based on feature engineering according to claim 1, wherein The reducing the dimension of the session-level statistical features using mutual information and XGBoost feature importance, and transforming to obtain session-level statistical features based on target encoding includes: Measure the correlation between the session-level statistical features and a preset target variable through mutual information, and filter out the first feature set; Use the XGBoost algorithm to evaluate the importance of the features in the first feature set, and perform weighted fusion according to the evaluation results to obtain a comprehensive score; Filter out a preset proportion of the second feature set according to the comprehensive score to achieve dimensionality reduction; Transform the selected feature set according to the selected encoding strategy.

3. The information system vulnerability detection method based on feature engineering according to claim 2, wherein After filtering out a preset proportion of the second feature set according to the comprehensive score, it further includes: Use a reinforcement learning strategy to adjust the weights to optimize feature selection, and filter out the third feature set.

4. The method for detecting information system vulnerabilities based on feature engineering according to claim 1, wherein The Stacking ensemble model includes a base model with LightGBM+BiLSTM in the first layer and a meta-model enhanced by the Attention mechanism in the second layer.

5. The information system vulnerability detection method based on feature engineering according to claim 4, characterized in that It also includes: When training the Stacking ensemble model, generate adversarial samples through FGSM to train the Stacking ensemble model.

6. The information system vulnerability detection method based on feature engineering according to claim 1, characterized in that, The multi-source heterogeneous data includes network traffic, system logs, API call sequences, and memory status data.

7. The information system vulnerability detection method based on feature engineering according to any one of claims 1-6, characterized in that, After outputting the vulnerability detection result, it further includes: Send known normal samples and abnormal samples to the information system to be tested respectively and record the multi-source heterogeneous data generated by the information system; Analyze the multi-source heterogeneous data generated by the information system to obtain vulnerability features; Feed back the vulnerability features to the Stacking ensemble model to dynamically update the Stacking ensemble model.

8. The information system vulnerability detection method based on feature engineering according to claim 7, wherein When a vulnerability is confirmed, generate corresponding proof-of-concept code.

9. An information system vulnerability detection device based on feature engineering according to any one of claims 1-7, characterized in that, Including: A data acquisition module for obtaining multi-source heterogeneous data of the information system; A statistical feature generation module for generating session-level statistical features based on the multi-source heterogeneous data through a time-series sliding window; A feature dimensionality reduction module for reducing the dimension of the session-level statistical features using mutual information and XGBoost feature importance, and transforming to obtain session-level statistical features based on target encoding; A model training module for inputting the session-level statistical features transformed into target encoding into a pre-constructed Stacking ensemble model for training; A prediction module for outputting the vulnerability detection result through the trained Stacking ensemble model.

10. The method according to claim 9, characterized in that, It also includes: The differential testing and feedback module is used to, after outputting the vulnerability detection results, send known normal samples and abnormal samples to the information system under test respectively and record the multi-source heterogeneous data generated by the information system under test, analyze the data to obtain vulnerability features, and feedback the vulnerability features to the Stacking integration model to dynamically update the Stacking integration model.

Citation Information

Cited By

  • Automatic code generation and optimization method and system based on large model

    CN120850301A

  • A large model-based automated code generation and optimization method and system

    CN120850301B