Security patch classification method and system based on pseudo label learning
Through a method based on pseudo-label learning, key semantic information of security patches is extracted and high-quality pseudo-labels are generated, which solves the problems of insufficient labeling data and low data quality in the prior art, and significantly improves the accuracy and robustness of security patch classification.
Patent Information
- Application Number
- CN202510694231.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-28
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-05-28
AI Technical Summary
The existing deep learning methods are not accurate in the classification of security patch vulnerability types, mainly due to insufficient labeling data and low data quality.
Using a pseudo-label learning method, key semantic information extraction is performed on security-related patch data sets and labelless patch data sets, high-quality pseudo-labels are generated, and high-quality pseudo-label samples are screened out using consensus algorithms, and the training set is extended to improve the learning effect and accuracy of the classification model.
It significantly improves the learning effect and accuracy of the security patch classification model, improves the accuracy and robustness of classification, reduces the risk of overfitting, and realizes the automation of the patch classification process, improving processing speed and efficiency.
Smart Images

Figure CN120217211A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of vulnerability type classification, and more specifically, to a security patch classification method and system based on pseudo-label learning. Background Art
[0002] In the current software engineering field, open source software has been widely used in the software supply chain of commercial / non-commercial products. At the same time, open source software vulnerabilities have also been widely spread, and downstream open source software users need to promptly discover and fix vulnerabilities in upstream open source software. In the vulnerability repair process, vulnerability type information is very important, which can help developers understand the root cause of the vulnerability, possible impact, and the type of mitigation measures to be deployed. Therefore, it is very important to classify security patches according to vulnerability types.
[0003] Researchers have proposed many methods to classify security patches. Among them, deep learning-based methods have attracted widespread attention because they can automatically extract features from code and identify complex patterns. These methods have made some progress in the classification of security patch vulnerability types, especially in reducing manual intervention and improving classification efficiency, showing their strong potential. However, although deep learning-based technologies have achieved certain results, their classification performance still has certain limitations, mainly due to the lack of high-quality training data, but there are many challenges in obtaining current labeled datasets.
[0004] On the one hand, manual code review often requires deep expert knowledge, which is not only time-consuming but also susceptible to human bias. On the other hand, although using existing static analysis tools to generate annotated datasets can accelerate the annotation process to a certain extent, the generated data has a high false positive rate, which further affects the quality of the data and the training effect of the model.
[0005] Therefore, in the current deep learning methods, how to use unlabeled data for effective learning, overcome the bottleneck of insufficient labeled data, and improve the accuracy and robustness of classification is a key problem that needs to be solved urgently in current technology. Summary of the invention
[0006] In view of the defects of the prior art, the purpose of this application is to provide a security patch classification method and system based on pseudo-label learning, aiming to solve the problem of low accuracy in the current classification of security patch vulnerability types.
[0007] To achieve the above objectives, in a first aspect, the present application provides a security patch classification method based on pseudo-label learning, comprising: Get security patches to be triaged; Extracting features of the security patch to obtain key semantic features; Input the key semantic features into the trained security patch model to obtain the classification result of the security patch; Among them, the security patch model is obtained by extracting key semantic information from a security-related patch dataset and an unlabeled patch dataset, and performing pseudo-label learning based on the key semantic information.
[0008] Optionally, the method for obtaining the security patch model includes: Exclude the repair content unrelated to security, extract the key variables related to the vulnerability, and use code slicing technology combined with data flow analysis to determine the vulnerability trigger point, and integrate the key variables and the vulnerability trigger point into key semantic information; Train an initial teacher model using the training set of labeled security patches, generate pseudo-labels for the unlabeled security patches using the initial teacher model, screen out high-quality pseudo-label samples through a consensus algorithm combined with the key semantic information and add them to the training set of labeled security patches to iteratively train the initial teacher model until a trained security patch classification model is obtained.
[0009] Optionally, the process of extracting the key semantic information specifically includes: Analyze the code modifications of the patches in the security-related patch dataset, label whether the code modification blocks are related to security repairs, and obtain a candidate sample set; Use the similar patch retrieval method to select similar samples of the input patches of the security-related patch dataset from the candidate sample set, and construct a few-shot prompt based on the similar samples; Use a large language model to exclude the code modification blocks unrelated to the vulnerability in the security patch according to the few-shot prompt; Analyze and summarize the repair features of the code modifications of the security patches, design extraction rules according to the repair type and the repair line type, and use the extraction rules to extract the key variables in the code modifications; Formulate corresponding vulnerability trigger point rules according to different vulnerability types, perform data flow analysis based on the key variables and the vulnerability trigger point rules until the analysis reaches a preset depth, and obtain code data that conforms to the vulnerability trigger point rules; Determine the vulnerability-related code based on the code data, key variables, and the data flow between the code data and the key variables as the key semantic information.
[0010] Optionally, the training process of the security patch classification model includes: Determine that the initial teacher model includes a code modification classification model and a text description classification model; Train the code modification classification model using the security patch code modifications labeled with vulnerability type labels, and train the text description classification model using the security patch text descriptions labeled with vulnerability type labels; Use the code modification classification model and the text description classification model to perform code modification prediction and text description prediction for the unlabeled security patches respectively, and generate pseudo-labels; Adopt a consensus algorithm based on code modification and text description and combine key semantic information to confirm the generated pseudo-labels, screen out high-quality pseudo-label samples, add them to the training set of labeled security patches, and iteratively train the initial teacher model until a trained security patch classification model is obtained.
[0011] Optionally, the process of data flow analysis includes: Construct a data flow graph of the code; Locate the target node of the key variable in the data flow graph; Traverse from the target node to match the vulnerability trigger point. In the case of encountering a function call statement, adjust to the program dependence graph of the called function and continue traversing until the preset depth is reached to obtain code data that conforms to the vulnerability trigger point rule.
[0012] Optionally, the calculation method of the loss function in the training process of the initial teacher model includes: Calculate the first loss between the predicted value of the labeled security patch and the actual label; Calculate the second loss between the predicted value of the pseudo-label sample and the pseudo-label; Take the weighted sum of the first loss and the second loss as the total loss of model training.
[0013] Optionally, the process of obtaining pseudo-labels includes: Perform probability sorting on the prediction results of the text description and code modification of the security patch, and select the top k categories respectively; If there are the same categories among the top k categories, select the category with the largest average probability as the pseudo-label; if there are no same categories among the top k categories, discard the current sample.
[0014] In a second aspect, the present application also provides a security patch classification system based on pseudo-label learning, including: An acquisition module for acquiring security patches to be classified; A feature extraction module for extracting features from the security patch to obtain key semantic features; A classification module for inputting the key semantic features into the trained security patch model to obtain the classification result of the security patch; Wherein, the security patch model is obtained by extracting key semantic information from a security-related patch data set and an unlabeled patch data set and performing pseudo-label learning according to the key semantic information.
[0015] Optionally, the security patch model includes: A critical semantic information extraction module, configured to exclude repair content unrelated to security, extract critical variables related to vulnerabilities, and use code slicing technology combined with data flow analysis to determine vulnerability trigger points, and integrate them into critical semantic information according to the critical variables and vulnerability trigger points; A pseudo-label learning module, configured to train an initial teacher model using a training set of labeled security patches, generate pseudo-labels for unlabeled security patches using the initial teacher model, filter out high-quality pseudo-label samples through a consensus algorithm in combination with the critical semantic information and add them to the training set of labeled security patches to iteratively train the initial teacher model until a trained security patch classification model is obtained.
[0016] In a third aspect, the present application provides an electronic device, including: at least one memory for storing a program; at least one processor for executing the program stored in the memory, and when the program stored in the memory is executed, the processor is configured to execute the method described in the first aspect or any one of the possible implementation manners of the first aspect.
[0017] In a fourth aspect, the present application provides a computer-readable storage medium storing a computer program, and when the computer program runs on a processor, it causes the processor to execute the method described in the first aspect or any one of the possible implementation manners of the first aspect.
[0018] In a fifth aspect, the present application provides a computer program product, and when the computer program product runs on a processor, it causes the processor to execute the method described in the first aspect or any one of the possible implementation manners of the first aspect.
[0019] It can be understood that the beneficial effects of the above second aspect to fifth aspect can refer to the relevant descriptions in the first aspect above, and will not be elaborated here.
[0020] Generally speaking, compared with the prior art through the above technical solutions conceived by the present application, the following beneficial effects are obtained: (1) By using the labeled security patch data set and the unlabeled patch data set to generate pseudo-labels, the present application can expand the training set when the data volume is limited, thereby significantly improving the learning effect and accuracy of the classification model. Combined with the learning of critical semantic information, the model can better understand and classify the characteristics of patches, especially the variables highly related to vulnerability repair. The present application improves the classification effectiveness of the model for security patches in real scenarios through a refined training method, and improves the classification accuracy and robustness of security patches.
[0021] (2) This application screens high-quality pseudo-label samples through a consensus algorithm. By extracting key features from patches from different sources and combining similarity retrieval with few-shot prompting, the diversity of training data is increased, enabling the model to handle more complex patch forms and repair types. While effectively enhancing the generalization ability of the model, the risk of overfitting can be reduced.
[0022] (3) This application automatically extracts key semantic features and uses code slicing technology and data flow analysis to identify vulnerability trigger points, realizing the automation of the patch classification process. The automated extraction and confirmation process not only reduces the need for manual intervention but also significantly improves the processing speed and efficiency. Compared with the traditional manual annotation process, the rapid iteration of model training and pseudo-label generation makes the entire patch classification workflow more efficient, capable of adapting to a rapidly changing security environment in a short time and improving the overall response ability.
[0023] (4) This application extracts key semantic information highly relevant to the vulnerability type as model input samples, thus significantly improving the sample quality and making it easier for the model to learn the features of the vulnerability type. Compared with traditional pseudo-label learning methods, this application uses a consensus algorithm based on patch text and code modification to screen pseudo-label samples, ensuring both the accuracy of the pseudo-labels and a sufficient number of pseudo-label samples. The pseudo-label learning module of this application can be combined with existing deep learning-based security patch classification methods to further enhance their classification effects. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 is one of the schematic flowcharts of the security patch classification method based on pseudo-label learning provided by an embodiment of this application; Figure 2 is a schematic diagram of key variable extraction in an embodiment of this application; Figure 3 is the second of the schematic flowcharts of the security patch classification method based on pseudo-label learning provided by an embodiment of this application; Figure 4 is one of the schematic structural diagrams of the security patch classification device based on pseudo-label learning provided by an embodiment of this application; Figure 5 is the second of the schematic structural diagrams of the security patch classification device based on pseudo-label learning provided by an embodiment of this application; Figure 6 is the schematic structural diagram of the electronic device provided by an embodiment of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0025] To make the objectives, technical solutions, and advantages of this application clearer and more understandable, the following further details this application in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not used to limit this application.
[0026] As used herein, the term "and / or" describes the association relationship of associated objects and indicates that there can be three relationships. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. The symbol " / " herein indicates that the associated objects are in an "or" relationship. For example, A / B represents A or B.
[0027] The terms "first", "second", etc. in the description and claims of this application are used to distinguish different objects, rather than to describe a specific order of the objects. For example, the first response message and the second response message are used to distinguish different response messages, rather than to describe the specific order of the response messages.
[0028] In the embodiments of this application, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary" or "for example" is intended to present relevant concepts in a specific manner.
[0029] In the description of the embodiments of this application, unless otherwise specified, the meaning of "a plurality of" refers to two or more. For example, a plurality of processing units refers to two or more processing units, and a plurality of elements refers to two or more elements.
[0030] The following describes the embodiments of this application in conjunction with the accompanying drawings in the embodiments of this application.
[0031] To achieve the above objectives, in a first aspect, this application provides a security patch classification method based on pseudo-label learning, including: S101. Obtain the security patch to be classified; S102. Extract features from the security patch to obtain key semantic features; S103. Input the key semantic features into the trained security patch model to obtain the classification result of the security patch; wherein, the security patch model is obtained by extracting key semantic information from a security-related patch dataset and an unlabeled patch dataset and performing pseudo-label learning based on the key semantic information.
[0032] Optionally, the method for obtaining the security patch model includes: Exclude security - unrelated repair content, extract key variables related to vulnerabilities, and use code slicing technology combined with data - flow analysis to determine the vulnerability trigger points, and integrate them into key semantic information according to the key variables and vulnerability trigger points; Train an initial teacher model using a training set of labeled security patches, generate pseudo - labels for unlabeled security patches using the initial teacher model, and screen out high - quality pseudo - label samples through a consensus algorithm combined with the key semantic information and add them to the training set of labeled security patches to iteratively train the initial teacher model until a trained security patch classification model is obtained.
[0033] Specifically, first obtain the security patch to be classified through S101 to ensure subsequent analysis and classification for a specific patch.
[0034] Then, through S102, extract features from these security patches with the goal of obtaining key semantic features. In this step, the algorithm analyzes the code modifications and text descriptions of the patches to refine important information related to vulnerabilities for the next classification operation.
[0035] Finally, through S103, input these extracted key semantic features into a pre - trained security patch model, and the model can output the classification result of the security patch, realizing an automated patch classification process.
[0036] It should be noted that the security patch model is obtained by extracting key semantic information from a security - related patch dataset and an unlabeled patch dataset. In the process of constructing this model, first exclude security - unrelated repair content and focus on extracting key variables related to vulnerabilities. To this end, the model combines code slicing technology and data - flow analysis to ensure accurate identification of vulnerability trigger points. After integrating these key variables and the determined vulnerability trigger points, complete key semantic information is formed for subsequent training and classification.
[0037] In the process of obtaining the security patch model, train an initial teacher model using a training set of labeled security patches. Then, use the initial model to generate pseudo - labels for unlabeled security patches, which provides an effective label prediction mechanism for unlabeled data. Combine the extracted key semantic information and use a consensus algorithm to screen out high - quality pseudo - label samples. By adding these high - quality pseudo - label samples to the training set of labeled security patches, the initial teacher model can be iteratively trained, gradually optimized, and finally a well - performing trained security patch classification model is obtained.
[0038] Embodiments of this application combine labeled and unlabeled data, optimize the model training process through the mechanism of pseudo-label learning, and improve the accuracy and efficiency of patch classification. Through an adaptive learning method, the model can not only process the labeled data but also effectively utilize a large amount of unlabeled data to more comprehensively identify and classify security patches.
[0039] Optionally, the process of extracting the key semantic information specifically includes: Analyze the code modifications of the patches in the security-related patch dataset, label whether the code modification blocks are related to security fixes, and obtain a candidate sample set; Use the similar patch retrieval method to select similar samples of the input patches of the security-related patch dataset from the candidate sample set, and construct a few-shot prompt according to the similar samples; Use a large language model to exclude the code modification blocks unrelated to vulnerabilities in the security patches according to the few-shot prompt; Analyze and summarize the patching features of the code modifications of the security patches, design extraction rules according to the patching type and patching line type, and use the extraction rules to extract the key variables in the code modifications; Formulate corresponding vulnerability trigger point rules according to different vulnerability types, perform data flow analysis based on the key variables and the vulnerability trigger point rules until the preset depth is reached, and obtain the code data that conforms to the vulnerability trigger point rules; Determine the vulnerability-related code as the key semantic information according to the code data, key variables, and the data flow between the code data and the key variables.
[0040] Specifically, the extraction of the key semantic information includes the following steps: (1.1) Exclude security-irrelevant fixes: Manually analyze the code modifications in the patch, label whether the code modification blocks (git-hunks) are related to security fixes, and form a candidate sample set. Use the similar patch retrieval technique to select several samples most similar to the input patch from the candidate sample set, construct a few-shot prompt, and input the few-shot prompt into a large language model to use the large language model (LLM) to exclude the code modification blocks unrelated to vulnerabilities in the security patches; The code for the specific implementation of constructing the prompt is: System prompt: Instruction: <instruction> Examples:<retrieval system examples> User prompt: Commit message:<commit message> Git-hunk: <git-hunk>。
[0041] Among them, the System prompt (i.e., the system instruction) consists of two parts: Instruction and Examples. The Instruction describes the task requirements, while the Examples part provides examples for the model to refer to. The User prompt (i.e., the user instruction) consists of two parts: Commit message and Git-hunk. The Commit message represents the text description of the submission, while the Git-hunk is the code block of the submitted modification.
[0042] (1.2) Key variable extraction: Summarize the patching features through manual analysis of the code modifications of security patches. Design corresponding rules for different patching types (variable assignment, function call, variable or function definition, control statement, etc.) and patching line types (only addition / deletion or modification) to extract the variables that have a relatively large correlation with the vulnerability in the code modification, that is, the key variables; The specific rules for key variable extraction are as Figure 2 shown. The specific implementation method is to use Tree-sitter to parse the diff file. For each git-hunk, if there are only plus lines or only minus lines, it is of the only addition / deletion type, otherwise it is of the modification type. Then extract the code before and after the modification, and extract the key variables according to the Figure 2 rules.
[0043] (1.3) Vulnerability-related code extraction: Develop corresponding vulnerability trigger point rules according to different vulnerability types, and then perform data flow analysis starting from the key variables, stopping at a predetermined depth. All the code that meets the rules during the process is recorded. Finally, these suspected vulnerability trigger point codes, key variables, and their data flows are used as vulnerability-related code, that is, key semantic information.
[0044] Optionally, the training process of the security patch classification model includes: Determine that the initial teacher model includes a code modification classification model and a text description classification model; Train the code modification classification model with the security patch code modifications labeled with vulnerability type tags, and train the text description classification model with the security patch text descriptions labeled with vulnerability type tags; Use the code modification classification model and the text description classification model respectively to predict the code modification and text description of the unlabeled security patch, and generate pseudo-labels; Use a consensus algorithm based on code modification and text description, combined with key semantic information, to confirm the generated pseudo-labels, screen out high-quality pseudo-label samples, add them to the training set with labeled security patches, and iteratively train the initial teacher model until a trained security patch classification model is obtained.
[0045] The training steps of the security patch classification model are as follows: (2.1) Model training: The initial teacher model includes a code modification classification model and a text description classification model; In this embodiment, a code modification classification model is trained using security patch code modifications labeled with vulnerability type tags, and a text description classification model is trained using security patch text descriptions labeled with vulnerability type tags; The specific implementation of model training is to calculate the loss between the sample class label and the model prediction value for labeled data . For unlabeled data, use the pseudo-label of the sample as the label and calculate the loss between the predicted class of the sample and the pseudo-label . The final loss function is:
[0046] where is a hyperparameter. During the training process, perform backpropagation of the gradient on to complete the update of the model.
[0047] (2.2) Pseudo-label generation: Use the code modification classification model and the text description classification model respectively to predict the code modification and text description of unlabeled security patches and generate pseudo-labels. Specifically, two score lists will be obtained. Assume that there are in total for the final vulnerability type categories, then each score list contains values, and each value represents the confidence score that the sample belongs to that vulnerability type.
[0048] (2.3) Pseudo-label confirmation: Use a consensus algorithm based on code modification and text description to confirm the generated pseudo-labels, so as to screen out high-quality pseudo-label samples and add them to the training set for the next round of training.
[0049] Optionally, the process of data flow analysis includes: Construct a data flow graph of the code; Locate the target node of the key variable in the data flow graph; Start traversing from the target node to match the vulnerability trigger point. In the case of encountering a function call statement, adjust to the program dependence graph of the called function and continue traversing until the preset depth is reached to obtain code data that conforms to the vulnerability trigger point rule.
[0050] Specifically, the specific implementation of extracting vulnerability-related code in this embodiment is as follows: First, use Joern to construct a data flow graph for the code, then locate the nodes where key variables are located in the data flow graph, start traversing from this node to match the vulnerability trigger points, and when encountering a function call statement during the process, jump to the program dependence graph of the called function and continue traversing, stop when reaching the set depth, and record all the code involved in the traversal process. For the cases of only addition and modification, extract the vulnerability-related code in the modified version, and for the case of only deletion, extract it in the version before modification.
[0051] Optionally, the process of obtaining pseudo-labels includes: Perform probability sorting on the text description of the security patch and the prediction results of code modification, and select the top k categories respectively; If there are the same categories among the top k categories, select the category with the maximum average probability as the pseudo-label; if there are no same categories among the top k categories, discard the current sample.
[0052] Specifically, the implementation manner of the consensus algorithm in this embodiment is that for each unlabeled sample After passing through the code modification classification model And the text description classification model After prediction, two score lists are obtained respectively. After sorting according to the score size, take the first Most likely categories, assuming after The set of predicted categories is:
[0053] After The set of predicted categories is:
[0054] If , then the current sample will not be added to the training set, otherwise, for Calculate the average probability for each category in, and select the category with the maximum average probability as the pseudo-label of this sample, and at the same time add this sample to the training set.
[0055] Repeat the above process until the model effect no longer improves, and select the model with the best final effect for output.
[0056] Refer to Figure 3 , Figure 3 Is the complete process schematic diagram of the embodiment of this application, including: Use the security-related patch data set and the unlabeled patch data set as inputs; Perform key semantic information extraction and pseudo-label learning, and output the key semantic information set and the security patch classification model respectively; Key semantic information extraction includes security - irrelevant repair exclusion, key variable extraction, and vulnerability - related code extraction; pseudo - label learning includes: pseudo - label generation, pseudo - label confirmation, and model training.
[0057] Referring Figure 4 , the present application also provides a security patch classification system based on pseudo - label learning, including: An acquisition module 410, configured to acquire security patches to be classified; A feature extraction module 420, configured to extract features from the security patches to obtain key semantic features; A classification module 430, configured to input the key semantic features into a trained security patch model to obtain the classification result of the security patches; Wherein, the security patch model is obtained by performing key semantic information extraction on a security - related patch dataset and an unlabeled patch dataset, and performing pseudo - label learning based on the key semantic information.
[0058] Referring Figure 5 , the security patch model includes: A key semantic information extraction module, configured to exclude repair content irrelevant to security, extract key variables related to vulnerabilities, and use code slicing technology combined with data - flow analysis to determine vulnerability trigger points, and integrate the key variables and vulnerability trigger points into key semantic information; A pseudo - label learning module, configured to train an initial teacher model using a training set of labeled security patches, generate pseudo - labels for unlabeled security patches using the initial teacher model, screen out high - quality pseudo - label samples through a consensus algorithm combined with the key semantic information and add them to the training set of labeled security patches to iteratively train the initial teacher model until a trained security patch classification model is obtained.
[0059] Optionally, the key semantic information extraction module specifically includes: A security - irrelevant repair exclusion sub - module, configured to analyze the code modifications of the patches in the security - related patch dataset, label whether the code modification blocks are related to security repairs, and obtain a candidate sample set; Use the similar patch retrieval method to select similar samples of the input patches in the security - related patch dataset from the candidate sample set, and construct few - shot prompts according to the similar samples; Use a large - language model to exclude code modification blocks irrelevant to vulnerabilities in the security patches according to the few - shot prompts; A key variable extraction sub - module, configured to analyze and summarize the patching features of the code modifications of the security patches, design extraction rules according to the patching type and patching line type, and use the extraction rules to extract key variables in the code modifications; A vulnerability-related code extraction sub-module, which is used to formulate corresponding vulnerability trigger point rules according to different vulnerability types, perform data flow analysis based on the key variables and the vulnerability trigger point rules until the analysis reaches a preset depth, and obtain code data that conforms to the vulnerability trigger point rules; Determine the vulnerability-related code as the key semantic information according to the code data, key variables, and the data flow between the code data and the key variables.
[0060] Optionally, the pseudo-label learning module specifically includes: A model training sub-module, which is used to determine that the initial teacher model includes a code modification classification model and a text description classification model; Use the security patch code modifications labeled with vulnerability type tags to train the code modification classification model, and use the security patch text descriptions labeled with vulnerability type tags to train the text description classification model; A pseudo-label generation sub-module, which is used to use the code modification classification model and the text description classification model respectively to predict the code modification and text description for the unlabeled security patch, and generate pseudo-labels; A pseudo-label confirmation sub-module, which is used to adopt a consensus algorithm based on code modification and text description and combine key semantic information to confirm the generated pseudo-labels, screen out high-quality pseudo-label samples, add them to the training set of labeled security patches, and perform iterative training on the initial teacher model until a trained security patch classification model is obtained.
[0061] Optionally, the process of data flow analysis includes: Construct a data flow graph of the code; Locate the target node of the key variable in the data flow graph; Start traversing from the target node to match the vulnerability trigger point. In the case of encountering a function call statement, adjust to the program dependence graph of the called function and continue traversing until the preset depth is reached, and obtain code data that conforms to the vulnerability trigger point rules.
[0062] Optionally, the calculation method of the loss function in the initial teacher model training process includes: Calculate the first loss between the predicted value and the actual label of the labeled security patch; Calculate the second loss between the predicted value and the pseudo-label of the pseudo-label sample; Take the weighted sum of the first loss and the second loss as the total loss of model training.
[0063] Optionally, the process of obtaining pseudo-labels includes: Perform probability sorting on the prediction results of the text description and code modification of the security patch, and select the top k categories respectively; If there are identical categories among the top k categories, select the category with the maximum average probability as the pseudo label; if there are no identical categories among the top k categories, discard the current sample.
[0064] It can be understood that for the detailed function implementation of each of the above units / modules, reference can be made to the introduction in the foregoing method embodiments, and details are not described herein again.
[0065] It should be understood that the above device is used to execute the method in the above embodiments. For the corresponding program modules in the device, their implementation principles and technical effects are similar to those described in the above method. The working process of the device can refer to the corresponding process in the above method, and details are not described herein again.
[0066] Referring to Figure 6 , based on the method in the above embodiments, an embodiment of the present application provides an electronic device, which may include: a processor (Processor) 610, a communication interface (Communications Interface) 620, a memory (Memory) 630, and a communication bus 640. Among them, the processor 610, the communication interface 620, and the memory 630 communicate with each other through the communication bus 640. The processor 610 can call the logical instructions in the memory 630 to execute the method in the above embodiments.
[0067] In addition, when the logical instructions in the above memory 630 are implemented in the form of software functional units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application.
[0068] Based on the method in the above embodiments, an embodiment of the present application provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program runs on a processor, the processor is caused to execute the method in the above embodiments.
[0069] Based on the method in the above embodiments, an embodiment of the present application provides a computer program product. When the computer program product runs on a processor, the processor is caused to execute the method in the above embodiments.
[0070] It can be understood that the processor in the embodiments of the present application may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. The general-purpose processor may be a microprocessor or any conventional processor.
[0071] The method steps in the embodiments of the present application may be implemented in a hardware manner or by a processor executing software instructions. The software instructions may be composed of corresponding software modules, and the software modules may be stored in a random access memory (RAM), flash memory, read-only memory (ROM), programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), registers, hard disks, removable hard disks, CD-ROMs, or any other form of storage medium well known in the art. An exemplary storage medium is coupled to the processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium may also be a component of the processor. The processor and the storage medium may be located in an ASIC.
[0072] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted through the computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)), etc.
[0073] It can be understood that the various numerical numbers involved in the embodiments of the present application are only for the convenience of description and are not used to limit the scope of the embodiments of the present application.
[0074] It is easy for those skilled in the art to understand that the above are only the preferred embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present application shall be included in the protection scope of the present application. < / instruction>
Claims
1. A security patch classification method based on pseudo-label learning, characterized in that, Including: Obtain security patches to be classified; Extract features from the security patches to obtain key semantic features; Input the key semantic features into a trained security patch model to obtain the classification result of the security patches; Among them, the security patch model is obtained by extracting key semantic information from a security-related patch dataset and an unlabeled patch dataset, and performing pseudo-label learning based on the key semantic information.
2. The security patch classification method based on pseudo-label learning according to claim 1, wherein The method for obtaining the security patch model includes: Exclude repair content unrelated to security, extract key variables related to vulnerabilities, and use code slicing technology combined with data flow analysis to determine vulnerability trigger points, and integrate the key variables and vulnerability trigger points into key semantic information; Train an initial teacher model using a training set of labeled security patches, use the initial teacher model to generate pseudo-labels for unlabeled security patches, and use a consensus algorithm combined with the key semantic information to screen out high-quality pseudo-label samples and add them to the training set of labeled security patches to iteratively train the initial teacher model until a trained security patch classification model is obtained.
3. The method for classifying security patches based on pseudo-label learning according to claim 2, wherein The process of extracting the key semantic information specifically includes: Analyze the code modifications of the patches in the security-related patch dataset, label whether the code modification blocks are related to security repairs, and obtain a candidate sample set; Use the similar patch retrieval method to select similar samples of the input patches in the security-related patch dataset from the candidate sample set, and construct a few-shot prompt according to the similar samples; Use a large language model to exclude code modification blocks unrelated to vulnerabilities in the security patches according to the few-shot prompt; Analyze and summarize the repair features of the code modifications of the security patches, design extraction rules according to the repair type and repair line type, and use the extraction rules to extract key variables in the code modifications; Formulate corresponding vulnerability trigger point rules according to different vulnerability types, perform data flow analysis based on the key variables and vulnerability trigger point rules until the analysis reaches a preset depth, and obtain code data that conforms to the vulnerability trigger point rules; Determine the vulnerability-related code as the key semantic information according to the code data, key variables, and the data flow between the code data and key variables.
4. The method for classifying security patches based on pseudo-label learning according to claim 2, wherein The training process of the security patch classification model includes: Determine that the initial teacher model includes a code modification classification model and a text description classification model; Train the code modification classification model using the code modifications of security patches labeled with vulnerability type labels, and train the text description classification model using the text descriptions of security patches labeled with vulnerability type labels; Use the code modification classification model and the text description classification model respectively to perform code modification prediction and text description prediction for unlabeled security patches, and generate pseudo-labels; Adopt a consensus algorithm based on code modification and text description and combine key semantic information to confirm the generated pseudo-labels, screen out high-quality pseudo-label samples, add them to the training set of labeled security patches, and iteratively train the initial teacher model until a trained security patch classification model is obtained.
5. The security patch classification method based on pseudo-label learning according to claim 3, wherein The process of the data flow analysis includes: Construct a data flow graph of the code; Locate the target nodes of the key variables in the data flow graph; Starting from the target node, traverse to match the vulnerability trigger point. In the case of encountering a function call statement, adjust to the program dependence graph of the called function and continue traversing until the preset depth is reached to obtain code data that conforms to the vulnerability trigger point rule.
6. The security patch classification method based on pseudo-label learning according to claim 2, characterized in that, The calculation method of the loss function in the initial teacher model training process includes: Calculating a first loss between the predicted value and the actual label of the labeled security patch; Calculating a second loss between the predicted value and the pseudo-label of the pseudo-label sample; Taking the weighted sum of the first loss and the second loss as the total loss for model training.
7. The method for classifying security patches based on pseudo-label learning according to claim 2, wherein The process of obtaining the pseudo-label includes: Performing probability ranking on the text description of the security patch and the prediction result of the code modification, and respectively selecting the top k categories; If there are the same categories among the top k categories, select the category with the maximum average probability as the pseudo-label; if there are no same categories among the top k categories, discard the current sample.
8. A security patch classification system based on pseudo-label learning, characterized in that, Including: An acquisition module for acquiring the security patch to be classified; A feature extraction module for extracting features from the security patch to obtain key semantic features; A classification module for inputting the key semantic features into the trained security patch model to obtain the classification result of the security patch; Wherein, the security patch model is obtained by extracting key semantic information from a security-related patch data set and an unlabeled patch data set, and performing pseudo-label learning according to the key semantic information.
9. The security patch classification system based on pseudo-label learning according to claim 8, wherein The security patch model includes: A key semantic information extraction module for excluding repair content unrelated to security, extracting key variables related to vulnerabilities, and using code slicing technology combined with data flow analysis to determine the vulnerability trigger point, and integrating the key variables and the vulnerability trigger point into key semantic information; A pseudo-label learning module for training an initial teacher model using a training set of labeled security patches, generating pseudo-labels for unlabeled security patches using the initial teacher model, screening out high-quality pseudo-label samples through a consensus algorithm in combination with the key semantic information and adding them to the training set of labeled security patches to iteratively train the initial teacher model until a trained security patch classification model is obtained.
10. An electronic device, characterized in that, Including: At least one memory for storing a computer program; At least one processor for executing the program stored in the memory. When the program stored in the memory is executed, the processor is used to execute the method according to any one of claims 1-7.
Citation Information
Patent Citations
Vulnerability classification method and device, equipment and medium
CN114117445A
Semantic segmentation network training method and device based on block patch learning and medium
CN114399638A
Lumbar vertebra disease semi-supervised classification method and system based on semantic comparison and fusion of uncertain perception
CN120047736A
Systems, methods, and apparatuses for implementing transferable visual words by exploiting the semantics of anatomical patterns for self-supervised learning
US20220309811A1
Automatic event graph construction method and device for multi-source vulnerability information
US20230035121A1
Cited By
Semi-supervised submission level vulnerability classification method based on reinforcement learning enhancement
CN122388765A
Systems, methods, and computer programs for determining a vulnerability of a network node
US12563079B2