A method, system, and storage medium for generating fixes for open-source components.
By preprocessing the original vulnerability dataset and generating remediation suggestions using a large language model, the problem of inaccurate open-source component remediation solutions in existing technologies is solved, achieving efficient and accurate vulnerability remediation.
Patent Information
- Application Number
- CN202511114344.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-11
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2045-08-11
AI Technical Summary
Existing software component analysis tools cannot provide precise remediation solutions for specific project environments and requirements, resulting in low efficiency and accuracy in fixing vulnerabilities in open-source components.
By acquiring and preprocessing the original vulnerability dataset, a standard vulnerability dataset is generated. Initial remediation suggestions are then generated using a large language model combined with target component information. Based on these initial suggestions, final remediation suggestions are determined, including multi-dimensional recommendations such as version upgrades and parameter adjustments.
It improves the efficiency and accuracy of vulnerability remediation for open-source components, provides targeted remediation solutions, and reduces the risks of dependency conflicts and feature changes.
Smart Images

Figure CN120611388B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of software security technology, and in particular to a method, system, and storage medium for generating repair opinions for open source components. Background Technology
[0002] In recent years, software development has relied heavily on open-source software, employing a wide variety of open-source components, and the open-source component ecosystem has become increasingly large and complex. While the widespread use of open-source components has accelerated software development and reduced costs, it has also introduced security risks. Vulnerabilities in open-source components are becoming increasingly prominent, allowing attackers to exploit them to compromise software systems, leading to data breaches, service endpoints, or other security issues. Therefore, identifying and fixing vulnerabilities in open-source components has become a crucial task in software security management.
[0003] Currently, software component analysis tools are widely used to detect known vulnerabilities in open-source components. Typically based on information from security databases, they can identify security vulnerabilities in open-source components and provide relevant vulnerability information. However, these tools only provide basic vulnerability detection information and suggest upgrading to a specific security version as a fix, without offering precise solutions tailored to the specific project environment and requirements. In practice, directly generating a new version can lead to dependency conflicts, feature changes, or even the introduction of new problems. Therefore, at present, after discovering vulnerabilities in open-source components, it is impossible to provide appropriate and secure fixes based on the current environment, resulting in low efficiency and accuracy in open-source component vulnerability remediation. Summary of the Invention
[0004] To improve the efficiency and accuracy of vulnerability remediation for open-source components, this application provides a method, system, and storage medium for generating remediation suggestions for open-source components.
[0005] Firstly, this embodiment provides a method for generating fix suggestions for open-source components, the method comprising:
[0006] Obtain the original vulnerability dataset, preprocess the original vulnerability dataset to obtain the standard vulnerability dataset, and store the standard vulnerability dataset;
[0007] Obtain the target component information corresponding to the standard vulnerability dataset, and substitute the standard vulnerability dataset and the target component information into a preset large language model to obtain initial remediation suggestion information;
[0008] The final repair recommendations are determined based on the initial repair recommendations.
[0009] In some embodiments, the preprocessing of the original vulnerability dataset to obtain a standard vulnerability dataset includes:
[0010] Vulnerability IDs, vulnerability content, release time, and data sources are selected from the original vulnerability dataset, and then sorted in a preset order to obtain a standard vulnerability dataset.
[0011] In some embodiments, selecting vulnerability content from the original vulnerability dataset includes:
[0012] Use keywords to select vulnerability IDs, release times, and data sources from the original vulnerability dataset to obtain reference vulnerability content;
[0013] Select vulnerability type interpretation content and vulnerability type level remediation suggestion content from the reference vulnerability content, and delete the vulnerability type interpretation content and vulnerability type level remediation suggestion content to obtain the vulnerability content.
[0014] In some embodiments, determining the final repair opinion information based on the initial repair opinion information includes:
[0015] Obtain the number of initial repair opinion sub-information contained in the initial repair opinion information, determine whether the number of information is greater than one, and if not, determine the initial repair opinion information as candidate repair opinion information;
[0016] If it is greater than, determine whether the initial repair opinion information contains disabled function category sub-information. If it does, determine the disabled function category sub-information as candidate repair opinion information.
[0017] If not, determine whether the initial repair opinion information contains version upgrade sub-information and / or parameter adjustment sub-information. If it does, determine the version upgrade sub-information or parameter adjustment sub-information as alternative repair opinion information.
[0018] If not included, any initial repair suggestion sub-information will be selected as a candidate repair suggestion information;
[0019] The alternative repair suggestions are sent to the reviewers to determine whether feedback information corresponding to the alternative repair suggestions is received. If not, the initial repair suggestions are determined as the final repair suggestions.
[0020] If so, the feedback information will be determined as the final repair suggestion information.
[0021] In some embodiments, the method further includes:
[0022] The final repair suggestions, along with the corresponding standard vulnerability dataset and target component information, are combined into a training dataset.
[0023] Obtain the information difference between the final repair opinion information and the alternative repair opinion information, determine whether the information difference meets the preset information difference, and if it does, store the training dataset;
[0024] If the conditions are not met, the preset large language model is adjusted using the stored training dataset, and the training dataset used to adjust the preset large language model is deleted.
[0025] In some embodiments, using the stored training dataset to adjust the preset large language model includes adjusting the learning rate or the number of training epochs in the preset large language model.
[0026] In some embodiments, the method further includes:
[0027] Obtain historical final remediation opinions, corresponding historical standard vulnerability datasets and historical target component information, as well as obtain local large language models;
[0028] The local large language model is trained using the historical final repair opinion information, historical standard vulnerability dataset, and historical target component information to obtain a preset large language model.
[0029] In some embodiments, the method further includes:
[0030] When determining the final repair opinion information based on the initial repair opinion information, the stored training dataset is used to adjust the preset large language model.
[0031] Secondly, this embodiment provides a system for generating fix suggestions for open-source components. The system includes: a data acquisition module, a data processing module, a fix suggestion generation module, and a fix optimization module; wherein...
[0032] The data acquisition module is used to obtain the original vulnerability dataset;
[0033] The data processing module is used to preprocess the original vulnerability dataset to obtain a standard vulnerability dataset, and to store the standard vulnerability dataset.
[0034] The data acquisition module is also used to obtain target component information corresponding to the standard vulnerability dataset;
[0035] The remediation suggestion generation module is used to input the standard vulnerability dataset and the target component information into a preset large language model to obtain initial remediation suggestion information;
[0036] The repair and optimization module is used to determine the final repair opinion information based on the initial repair opinion information.
[0037] Thirdly, this embodiment provides a computer-readable storage medium having a computer program stored thereon that can run on a processor, wherein when the computer program is executed by the processor, it implements a method for generating open-source component repair opinions as described in the first aspect.
[0038] By employing the above method, this application first obtains the original vulnerability dataset, preprocesses it to obtain a standard vulnerability dataset, and stores the standard vulnerability dataset. Then, it obtains the target component information corresponding to the standard vulnerability dataset, substitutes the potential vulnerability dataset and target component information into a pre-defined large language model, and obtains initial remediation suggestions. Finally, it determines the final remediation suggestions based on the initial remediation suggestions. This approach utilizes the natural language understanding capabilities of the large language model, combined with multi-dimensional information such as vulnerability scanning, remediation recommendations, and impact scope, to automatically analyze and generate targeted remediation suggestions. It can not only recommend secure version upgrade solutions but also provide other feasible remediation measures, helping to select the optimal remediation solution in different application scenarios and improving the efficiency and accuracy of vulnerability remediation. Attached Figure Description
[0039] Figure 1 This is a flowchart of a method for generating repair suggestions for open-source components, provided in an embodiment of this application.
[0040] Figure 2 This is a flowchart of a method for determining final repair opinion information based on initial repair opinion information, provided in an embodiment of this application.
[0041] Figure 3 This is a schematic diagram of a system connection for generating repair opinions for open-source components, provided in an embodiment of this application. Detailed Implementation
[0042] To better understand the purpose, technical solutions, and advantages of this application, it has been described and illustrated below with reference to the accompanying drawings and embodiments. However, those skilled in the art should understand that this application can be implemented without these details. It will be apparent to those skilled in the art that various modifications can be made to the embodiments disclosed in this application, and the general principles defined in this application can be applied to other embodiments and application scenarios without departing from the principles and scope of this application. Therefore, this application is not limited to the illustrated embodiments, but is consistent with the broadest scope claimed in this application.
[0043] The embodiments of this application will now be described in further detail with reference to the accompanying drawings.
[0044] Figure 1 This is a block diagram of a method for generating repair suggestions for open-source components, provided in an embodiment of this application. Figure 1 As shown, a method for generating fixes for open-source components includes the following steps:
[0045] Step S100: Obtain the original vulnerability dataset, preprocess the original vulnerability dataset to obtain the standard vulnerability dataset, and store the standard vulnerability dataset.
[0046] This application describes the embodiments from the perspective of the processing end. When the processing end has a task to fix open-source components, it first obtains the original vulnerability dataset. The timing of when the processing end has a task to fix open-source components can be determined based on the actual situation; therefore, this application does not further limit the method for determining when the processing end has a task to fix open-source components.
[0047] The aforementioned original vulnerability dataset specifically includes descriptions of open-source component vulnerabilities, the scope of their impact, security announcements, a list of detected affected components, and the runtime environment information of the tested applications, such as operating system version and middleware version. The original vulnerability dataset can be obtained by viewing publicly available vulnerability databases, including osv.dev and nvd. Alternatively, it can be obtained by viewing security announcements from software vendors; for example, Spring and Apache regularly release vulnerability announcements.
[0048] After obtaining the raw vulnerability dataset, the processing unit immediately preprocesses it to obtain a standard vulnerability dataset, which is then stored. Preprocessing includes data cleaning, standardization, and structured storage. Specifically, preprocessing the raw vulnerability dataset to obtain the standard vulnerability dataset involves selecting vulnerability IDs, vulnerability content, release times, and data sources from the raw dataset, and then sorting these elements in a preset order to obtain the standard vulnerability dataset.
[0049] The process of selecting vulnerability content from the original vulnerability dataset includes: using keywords to select vulnerability IDs, release times, and data sources from the original vulnerability dataset to obtain reference vulnerability content; selecting vulnerability type interpretation content and vulnerability type-level remediation suggestions from the reference vulnerability content, and then deleting the vulnerability type interpretation content and vulnerability type-level remediation suggestions to obtain the final vulnerability content.
[0050] Specifically, each vulnerability ID, release time, and data source corresponds to a unique keyword. Therefore, by examining the content corresponding to these keywords in the original vulnerability dataset and extracting or cutting them out, we can prioritize obtaining the vulnerability ID, release time, and data source. The information in the original vulnerability dataset after removing the vulnerability ID, release time, and data source constitutes the reference vulnerability content. Then, semantic recognition or keyword technology is used to select meaningless information such as vulnerability type interpretations and vulnerability type-level remediation suggestions from the reference vulnerability content, thus identifying the remaining content as the vulnerability content. In this way, by first selecting easily identifiable vulnerability IDs, release times, and data sources from the original vulnerability dataset to obtain the remaining reference vulnerability content, and then deleting meaningless and lengthy vulnerability type interpretations and vulnerability type-level remediation suggestions from the remaining reference vulnerability content, we can quickly and accurately obtain the vulnerability content, thereby completing the preprocessing data cleaning.
[0051] Then, the obtained vulnerability IDs, vulnerability content, release time, and data source are sorted according to a pre-determined order, ensuring that the attribute order of the content in each standard vulnerability dataset is consistent, thus standardizing the standard vulnerability dataset. The vulnerability IDs, vulnerability content, release time, and data source are then stored in a blank one-dimensional array according to the aforementioned pre-determined order, completing the structured storage of the standard vulnerability dataset. Each one-dimensional array includes four cells, each corresponding to specific attribute information. For example, the first cell corresponds to the vulnerability ID attribute, the second cell to the vulnerability content attribute, the third cell to the release time attribute, and the fourth cell to the data source attribute. This preprocessing of the original vulnerability dataset results in a standardized standard vulnerability dataset free of meaningless data, improving data quality and usability. It also provides concise and standardized information for subsequent pre-defined large language models, reducing the processing of meaningless data by the pre-defined large language models, indirectly improving the efficiency and accuracy of determining remediation recommendations for open-source components.
[0052] The obtained standard vulnerability dataset will then be stored so that it can be easily retrieved when needed in the future.
[0053] Step S200: Obtain the target component information corresponding to the standard vulnerability dataset, and substitute the standard vulnerability dataset and target component information into the preset large language model to obtain initial remediation suggestions.
[0054] The target component information mentioned above refers to the information of the component associated with the vulnerability. For example, the component associated with a vulnerability CVE-2020-23811 is Maven's com.xuxueli:xxl-job. The SCA product can obtain information about the component associated with the vulnerability; you can obtain the target component information corresponding to the standard vulnerability dataset by viewing the SCA product.
[0055] The standard vulnerability dataset and target component information obtained above are then sent to a pre-defined large language model. The pre-defined large language model is used to analyze the impact of the vulnerability and the remediation strategy, and to generate initial remediation suggestions.
[0056] Preferably, the preset large language model is obtained by training a local large language model. Obtaining the preset large language model includes: acquiring historical final remediation information and corresponding historical standard vulnerability datasets and historical target component information; acquiring a local large language model; and training the local large language model using the historical final remediation information, historical standard vulnerability datasets, and historical target component information to obtain the preset large language model.
[0057] The aforementioned historical final remediation information refers to the final remediation information obtained by the processing end in the past. The historical standard vulnerability dataset refers to the standard vulnerability dataset obtained by the processing end in the past. The historical target component information refers to the target component information obtained by the processing end in the past. The aforementioned local large language model refers to a general-purpose large language model, obtained through pre-training and fine-tuning based on the architecture of a large language model. Specifically, the local large language model can be one of LlaMA, CodeGeeX2, or Deepseek-Coder. The local large language model is stored on the processing end, and the historical final remediation information, corresponding historical standard vulnerability datasets, historical target component information, and the local large language model can be obtained by viewing the information stored on the processing end. Then, the historical standard vulnerability dataset, historical target component information, and historical final remediation information corresponding to the same detection and remediation operation are used as a set of training model information. Finally, the information from multiple sets of training models is used to train the local large language model to obtain a preset large language model. This makes the preset large language model more adaptable to the processing of component remediation opinions, thereby improving the accuracy of the initial remediation opinions obtained subsequently through the preset large language model.
[0058] Upon receiving a standard vulnerability dataset and corresponding target component information, the pre-defined large language model generates prompts and inputs them into the base model to produce a remediation plan. By using this pre-defined large language model, which better matches the generated remediation suggestions, and then inputting a standard vulnerability dataset containing environmental information and free of meaningless information into it, the pre-defined large language model can generate initial remediation suggestions more quickly and accurately. The types of initial remediation suggestions include, but are not limited to, recommending the best-case security upgrade version, using a configuration remediation plan, parameter adjustment strategies, and plans to disable or shut down affected functions. Initial remediation suggestions must include at least one of these types.
[0059] Step S300: Determine the final repair opinion information based on the initial repair opinion information.
[0060] The initial repair suggestions mentioned above are only the suggestions provided by the preset large language model. The initial repair suggestions may include multiple suggestions and there may be some areas that need adjustment. In order to obtain more accurate repair suggestions, it is necessary to determine the final repair suggestions based on the initial repair suggestions. Figure 2 This is a flowchart illustrating a method for determining final repair opinion information based on initial repair opinion information, provided in an embodiment of this application. For example... Figure 2 As shown, determining the final repair opinion information based on the initial repair opinion information includes the following steps:
[0061] Step S301: Obtain the number of initial repair opinion sub-information contained in the initial repair opinion information, and determine whether the number of information is greater than one. If it is not greater than one, determine the initial repair opinion information as candidate repair opinion information.
[0062] Step S302: If the value is greater than the value, determine whether the initial repair opinion information contains disabled function category sub-information. If it does, determine the disabled function category sub-information as alternative repair opinion information.
[0063] Step S303: If not, determine whether the initial repair opinion information contains version upgrade sub-information and / or parameter adjustment sub-information. If it does, determine the version upgrade sub-information or parameter adjustment sub-information as alternative repair opinion information.
[0064] Step S304: If not included, determine any initial repair opinion sub-information as alternative repair opinion information.
[0065] Step S305: Send the alternative repair suggestions to the reviewers and determine whether feedback information corresponding to the alternative repair suggestions has been received. If not, determine the initial repair suggestions as the final repair suggestions.
[0066] Step S306: If so, confirm the feedback information as the final repair suggestion information.
[0067] Specifically, the pre-defined large language model integrates each completed repair solution into an initial repair opinion sub-information, i.e., a file. Therefore, when multiple repair solutions are obtained, a corresponding number of initial repair opinion sub-information will be generated. The number of information sub-information is determined by the number of files read. When the number of information sub-information is no more than one, it indicates that the initial repair opinion information generated by the pre-defined large language model contains only one repair solution. In this case, this initial repair opinion information is identified as a candidate repair opinion information; that is, this initial repair opinion information represents the solution that needs to be repaired, as determined by the model.
[0068] When the amount of information is greater than one, it indicates that the initial repair suggestions generated by the preset large language model have more than one repair solution. At this time, it is necessary to determine one repair solution from multiple repair solutions as the solution that needs to be repaired by the model.
[0069] Specifically, the type of initial fix suggestion sub-information represented by each initial fix suggestion can be determined through techniques such as viewing or semantic recognition. If the obtained type includes disabled function sub-information, then the initial fix suggestion contains disabled function sub-information. If the obtained type does not include disabled function sub-information, it is further examined whether the obtained type includes version upgrade sub-information and / or parameter adjustment sub-information. If it does, then the initial fix suggestion contains version upgrade sub-information and / or parameter adjustment sub-information. If it does not, then the initial fix suggestion contains neither disabled function sub-information nor version upgrade sub-information and / or parameter adjustment sub-information.
[0070] Since disabled feature information is considered the most secure, followed by version upgrade information and parameter adjustment information, the initial fix suggestion information, if containing disabled feature information, prioritizes the fix corresponding to that disabled feature information as the model-generated fix. If the initial fix suggestion information does not contain disabled feature information, it then checks for version upgrade information and parameter adjustment information, which are considered less secure. If so, the fix corresponding to either is selected as the model-generated fix. If none of these three types of information are present, any one of the sub-information items in the initial fix suggestion information is selected as the model-generated fix. This approach, by first determining the number of items in the initial fix suggestion information and then the types of sub-information, allows for the rapid generation of the model-generated fix, i.e., alternative fix suggestions, indirectly improving the efficiency of generating fix suggestions for open-source components.
[0071] The proposed fixes derived from the model are then sent to the reviewers. This allows the reviewers to determine the final fixes required for the open-source component without directly generating fix suggestions. If adjustments are needed based on the alternative fix suggestions, these adjustments are made accordingly, and the adjusted information is fed back to the processing end as feedback. This feedback information, received by the processing end, represents the final fix for the open-source component.
[0072] If no adjustments are made based on the alternative fix suggestions, the reviewers will not send feedback to the processing end. In this case, the processing end will determine the alternative fix suggestions as the final fixes required for the open-source component. This approach, on the one hand, sends only one fix suggestion to the reviewers, reducing their workload and increasing the efficiency of determining the final fix. On the other hand, by leveraging the accurate and efficient processing based on the pre-defined large language model, and then having the reviewers make the final determination, the accuracy of the fix suggestions can be improved.
[0073] Preferably, after obtaining the final repair opinion information, this application further includes the following steps:
[0074] Step S400: Combine the final repair suggestions, the corresponding standard vulnerability dataset, and the target component information into a training dataset.
[0075] Step S500: Obtain the information difference between the final repair opinion information and the alternative repair opinion information, determine whether the information difference meets the preset information difference, and if it does, store the training dataset.
[0076] Step S600: If satisfied, adjust the preset large language model using the stored training dataset, and delete the training dataset used to adjust the preset large language model.
[0077] The aforementioned information difference refers to the number of differences between the final repair opinion and the alternative repair opinion. The preset information difference refers to a preset range of differences that characterizes the final repair opinion and the alternative repair opinion as having no significant difference. Whether the information difference meets the preset information difference can be determined by comparing the number of differences with the minimum and maximum values of the preset difference range. If the number of differences is less than the minimum value or greater than the maximum value of the preset difference range, the information difference does not meet the preset information difference. If the number of differences is not less than the minimum value and not greater than the maximum value of the preset difference range, the information difference meets the preset information difference.
[0078] If the information difference meets the preset information difference, it indicates that the repair scheme obtained by the preset large language model is highly accurate. At this time, there is no need to adjust the preset large language model. That is, the training dataset is stored so that there is a training set when the preset large language model needs to be adjusted later.
[0079] If the information difference does not meet the preset information difference, it indicates that the accuracy of the repair scheme obtained by the preset large language model is low. At this time, it is necessary to adjust the preset large language model. The preset large language model is adjusted using the already stored training dataset, and the relevant training dataset used to train the preset large language model is deleted. On the one hand, this can relieve the storage pressure on the processing end, and on the other hand, the same training set is not used when training the preset large language model again, which can improve the accuracy of the preset large language model after training.
[0080] This involves using the stored training dataset to adjust the preset large language model, including adjusting the learning rate or the number of training epochs. By adjusting only some parameters without modifying the overall model framework, training time can be shortened, reducing the impact on subsequent determination of fixes for other open-source components.
[0081] Furthermore, when determining the final repair opinions based on the initial repair opinions, the stored training dataset is used to adjust the preset large language model. The processing end does not use the preset large language model when determining the final repair opinions based on the initial repair opinions. Therefore, when there is a need to adjust the preset large language model, this timeframe can be used for adjustment, reducing the impact on subsequent determination of repair opinions for other open-source components.
[0082] Figure 3This is a schematic diagram of a system connection for generating repair suggestions for open-source components, provided in an embodiment of this application. Figure 3 As shown, a system for generating fixes for open-source components includes: a data acquisition module, a data processing module, a fix suggestion generation module, and a fix optimization module.
[0083] The system comprises the following modules: a data acquisition module to obtain the raw vulnerability dataset; a data processing module to preprocess the raw vulnerability dataset to obtain a standard vulnerability dataset, which is then stored; a data acquisition module to obtain the target component information corresponding to the standard vulnerability dataset; a remediation suggestion generation module to input the standard vulnerability dataset and target component information into a pre-defined large language model to obtain initial remediation suggestions; and a remediation optimization module to determine the final remediation suggestions based on the initial suggestions.
[0084] The other functions performed by the aforementioned data acquisition module, data processing module, repair suggestion generation module, and repair optimization module, as well as the technical details of each function, are the same as or similar to the corresponding features in the previously described method for generating repair suggestions for open-source components, and therefore will not be repeated here.
[0085] This application also provides a computer storage medium storing a computer program that, when run on a computer, enables the computer to perform the steps in the previously described method for generating open-source component repair suggestions.
[0086] It should be understood that although the steps in the flowcharts in the accompanying figures are shown sequentially as indicated by the arrows, these steps are not necessarily performed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order requirement for the execution of these steps, and they may be performed in other orders.
[0087] The above description is only a partial embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.
Claims
1. A method for generating fix suggestions for open-source components, characterized in that, The method includes: Obtain the original vulnerability dataset, preprocess the original vulnerability dataset to obtain the standard vulnerability dataset, and store the standard vulnerability dataset; Obtain the target component information corresponding to the standard vulnerability dataset, and substitute the standard vulnerability dataset and the target component information into a preset large language model to obtain initial remediation suggestion information; The final repair recommendations are determined based on the initial repair recommendations. The determination of final repair opinion information based on initial repair opinion information includes: Obtain the number of initial repair opinion sub-information contained in the initial repair opinion information, determine whether the number of information is greater than one, and if not, determine the initial repair opinion information as candidate repair opinion information; If it is greater than, determine whether the initial repair opinion information contains disabled function category sub-information. If it does, determine the disabled function category sub-information as candidate repair opinion information. If not, determine whether the initial repair opinion information contains version upgrade sub-information and / or parameter adjustment sub-information. If it does, determine the version upgrade sub-information or parameter adjustment sub-information as alternative repair opinion information. If not included, any initial repair suggestion sub-information will be selected as a candidate repair suggestion information; The alternative repair suggestions are sent to the reviewers to determine whether feedback information corresponding to the alternative repair suggestions is received. If not, the alternative repair suggestions are determined as the final repair suggestions. If so, the feedback information will be determined as the final repair suggestion information; The method further includes: The final repair suggestions, along with the corresponding standard vulnerability dataset and target component information, are combined into a training dataset. Obtain the information difference between the final repair opinion information and the alternative repair opinion information, determine whether the information difference meets the preset information difference, and if it does, store the training dataset; If the conditions are not met, the preset large language model is adjusted using the already stored training dataset, and the relevant training dataset used to train this preset large language model is deleted.
2. The method according to claim 1, characterized in that, The preprocessing of the original vulnerability dataset to obtain the standard vulnerability dataset includes: Select vulnerability ID, vulnerability content, release time, and data source from the original vulnerability dataset, and sort the vulnerability ID, vulnerability content, release time, and data source in a preset order to obtain a standard vulnerability dataset; The vulnerability content selected from the original vulnerability dataset includes: Use keywords to select vulnerability IDs, release times, and data sources from the original vulnerability dataset to obtain reference vulnerability content; Select vulnerability type interpretation content and vulnerability type level remediation suggestion content from the reference vulnerability content, and delete the vulnerability type interpretation content and vulnerability type level remediation suggestion content to obtain the vulnerability content.
3. The method according to claim 1, characterized in that, The step of adjusting the preset large language model using the already stored training dataset includes adjusting the learning rate or the number of training epochs in the preset large language model.
4. The method according to claim 1, characterized in that, The method further includes: Obtain historical final remediation opinions, corresponding historical standard vulnerability datasets and historical target component information, as well as obtain local large language models; The local large language model is trained using the historical final repair opinion information, historical standard vulnerability dataset, and historical target component information to obtain a preset large language model.
5. The method according to claim 1, characterized in that, The method further includes: When determining the final repair opinion information based on the initial repair opinion information, the stored training dataset is used to adjust the preset large language model.
6. A system for generating fix suggestions for open-source components, characterized in that, The system includes: a data acquisition module, a data processing module, a repair suggestion generation module, and a repair optimization module; wherein... The data acquisition module is used to obtain the original vulnerability dataset; The data processing module is used to preprocess the original vulnerability dataset to obtain a standard vulnerability dataset, and to store the standard vulnerability dataset. The data acquisition module is also used to obtain target component information corresponding to the standard vulnerability dataset; The remediation suggestion generation module is used to input the standard vulnerability dataset and the target component information into a preset large language model to obtain initial remediation suggestion information; The repair and optimization module is used to determine the final repair opinion information based on the initial repair opinion information; The determination of final repair opinion information based on initial repair opinion information includes: Obtain the number of initial repair opinion sub-information contained in the initial repair opinion information, determine whether the number of information is greater than one, and if not, determine the initial repair opinion information as candidate repair opinion information; If it is greater than, determine whether the initial repair opinion information contains disabled function category sub-information. If it does, determine the disabled function category sub-information as candidate repair opinion information. If not, determine whether the initial repair opinion information contains version upgrade sub-information and / or parameter adjustment sub-information. If it does, determine the version upgrade sub-information or parameter adjustment sub-information as alternative repair opinion information. If not included, any initial repair suggestion sub-information will be selected as a candidate repair suggestion information; The alternative repair suggestions are sent to the reviewers to determine whether feedback information corresponding to the alternative repair suggestions is received. If not, the alternative repair suggestions are determined as the final repair suggestions. If so, the feedback information will be determined as the final repair suggestion information; The final repair suggestions, along with the corresponding standard vulnerability dataset and target component information, are combined into a training dataset. Obtain the information difference between the final repair opinion information and the alternative repair opinion information, determine whether the information difference meets the preset information difference, and if it does, store the training dataset; If the conditions are not met, the preset large language model is adjusted using the already stored training dataset, and the relevant training dataset used to train this preset large language model is deleted.
7. A computer-readable storage medium having a computer program stored thereon that can run on a processor, characterized in that, When the computer program is executed by the processor, it implements a method for generating repair suggestions for open-source components as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Vulnerability description and repair suggestion generation method based on big language model reasoning and retrieval enhancement
CN120145397A