High-quality data set construction method and device
By building a high-quality dataset that conforms to the characteristics of the Python language, we have solved the problems of insufficient language specificity and poor data integrity of existing general datasets, achieved effective detection and repair of Python language code vulnerabilities, and improved development efficiency and project quality.
Patent Information
- Application Number
- CN202510550469.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-09-19
AI Technical Summary
Existing general vulnerability datasets cannot meet the specific requirements of the Python language. The language is not targeted enough, and the data integrity and accuracy are poor, making it difficult to meet the fine-tuning requirements of large models.
By obtaining the first data set of the vulnerability list, filtering out vulnerability data that conforms to the target language and field, obtaining patch information and patch source code, and using the target classifier function for classification, combining large model interpretation and repair interpretation, a high-quality dataset in SFT format is constructed.
It provides strong support for large models, enables pre-detection and repair of Python language code vulnerabilities, and improves development efficiency and project quality.
Smart Images

Figure CN120671136A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of next-generation Internet application security and cyberspace security technology, and in particular to a method and device for constructing a high-quality data set. Background Art
[0002] With the booming development of artificial intelligence and software development, data quality and adaptability play a key role in technological advancement. Python, with its concise syntax, rich library resources, and powerful functionality, has been widely used in cutting-edge fields such as AI supply chain, large-scale model development and deployment, and deep learning.
[0003] However, with the increasing adoption of Python in various complex systems, code vulnerabilities have become increasingly prominent. These vulnerabilities can not only cause system malfunctions and performance degradation, but can also pose serious security risks, posing threats to data security and system stability. Therefore, high-quality datasets on Python vulnerability fixes are needed to provide strong support for fine-tuning large models. This allows for the pre-detection and repair of Python code vulnerabilities, ensuring development efficiency and project quality.
[0004] In the related art, existing general vulnerability datasets have many defects and are difficult to meet the specific needs of the Python language. Based on this, how to build a high-quality dataset for vulnerability repair in the Python language is an urgent problem to be solved. Summary of the Invention
[0005] The present invention aims to solve one of the technical problems in the related art at least to a certain extent.
[0006] To this end, one purpose of the present invention is to propose a method for constructing a high-quality dataset, which obtains a high-quality dataset for the Python language, thereby providing strong support for fine-tuning of large models and pre-detecting and repairing code vulnerabilities in the Python language, thereby ensuring development efficiency and project quality.
[0007] Another object of the present invention is to provide a device for constructing a high-quality dataset.
[0008] To achieve the above objectives, an embodiment of the present invention provides a method for constructing a high-quality dataset, including:
[0009] Obtaining a first data set of a vulnerability list;
[0010] Filtering the first data set by a filtering function to obtain a second data set of the vulnerability list;
[0011] Determining a first target data set for constructing a high-quality data set based on the second data set;
[0012] Determining, based on the first target data set, a second target data set for constructing a high-quality data set;
[0013] Based on the second target data set, a high-quality data set in a target format is determined.
[0014] The method for constructing a high-quality dataset according to the embodiment of the present invention may also have the following additional technical features:
[0015] Furthermore, the first filtering function includes a target language and a field set; and filtering the first data set by the filtering function to obtain the second data set of the vulnerability list includes:
[0016] Each vulnerability data in the first data set is screened by the screening function, and the vulnerability data that conforms to the target language and exists in all fields of the field set is determined as the second data set of the vulnerability list.
[0017] Furthermore, determining a first target data set for constructing a high-quality data set based on the second data set includes:
[0018] Obtaining patch information for each piece of vulnerability data based on the address field of each piece of vulnerability data in the second data set;
[0019] Based on the patch information of each vulnerability data, obtaining the patch source code of each vulnerability data;
[0020] Each vulnerability data in the second data set, together with the patch information and patch source code of each vulnerability data, is determined as a first target data set for constructing a high-quality data set.
[0021] Furthermore, determining a second target data set for constructing a high-quality data set based on the first target data set includes:
[0022] Classify each vulnerability data in the first target data set by using a target classifier function to obtain a first target vulnerability type for each vulnerability data;
[0023] Interpreting the patch source code of each vulnerability data in the first target data set using the large model to obtain a corresponding patch source code interpretation;
[0024] Determining, based on the first target vulnerability type of each piece of vulnerability data, a target vulnerability repair explanation corresponding to each piece of vulnerability data;
[0025] Each vulnerability data in the first target data set is compared with the patch source code interpretation and the target vulnerability repair interpretation of each vulnerability data to determine a second target data set for constructing a high-quality data set.
[0026] Furthermore, determining a target vulnerability repair explanation corresponding to each piece of vulnerability data based on the first target vulnerability type of each piece of vulnerability data includes:
[0027] Obtain a first vulnerability repair explanation set and a second vulnerability repair explanation set;
[0028] Calculating a similarity between a first vulnerability repair explanation in the first vulnerability repair explanation set and a second vulnerability repair explanation in the second vulnerability repair explanation set, and determining a second target vulnerability type and a third vulnerability repair explanation corresponding to the first vulnerability repair explanation based on the obtained similarity result;
[0029] Processing the first vulnerability repair interpretation based on the second target vulnerability type and the third vulnerability repair interpretation to obtain a fourth vulnerability repair interpretation for each vulnerability type;
[0030] Based on the fourth vulnerability repair interpretation of each vulnerability type and the first target vulnerability type of each piece of vulnerability data, a target vulnerability repair interpretation corresponding to each piece of vulnerability data is determined.
[0031] Furthermore, determining a high-quality data set in a target format based on the second target data set includes:
[0032] Obtaining a prompt word for each vulnerability data in the second target data set;
[0033] interpreting the prompt word, basic information, and patch source code of each vulnerability data in the second target data set to determine the task instruction;
[0034] Determine the target vulnerability repair explanation and patch information for each vulnerability data in the second target data set as the task output;
[0035] A data set consisting of a task instruction and a task output based on each vulnerability data in the second target data set is determined as a high-quality data set in SFT format.
[0036] Furthermore, the method further includes:
[0037] The large model is fine-tuned based on the high-quality dataset in the target format to obtain a fine-tuned large model.
[0038] To achieve the above-mentioned object, another embodiment of the present invention provides a high-quality dataset construction device, the device comprising:
[0039] An acquisition module, configured to acquire a first data set of a vulnerability list;
[0040] a screening module, configured to screen the first data set using a first screening function to obtain a second data set of the vulnerability list;
[0041] A first determining module, configured to determine, based on the second data set, a first target data set for constructing a high-quality data set;
[0042] A second determining module is configured to determine, based on the first target data set, a second target data set for constructing a high-quality data set;
[0043] The third determination module is configured to determine a high-quality data set in a target format based on the second target data set.
[0044] The high-quality dataset construction method and device proposed in the present invention obtain a high-quality dataset for the Python language, which can provide strong support for the fine-tuning of large models, and pre-detect and repair code vulnerabilities in the Python language, ensuring development efficiency and project quality.
[0045] Additional aspects and advantages of the present invention will be set forth in part in the description which follows and, in part, will be obvious from the description which follows, or may be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:
[0047] Figure 1 A flowchart of a method for constructing a high-quality dataset according to an embodiment of the present invention;
[0048] Figure 2 A schematic structural diagram of an apparatus for constructing a high-quality dataset according to an embodiment of the present invention. DETAILED DESCRIPTION
[0049] The following describes embodiments of the present invention in detail, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present invention, and are not to be construed as limiting the present invention.
[0050] In the existing technology, high-quality datasets for the Python language are extremely scarce. While there are common vulnerability datasets in the language community, there are very few for the Python language. Furthermore, existing common datasets have the following deficiencies:
[0051] 1. Insufficient language specificity: General datasets fail to fully consider Python's unique syntax, programming conventions, and its application characteristics in AI and deep learning scenarios. For example, Python's unique dynamic type system, decorators, and generators are not effectively reflected in general datasets, making it difficult for models trained on general datasets to accurately identify and address vulnerabilities in Python code.
[0052] 2. Poor data integrity and accuracy: Common datasets lack comprehensive and accurate documentation of key information designed for Python, such as vulnerabilities in specific library functions and potential risks when integrated with deep learning frameworks. Vulnerabilities in common datasets often have vague descriptions, incomplete fix information, and are often mislabeled. This can lead to bias and misjudgment when developers and researchers use common datasets for model training and vulnerability analysis.
[0053] 3. Difficulty meeting fine-tuning requirements: With the development of large-scale model technology, the demand for high-quality, fine-tunable datasets has become increasingly urgent. However, existing general datasets do not fully consider the requirements of large-scale model fine-tuning in terms of format and content organization. They cannot provide models with high-quality datasets with clear structure and logical coherence. This makes it difficult for models to learn effective vulnerability remediation patterns during fine-tuning, limiting their performance in Python code vulnerability handling tasks.
[0054] Based on the above description, the present invention proposes a method for constructing a high-quality dataset to obtain a high-quality dataset for the Python language, which can provide strong support for fine-tuning of large models and pre-detect and repair code vulnerabilities in the Python language, ensuring development efficiency and project quality.
[0055] The following describes a method and apparatus for constructing a high-quality dataset according to an embodiment of the present invention with reference to the accompanying drawings.
[0056] First, a method for constructing a high-quality dataset according to an embodiment of the present invention will be described with reference to the accompanying drawings.
[0057] Figure 1 The figure is a flowchart of a method for constructing a high-quality dataset according to an embodiment of the present invention.
[0058] like Figure 1 As shown, the method for constructing a high-quality dataset may include the following steps:
[0059] Step S1: Obtain a first data set of a vulnerability list.
[0060] In one embodiment of the present invention, the first data set D of the vulnerability list can be periodically obtained from a third-party website (such as the Mend website). raw , where the first data set D raw Each element d i ∈D raw Represents a vulnerability record, i=1, 2, ..., n, and n is the total number of the first data set.
[0061] In one embodiment of the present invention, a time range for obtaining the first data set can be set. The time range may include the start time and end time for obtaining the first data set, and is expressed in months (for example: the start time is January 2017 and the end time is December 2024).
[0062] Step S2: Filter the first data set using a filtering function to obtain a second data set of a vulnerability list.
[0063] In one embodiment of the present invention, after obtaining the first data set through the above steps, the first data set can be filtered through a filtering function to obtain a second data set in the vulnerability list for subsequent construction of a high-quality data set.
[0064] Furthermore, in one embodiment of the present invention, the method of filtering the first data set using a filtering function to obtain a second data set of vulnerability lists may include: filtering each vulnerability data item in the first data set using the filtering function, and determining vulnerability data items that match the target language and exist in all fields of the field set as the second data set of the vulnerability list. In one embodiment of the present invention, the target language and field set may be set as needed. For example, the target language is Python, and the field set includes {CWE, language type, fix link}.
[0065] In one embodiment of the present invention, the above screening function f filter (d) Access to vulnerability data d in the first data set i Each field and determine the vulnerability data d i Is it the target language and does it exist in all fields in the field set? If d i is the target language and all fields in the field set exist, then f filter (d)=1; otherwise f filter (d)=0.
[0066] And, in one embodiment of the present invention, through the above screening function f filter (d) After filtering the first data set, a second data set D of vulnerability lists can be obtained. python ={d i∈D raw |f filter (d i )=1}, that is, D python It is by D raw The set of all fields in the target language and existing field set is filtered by the filtering function, thus providing basic data for the subsequent construction of a high-quality dataset.
[0067] Step S3: Determine a first target data set for constructing a high-quality data set based on the second data set.
[0068] In one embodiment of the present invention, after obtaining the second data set through the above steps, a first target data set for constructing a high-quality data set can be determined based on the second data set.
[0069] Specifically, in one embodiment of the present invention, the method for determining the first target data set for constructing the high-quality data set based on the second data set may include the following steps:
[0070] Step S31, based on the address field of each vulnerability data in the second data set, obtaining patch information of each vulnerability data;
[0071] Step S32, based on the patch information of each vulnerability data, obtaining the patch source code of each vulnerability data;
[0072] Step S33: Determine each vulnerability data item, the patch information and the patch source code of each vulnerability data item in the second data set as a first target data set for constructing a high-quality data set.
[0073] In one embodiment of the present invention, the patch information P of each vulnerability data can be obtained through a third-party API based on the URL address corresponding to the address field of each vulnerability data in the second data set, wherein each element p in the patch information P is k ∈P,k=1,2,…,m (m is the total number of elements in the patch information), which can include obtaining the patch repair process (detailed addition and deletion code content), source code address, and the developer's explanation of the vulnerability repair.
[0074] And, in one embodiment of the present invention, after obtaining the patch information of each vulnerability data through the above steps, the function h can be obtained through the source code. source (p), according to the source code address provided in the patch information, obtain the corresponding patch source code file from the third-party code library website (such as Github). Among them, for each source address patch element p in the vulnerability data k , get the function h through the source code source (p k) to obtain the patch source code, thereby enriching the context of the patch repair process.
[0075] Furthermore, in one embodiment of the present invention, after obtaining the patch information and patch source code of each vulnerability data through the above steps, each vulnerability data and the patch information and patch source code of each vulnerability data in the second data set can be determined as the first target data set D for constructing the high-quality data set. total .
[0076] Step S4: Based on the first target data set, determine a second target data set for constructing a high-quality data set.
[0077] In one embodiment of the present invention, after obtaining the first target data set through the above steps, a second target data set for constructing a high-quality data set can be determined based on the first target data set.
[0078] Specifically, in one embodiment of the present invention, the method for determining the second target data set for constructing the high-quality data set based on the first target data set may include the following steps:
[0079] Step S41, classifying each vulnerability data in the first target data set by using a target classifier function to obtain a first target vulnerability type for each vulnerability data;
[0080] Step S42, interpreting the patch source code of each vulnerability data in the first target data set using the large model to obtain a corresponding patch source code interpretation;
[0081] Step S43, determining a target vulnerability repair explanation corresponding to each piece of vulnerability data based on the first target vulnerability type of each piece of vulnerability data;
[0082] Step S44 , combining each vulnerability data in the first target data set with the patch source code interpretation and the target vulnerability repair interpretation of each vulnerability data, to determine a second target data set for constructing a high-quality data set.
[0083] In one embodiment of the present invention, the target classifier function may be a CWE classifier function C classify (d t ), and through the target classifier function C classify (d t ) for the first target data set D total Perform CWE classification on each vulnerability data and obtain the first target vulnerability type of each vulnerability data.
[0084] And, in one embodiment of the present invention, the result set after classification by the target classifier function can be C cwe, where each element c i ∈C cwe Represents a CWE classification label, then C cwe ={C classify (d t )|d t ∈D total}, so that the CWE category to which each vulnerability data belongs can be determined, making the first target data set clearer and more orderly in terms of vulnerability types.
[0085] Furthermore, in one embodiment of the present invention, a large model can be used to modularize and functionalize the patch source code for each vulnerability data item in the first target data set. The corresponding modularization and functionalization interpretations can then be used to determine the corresponding patch source code interpretations. This can optimize the code context and allow users to gain a deeper understanding of the original code logic and the causes of the vulnerabilities. In one embodiment of the present invention, the large model can be an existing publicly available large model.
[0086] Furthermore, in one embodiment of the present invention, the method for determining the target vulnerability remediation interpretation corresponding to each piece of vulnerability data based on the first target vulnerability type of each piece of vulnerability data may include the following steps:
[0087] Step 1: Obtain a first vulnerability repair explanation set and a second vulnerability repair explanation set;
[0088] Step 2: Calculate the similarity between the first vulnerability repair explanation in the first vulnerability repair explanation set and the second vulnerability repair explanation in the second vulnerability repair explanation set, and determine the second target vulnerability type and the third vulnerability repair explanation corresponding to the first vulnerability repair explanation based on the obtained similarity result;
[0089] Step 3: Process the first vulnerability repair explanation based on the second target vulnerability type and the third vulnerability repair explanation to obtain a fourth vulnerability repair explanation for each vulnerability type;
[0090] Step 4: Based on the fourth vulnerability repair explanation of each vulnerability type and the first target vulnerability type of each vulnerability data, determine the target vulnerability repair explanation corresponding to each vulnerability data.
[0091] In one embodiment of the present invention, the first vulnerability repair explanation set Commit may be an existing vulnerability repair explanation, which is represented by C={c i |c i ∈Commit}.
[0092] And, in one embodiment of the present invention, a second vulnerability repair explanation set VKB is obtained, wherein the second vulnerability repair explanation set includes a standard vulnerability repair description corresponding to each vulnerability type. Based on this, the second vulnerability repair explanation set VKB={(k j , d j )|(k j is the vulnerability type, d j k j Vulnerability type corresponding to the vulnerability repair standard description}.
[0093] Further, in one embodiment of the present invention, the method of calculating the similarity between the first vulnerability repair interpretation in the first vulnerability repair interpretation set and the second vulnerability repair interpretation in the second vulnerability repair interpretation set, and determining the second target vulnerability type and the third vulnerability repair interpretation corresponding to the first vulnerability repair interpretation based on the obtained similarity result may include: calculating the cosine similarity between the first vulnerability repair interpretation in the first vulnerability repair interpretation set and the second vulnerability repair interpretation in the second vulnerability repair interpretation set, and sorting the obtained similarity results in descending order, determining the second vulnerability repair interpretation ranked first as the third vulnerability repair interpretation of the first vulnerability repair interpretation, and determining the vulnerability type corresponding to the third vulnerability repair interpretation as the second target vulnerability type of the first vulnerability repair interpretation.
[0094] Furthermore, in one embodiment of the present invention, after determining the second target vulnerability type and the third vulnerability repair interpretation corresponding to the first vulnerability repair interpretation through the above steps, the first vulnerability repair interpretation can be processed based on the second target vulnerability type and the third vulnerability repair interpretation to obtain a fourth vulnerability repair interpretation for each vulnerability type. Specifically, in one embodiment of the present invention, the first vulnerability repair interpretation can be manually processed based on the second target vulnerability type and the third vulnerability repair interpretation to obtain the fourth vulnerability repair interpretation for each vulnerability type.
[0095] Furthermore, in one embodiment of the present invention, after obtaining the fourth vulnerability repair explanation for each vulnerability type through the above steps, the fourth vulnerability repair explanation corresponding to the first target vulnerability type can be determined as the target vulnerability repair explanation for the corresponding vulnerability data.
[0096] Step S5: determining a high-quality data set in a target format based on the second target data set.
[0097] In one embodiment of the present invention, after obtaining the second target data set through the above steps, a high-quality data set in a target format can be determined based on the second target data set.
[0098] Specifically, in one embodiment of the present invention, the method for determining a high-quality dataset in a target format based on the second target dataset may include the following steps:
[0099] Step S51, obtaining the prompt word I of each vulnerability data in the second target data set prompt (d i );
[0100] Step S52: the prompt word I of each vulnerability data in the second target data set is prompt (d i ), basic information I vul (d i ) and patch source code explanation I raw-exp (d i ), determined as the task instruction i(d i )={I prompt (d i ),I prompt (d i ),I raw-exp (d i )};
[0101] Step S53: fix the target vulnerability of each vulnerability data in the second target data set. commit (d i ) and patch information R patch (d i ), determined as the task output R(d i )={R patch (d i ),R commit (d i )};
[0102] Step S54: The data set {I(d i ),R(d i )}, determined to be a high-quality dataset in SFT format.
[0103] In one embodiment of the present invention, a guiding prompt can be manually designed based on the vulnerability scenario, such as "You are a code repair expert. Please repair the code according to the following information. The repair solution should not change the code logic." Furthermore, in one embodiment of the present invention, the content of each field of each vulnerability data item in the second target data set can be the basic information of each vulnerability data item.
[0104] Furthermore, in one embodiment of the present invention, after obtaining a high-quality dataset in the target format through the above steps, the large model can be fine-tuned based on the high-quality dataset in the target format to obtain a fine-tuned large model.
[0105] Furthermore, in one embodiment of the present invention, the high-quality dataset constructed through the above steps not only contains rich vulnerability explanations and vulnerability code contexts, but also has detailed repair explanations and repair information, which can effectively support the training and fine-tuning of large models in Python code vulnerability processing, and can be widely used in scenarios such as vulnerability analysis and repair in Python project development and large model training based on the Python language, significantly improving the efficiency and quality of Python language-related development and research work.
[0106] According to the high-quality dataset construction method proposed in an embodiment of the present invention, a high-quality dataset for the Python language is obtained, which can provide strong support for the fine-tuning of large models and pre-detect and repair code vulnerabilities in the Python language, ensuring development efficiency and project quality.
[0107] Next, a high-quality dataset construction apparatus according to an embodiment of the present invention will be described with reference to the accompanying drawings.
[0108] Figure 2 A schematic diagram of the structure of a device for constructing a high-quality dataset according to an embodiment of the present invention.
[0109] like Figure 2 As shown, the high-quality data set construction device 10 includes: an acquisition module 201, a screening module 202, a first determination module 203, a second determination module 204 and a third determination module, wherein:
[0110] An acquisition module 201 is configured to acquire a first data set of a vulnerability list;
[0111] A screening module 202 is configured to screen the first data set using a first screening function to obtain a second data set of a vulnerability list;
[0112] A first determining module 203 is configured to determine a first target data set for constructing a high-quality data set based on the second data set;
[0113] A second determining module 204 is configured to determine, based on the first target data set, a second target data set for constructing a high-quality data set;
[0114] The third determining module 205 is configured to determine a high-quality data set in a target format based on the second target data set.
[0115] Furthermore, the first screening function includes a target language and a field set; the screening module 202 is specifically configured to:
[0116] Each vulnerability data in the first data set is screened by a screening function, and the vulnerability data that conforms to the target language and exists in all fields in the field set is determined as a second data set of the vulnerability list.
[0117] Furthermore, the first determining module 203 is specifically configured to:
[0118] Based on the address field of each vulnerability data in the second data set, obtaining patch information for each vulnerability data;
[0119] Based on the patch information of each vulnerability data, obtain the patch source code of each vulnerability data;
[0120] Each vulnerability data, patch information and patch source code of each vulnerability data in the second data set are determined as a first target data set for constructing a high-quality data set.
[0121] Furthermore, the second determining module 204 is specifically configured to:
[0122] Classify each vulnerability data in the first target data set by using a target classifier function to obtain a first target vulnerability type for each vulnerability data;
[0123] Interpreting the patch source code of each vulnerability data in the first target data set through the large model to obtain a corresponding patch source code interpretation;
[0124] Determine, based on the first target vulnerability type of each piece of vulnerability data, a target vulnerability remediation explanation corresponding to each piece of vulnerability data;
[0125] Each vulnerability data in the first target data set is compared with the patch source code interpretation and the target vulnerability repair interpretation of each vulnerability data to determine a second target data set for constructing a high-quality data set.
[0126] Furthermore, the second determining module 204 is further configured to:
[0127] Obtain a first vulnerability repair explanation set and a second vulnerability repair explanation set;
[0128] Calculating the similarity between the first vulnerability repair explanation in the first vulnerability repair explanation set and the second vulnerability repair explanation in the second vulnerability repair explanation set, and determining the second target vulnerability type and the third vulnerability repair explanation corresponding to the first vulnerability repair explanation based on the obtained similarity result;
[0129] Processing the first vulnerability repair explanation based on the second target vulnerability type and the third vulnerability repair explanation to obtain a fourth vulnerability repair explanation for each vulnerability type;
[0130] Based on the fourth vulnerability repair explanation of each vulnerability type and the first target vulnerability type of each vulnerability data, a target vulnerability repair explanation corresponding to each vulnerability data is determined.
[0131] Furthermore, the third determining module 205 is specifically configured to:
[0132] Obtaining a prompt word for each vulnerability data in the second target data set;
[0133] interpreting the prompt word, basic information, and patch source code of each vulnerability data in the second target data set to determine the task instruction;
[0134] Determine the target vulnerability repair explanation and patch information for each vulnerability data in the second target data set as the task output;
[0135] A data set consisting of task instructions and task outputs based on each vulnerability data in the second target data set is determined as a high-quality data set in SFT format.
[0136] Furthermore, the above device is also used for:
[0137] Fine-tune the large model based on a high-quality dataset in the target format to obtain a fine-tuned large model.
[0138] According to the high-quality dataset construction device proposed in the embodiment of the present invention, a high-quality dataset for the Python language is obtained, which can provide strong support for the fine-tuning of large models, and pre-detect and repair code vulnerabilities in the Python language, ensuring development efficiency and project quality.
[0139] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of the technical features being referred to. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one such feature. In the description of the present invention, "plurality" means at least two, such as two, three, etc., unless otherwise specifically defined.
[0140] In the description of this specification, the reference terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" mean that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification and features of different embodiments or examples without contradiction.
[0141] Although the embodiments of the present invention have been shown and described above, it will be understood that the above embodiments are illustrative and are not to be construed as limitations on the present invention. A person skilled in the art may change, modify, replace and modify the above embodiments within the scope of the present invention.
Claims
1. A method for constructing a high-quality dataset, characterized in that: The method comprises: Obtaining a first data set of a vulnerability list; Filtering the first data set by a filtering function to obtain a second data set of the vulnerability list; Determining a first target data set for constructing a high-quality data set based on the second data set; Determining, based on the first target data set, a second target data set for constructing a high-quality data set; Based on the second target data set, a high-quality data set in a target format is determined.
2. The method according to claim 1, characterized in that The first filtering function includes a target language and a field set; the filtering function is used to filter the first data set to obtain the second data set of the vulnerability list, including: Each vulnerability data in the first data set is screened by the screening function, and the vulnerability data that conforms to the target language and exists in all fields of the field set is determined as the second data set of the vulnerability list.
3. The method according to claim 1, characterized in that The determining, based on the second data set, a first target data set for constructing a high-quality data set comprises: Obtaining patch information for each piece of vulnerability data based on the address field of each piece of vulnerability data in the second data set; Based on the patch information of each vulnerability data, obtaining the patch source code of each vulnerability data; Each vulnerability data in the second data set, together with the patch information and patch source code of each vulnerability data, is determined as a first target data set for constructing a high-quality data set.
4. The method according to claim 3, characterized in that The determining, based on the first target data set, a second target data set for constructing a high-quality data set includes: Classify each vulnerability data in the first target data set by using a target classifier function to obtain a first target vulnerability type for each vulnerability data; Interpreting the patch source code of each vulnerability data in the first target data set using the large model to obtain a corresponding patch source code interpretation; Determining, based on the first target vulnerability type of each piece of vulnerability data, a target vulnerability repair explanation corresponding to each piece of vulnerability data; Each vulnerability data in the first target data set is compared with the patch source code interpretation and the target vulnerability repair interpretation of each vulnerability data to determine a second target data set for constructing a high-quality data set.
5. The method according to claim 4, characterized in that The determining, based on the first target vulnerability type of each piece of vulnerability data, a target vulnerability remediation explanation corresponding to each piece of vulnerability data includes: Obtain a first vulnerability repair explanation set and a second vulnerability repair explanation set; Calculating a similarity between a first vulnerability repair explanation in the first vulnerability repair explanation set and a second vulnerability repair explanation in the second vulnerability repair explanation set, and determining a second target vulnerability type and a third vulnerability repair explanation corresponding to the first vulnerability repair explanation based on the obtained similarity result; Processing the first vulnerability repair interpretation based on the second target vulnerability type and the third vulnerability repair interpretation to obtain a fourth vulnerability repair interpretation for each vulnerability type; Based on the fourth vulnerability repair interpretation of each vulnerability type and the first target vulnerability type of each piece of vulnerability data, a target vulnerability repair interpretation corresponding to each piece of vulnerability data is determined.
6. The method according to claim 4, characterized in that The determining of a high-quality data set in a target format based on the second target data set includes: Obtaining a prompt word for each vulnerability data in the second target data set; interpreting the prompt word, basic information, and patch source code of each vulnerability data in the second target data set to determine them as task instructions; Determine the target vulnerability repair explanation and patch information for each vulnerability data in the second target data set as the task output; A data set consisting of a task instruction and a task output based on each vulnerability data in the second target data set is determined as a high-quality data set in SFT format.
7. The method according to claim 1, characterized in that The method further comprises: The large model is fine-tuned based on the high-quality dataset in the target format to obtain a fine-tuned large model.
8. A high-quality data set construction device, characterized in that: The device comprises: An acquisition module, configured to acquire a first data set of a vulnerability list; a screening module, configured to screen the first data set using a first screening function to obtain a second data set of the vulnerability list; A first determining module, configured to determine, based on the second data set, a first target data set for constructing a high-quality data set; A second determining module is configured to determine, based on the first target data set, a second target data set for constructing a high-quality data set; The third determination module is configured to determine a high-quality data set in a target format based on the second target data set.
9. An electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.