Version controllable data set automatic evolution method based on large language model

Through the automation method of large language models and multi-agent systems, the problem of difficulty in real-time update of Python library versions and API changes is solved, efficient and accurate data set maintenance is achieved, and the evaluation ability of LLMs in dynamic development environments is improved.

CN120335855AActive Publication Date: 2025-07-18NANJING UNIV OF POSTS & TELECOMM

Patent Information

Application Number
CN202510417539.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-07-18
Estimated Expiration
2045-04-03

AI Technical Summary

Technical Problem

The existing code generation data set maintenance methods rely on manual labor, are inefficient and difficult to reflect changes in Python library versions and APIs in real time, resulting in deviations from the actual application effect, and it is impossible to accurately evaluate the performance of large language models (LLMs) in dynamic development environments.

Method used

The automation method based on large language model and multi-agent system is adopted, and the data crawling, parsing, annotation and task processing modules are used to realize the automatic monitoring and data set update of third-party library version updates in the Python ecosystem. The AexPy tool and multi-agent work together to build and maintain a controlled version of the data set.

Benefits of technology

Improve the efficiency and accuracy of data set maintenance, ensure that the data set is synchronized with the latest library versions and API changes, improve the adaptability evaluation capabilities of LLMs in dynamic development environments, and reduce the need for manual intervention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120335855A_ABST
    Figure CN120335855A_ABST
Patent Text Reader

Abstract

The invention discloses a version controllable data set automatic evolution method based on a large language model, which mainly comprises a data crawling module, a data analysis module, a data annotation module and a task form processing module, and reduces the demand of manual intervention by introducing an automatic updating mechanism. Specifically, the multi-agent system can monitor version update of a third-party library in Python ecology and update a data set according to the latest library version. The automatic updating mechanism not only improves the maintenance efficiency of the data set, but also ensures the timeliness of the data set, and provides reliable technical support for the adaptability evaluation of the LLMs in a dynamic development environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of code generation, and particularly to an automatic evolution method for a version-controlled dataset based on a large language model. Background Art

[0002] In recent years, large language models represented by ChatGPT and CodeLLaMA have made remarkable breakthroughs in the field of code generation. Their performance evaluation mainly relies on static benchmark test sets such as HumanEval and MBPP. However, software development inherently has the characteristic of dynamic evolution, which is specifically manifested as the frequent version updates of software libraries and the continuous iteration of API interfaces.

[0003] Traditional evaluation methods have obvious limitations: existing datasets fail to accurately simulate the dynamic changes of library version dependencies in real development scenarios, resulting in the inability to effectively test the adaptation ability of LLMs to specific version libraries. VersiCode is a benchmark dataset specifically designed to evaluate the code generation ability of LLMs under version control, covering 300 Python libraries and more than 2,000 versions, which provides a comprehensive and realistic test platform for version-aware code generation of LLMs.

[0004] In addition, LLMs have shown reasoning and planning abilities comparable to those of humans, which has also led to the rapid development of LLM-based agents (also known as LLM agents). Such agents can understand and generate human-like instructions, supporting complex interactions and decision-making. To further improve the capabilities of agents, researchers have proposed LLM-based multi-agent systems, which solve complex problems through the cooperation of multiple specialized agents. Compared with single agents, multi-agent systems can simulate cooperation and competition in complex environments, showing stronger task-solving capabilities and being widely applied in fields such as software development, robotic systems, social simulation, policy making, and games.

[0005] With the continuous evolution of Python software libraries and APIs, it is crucial to ensure that version-controlled datasets such as VersiCode can timely reflect the latest library version updates and API changes for accurately evaluating the performance of LLMs in a dynamic development environment. Traditional dataset maintenance methods often rely on manual updates, which are not only time-consuming and laborious but also prone to missing important updates, affecting the accuracy of model performance evaluation. To solve this problem, adopting an automated evolution method based on large language models and multi-agent systems has become a cutting-edge solution.

[0006] The maintenance method of traditional code generation datasets relies on manual writing, which is inefficient and difficult to control in terms of quality, affecting the accuracy and timeliness of evaluation. In addition, as the first benchmark dataset specifically for the dynamic changes in Python library versions, VersiCode faces challenges such as low efficiency of manual updates and difficulty in reflecting the changes of the latest library versions in real time. Therefore, it is urgent to expand and enrich VersiCode through automated means to improve its dynamic update ability.

[0007] Existing evaluation benchmarks have certain limitations: (1) The construction of the dataset highly depends on manual work. Experts are required to write tasks, test cases, and standard answers, which is time-consuming and laborious and has poor scalability. The quality of the dataset is likely to vary due to the subjective preferences of the designers. (2) The version information of the code is ignored, and the version differences of Python third-party libraries (such as API changes in pandas or numpy) are not considered. It is difficult to evaluate the performance of the model in generating code in a specific library version environment and to reflect the version control issues in actual engineering scenarios. These problems lead to a deviation between the evaluation results and the actual application effects, and may overestimate the practicality of the model in the real environment.

[0008] The present invention aims to solve the following key problems: (1) How to achieve the automated maintenance of the VersiCode version-controlled dataset, reduce manual intervention, and ensure that the dataset is synchronized with the latest and more popular library versions and API changes; (2) How to further expand and enrich the API evolution types of its Python libraries on the basis of the existing VersiCode benchmark dataset to more comprehensively reflect the changes in library version APIs and accurately evaluate the code generation ability of LLMs in a dynamic development environment; (3) How to improve the efficiency and flexibility of the VersiCode dataset maintenance through the collaboration of multiple agents to support the adaptive evaluation of LLMs in a rapidly changing software development environment. Summary of the Invention

[0009] To solve the above problems, the present invention discloses a method for automatically evolving a version-controlled dataset based on a large language model. Through the collaborative work of multiple agents, it realizes the automated monitoring of the version updates of third-party libraries in the Python ecosystem and the update of the dataset, effectively improving the maintenance efficiency and accuracy of the dataset and providing reliable technical support for the evaluation of the dynamic environment adaptation ability of LLMs.

[0010] The present invention adopts the following technical solutions to achieve the above object:

[0011] A method for automatically evolving a version-controlled dataset based on a large language model mainly includes a data crawling module, a data parsing module, a data annotation module, and a task form processing module. The specific content of each module is as follows:

[0012] (1) Data Crawling Module

[0013] This module collects metadata of Python library versions from multiple data sources (such as GitHub, PyPI website, etc.) through automated means, ensuring the integrity and real-time nature of the collected data, and conducts preliminary screening on the data. Denote the process steps of this module as S1, and the specific steps are as follows:

[0014] In step S101, use a crawler script to retrieve the legal Python libraries published on the PyPI website and obtain a list of library names;

[0015] In step S102, based on the list of library names obtained in step S101, retrieve the projects of the Python libraries in the list on the Github website to obtain meta-information, including the number of stars of the project, release time, etc.;

[0016] In step S103, based on the meta-information obtained in step S102, filter out the Python libraries whose number of stars of the project meets the threshold conditions as the objects to be crawled for constructing the dataset. The threshold is set manually. For example, the number of stars of the project is greater than 10000 or the number of increased stars in the recent month is greater than 1000. After filtering, obtain the list of Python libraries to be crawled;

[0017] In step S104, based on the list of libraries to be crawled obtained in step S103, classify the libraries to be crawled in combination with the Python library version situation included in the local VersiCode dataset: if the local library already exists in the list of libraries to be crawled, then crawl the new source code packages of this library except the locally existing versions; if not, then crawl all historical source code packages of this library within the specified time range;

[0018] In step S105, based on the meta-information of the two types of Python libraries obtained in step S104, including library name, version number, release time, and the requires_python field (the Python version environment required to install this library version), after sorting, hand it over to the data parsing module for subsequent data processing.

[0019] (2) Data Parsing Module

[0020] This module conducts structured processing and preliminary cleaning on the crawled data through automated means, extracts key fields and standardizes the data format, ensuring the consistency and accuracy of the data, and providing reliable input for subsequent annotation and processing. Denote the process steps of this module as S2, and the specific steps are as follows:

[0021] In step S201, the AexPy tool proposed in the literature "AexPy: Detecting API Breaking Changes in Python Packages" is optimized. AexPy is a tool for detecting API breaking changes in third-party Python libraries. Its core functions involve (1) crawling Python library meta-information, (2) extracting internal library APIs, (3) detecting API changes between library versions, and (4) generating API change detection reports. The four modules corresponding to the above functions are "preprocess", "extract", "diff", and "report" respectively. Among them, the "preprocess" module is used to obtain the meta-information of Python libraries, etc., and the "extract" module mainly implements the dynamic extraction of API information based on three built-in Python libraries, namely importlib, pkgutil, and inspect. The modifications to AexPy are as follows:

[0022] (1) The original AexPy could only run in the Python 3.12 environment. Modify the internal code of AexPy to make it support multiple versions of the Python environment (such as Python 3.6 to 3.12);

[0023] (2) When the original AexPy extracts the API of the library version, pip installs the.whl file compatible with the latest Python version supported by the library by default (such as 3.12), but this may lead to incompatible dependency versions. For example, after installing a certain library, pip may select the latest version of the dependent library compatible with Python 3.12, resulting in a version conflict between the installed library and the dependent library. By parsing the requires_python field of the library version meta-information, obtain the specified Python version and pass it into the AexPy tool, thereby effectively alleviating the problem of incompatible library and dependency versions;

[0024] (3) When some Python libraries are installed with pip, requirements.txt may not contain all their required dependencies, resulting in the failure of the Extract module to extract the API due to missing dependencies. By modifying the internal code of AexPy and using the predefined library-dependency mapping relationship, automatically identify and install the missing dependencies, thereby effectively reducing the extraction failure problem caused by missing dependencies.

[0025] In step S202, based on the Python library metadata finally obtained in step S1, parse the requires_python field corresponding to each library version to obtain the Python version environment required to install this library version. The parsing method is based on regular expressions. For example, if the value of requires_python is ">=3.6.1" or ">3.7", and the version number pattern is denoted as major.minor.patch, write a regular expression to match this type of pattern and capture the version number in the match. If it is ">=", the parsed Python version number is itself. If it is ">", if the patch part is included, return major.minor.(patch + 1), and if the patch is not included, return major.(minor + 1).

[0026] In step S203, based on the Python version to be installed obtained in step S202, use the preprocess function module of the AexPy tool to crawl the source code package of the library version corresponding to the Python version for subsequent extraction of APIs in the library;

[0027] In step S204, based on the source code package of the library version crawled in step S203, use the extract function module of the AexPy tool to perform API extraction. The working process of the AexPy tool is as follows: First, create a virtual environment according to the runnable Python version of the current library version, and install the previously crawled source code package into this virtual environment. Then, based on the three Python built-in libraries importlib, pkgutil, and inspect, extract the APIs in the Python third-party library. Among them, importlib is used to dynamically load modules and support importing the target library and its sub-modules as needed. pkgutil is responsible for traversing the package structure and recursively obtaining all modules and sub-modules in the library. Inspect is used to extract information such as functions and classes in the module and obtain their metadata (such as parameter signatures and docstrings). At this point, the APIs and related information in the current library version can be obtained, denoted as the internal knowledge base of the library version, and stored in JSON format.

[0028] In step S205, based on the internal knowledge base of the library version obtained in step S204, according to the Python API evolution patterns proposed in the literature "How Do Python Framework APIs Evolve? An Exploratory Study", using the written script tool, find the APIs that have evolved between two consecutive library versions (i.e., the internal knowledge bases of two library versions). Each or several functions in the script correspond to an API evolution rule mentioned in the above literature. So far, the API evolution relationships that occur between library versions can be obtained. Each pair of evolution relationships consists of the APIs corresponding to the old and new versions and their related information. These evolution relationships can be recorded as the API evolution knowledge graph between library versions and stored in JSON format.

[0029] (3) Data annotation module

[0030] This module plans to automatically annotate the parsed data through an LLM-based agent and preset rules, adjust the annotation strategy according to different task scenarios, and the annotation module ensures the accuracy of data annotation, laying a foundation for subsequent task processing and evaluation. Denote the process steps of this module as S3, and its specific steps are as follows:

[0031] In step S301, based on the evolved APIs and their related information obtained in step S20, construct code usage examples for these APIs as the ground truth in the dataset, which is used to measure the correctness of the answers generated by the large model. Use an LLM-based agent to extract code usage examples from the docstring of the API metadata. The extraction method is to write the possible code example patterns in the docstring into the prompt as context to inform the LLM agent, so as to guide it to identify and extract valid code snippets.

[0032] In step S302, based on the successfully extracted code snippets in step S301, use an LLM-based agent to generate function descriptions for them. The role of the function description is to be written into the prompt as part of the dataset, together with related information such as library versions, as context to inform the LLM agent, and combined with methods such as few-shot learning, so as to guide it to generate code corresponding to the version and function, that is, this is a code generation task with controllable versions.

[0033] In step S303, based on the data obtained from the processing in step S302, an automated script is used to perform data cleaning according to heuristic rules to obtain the successfully processed data, which is recorded as metadata. The form of the metadata is a multi-tuple, which at least includes three core elements: library version, function description, and code snippet, and may also include other relevant information, such as API name, etc.

[0034] (4) Task form processing module

[0035] This module plans to convert the labeled data into the form required for specific tasks through a large language model-based agent and automated script tools, making it more suitable for the evaluation process, ensuring that the data can meet the LLM performance evaluation scenario, and improving the quality of the final dataset. The process steps of this module are denoted as S4, and the specific steps are as follows:

[0036] In step S401, based on the metadata obtained in step S3, it is further processed into two task forms: code completion and code migration included in the VersiCode dataset. For the code completion task, for the evolution types that only involve a single library version, such as the addition and deletion of APIs, etc. Use the script to use the AST (Abstract Syntax Tree) package to locate the version-sensitive parts of each code snippet (i.e., the API names marked in the metadata), and replace them with <mask>The tag, as the content to be inferred, enables the LLM to predict the masked part based on the context. For the code migration task, for the evolution types involving cross-versions, such as the renaming and relocation of APIs within a Python library. Based on the API evolution knowledge graph between library versions, each set of data involving APIs of two versions is used as a test instance. The Python library name, source version, source version API name, and source version usage example code snippet are informed to the LLM, allowing the LLM to infer the code snippet containing the evolved API in the target version, that is, to achieve code migration from one version to another. The version before migration is called the source version, and the version after migration is called the target version.

[0037] The beneficial effects of the present invention:

[0038] (1) Automatically update the dataset, reduce manpower, and dynamically reflect library evolution;

[0039] The present invention reduces the need for manual intervention by introducing an automated update mechanism. Specifically, the multi-agent system can monitor the version updates of third-party libraries in the Python ecosystem and update the dataset in real time according to the latest library versions. This automated update mechanism not only improves the efficiency of dataset maintenance but also ensures the timeliness of the dataset, providing reliable technical support for the adaptability evaluation of LLMs in a dynamic development environment.

[0040] (2) Expand the API evolution patterns of the VersiCode dataset;

[0041] Although VersiCode is the first benchmark dataset specifically for the dynamic changes of library versions, the Python API evolution patterns it contains still have limitations. The present invention expands the coverage of the evolution patterns of VersiCode by introducing the use of the AexPy tool and referring to the richer API evolution patterns summarized in previous research work, enabling it to more comprehensively reflect the actual situation of Python library version evolution. The present invention realizes the automation of VersiCode dataset update through the cooperation of the multi-agent system, monitors the version updates of third-party libraries in the Python ecosystem in real time, and ensures that the dataset is always consistent with the latest library versions and API changes.

[0042] (3) The multi-agent updates the dataset, improving efficiency and accuracy;

[0043] Through the collaboration of a multi-agent system, the present invention realizes the efficient maintenance of a dataset in a pipeline manner. Each agent is specifically designed for a particular task and executes step by step in a predefined order. Multiple agents work together sequentially to complete the construction and update of the dataset. For example, one agent constructs code usage examples for the library version APIs obtained through automated scripts, then another agent generates corresponding function descriptions for these code snippets to enrich the context information informed to LLMs in the dataset, and finally a third agent generates test cases for each data instance. This multi-agent pipeline collaboration mechanism not only improves the efficiency of dataset maintenance but also enhances the accuracy and integrity of the dataset. In addition, the multi-agent system exhibits high flexibility and scalability, capable of dynamically adjusting the pipeline strategy according to the characteristics of different libraries and their APIs to ensure that the dataset remains consistent with the latest technological developments.

[0044] In summary, through the extension of the API evolution mode of the VersiCode dataset, the introduction of an automated mechanism, and the multi-agent pipeline-style collaboration, the present invention significantly improves the practicality, timeliness, and accuracy of the dataset, providing comprehensive and reliable technical support for the code generation ability evaluation of LLMs in a dynamic development environment. At the same time, the present invention greatly reduces the need for manual intervention, providing technical guarantees for efficient and long-term maintenance. Brief Description of the Drawings

[0045] Figure 1 Flowchart of the data crawling module of the present invention;

[0046] Figure 2 Flowchart of the data parsing module of the present invention;

[0047] Figure 3 Flowchart of the data annotation module of the present invention;

[0048] Figure 4 Flowchart of the task form processing module of the present invention. Detailed Embodiments

[0049] The following further clarifies the present invention in conjunction with the drawings and specific embodiments. It should be understood that the following specific embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. It should be noted that the terms "front", "rear", "left", "right", "upper", and "lower" used in the following description refer to the directions in the drawings, and the terms "inner" and "outer" refer to the directions towards or away from the geometric center of a specific component respectively.

[0050] A method for automatically evolving a version-controlled dataset based on a large language model, and its specific implementation cases are as follows:

[0051] Assume that the Python library for the data to be crawled is "ray", and use this as an example to show the execution process of the method for automatically maintaining a version-controlled dataset based on a multi-agent system. The specific steps are as follows:

[0052] As Figure 1 shown, step S1: According to the list of legal Python libraries retrieved from the PyPI website, retrieve the projects of these libraries on the GitHub website to obtain the meta-information of the libraries, including the number of stars of the project, download address, latest version, release time, etc.; filter out Python libraries with a total number of stars of the project greater than 10,000 or an increase in the number of stars greater than 1,000 in the past month. Combine the library version information included in the local VersiCode dataset. For example, for the ray library, the number of stars of the project is greater than 10,000. Assume that the latest version of the ray library in the VersiCode dataset is 1.12.0, and the new released versions are 1.12.1 and 1.13.0. By comparing the latest version in the meta-information and the records of the library versions in the locally maintained dataset, determine that the data to be crawled is the source code packages of ray-1.12.1 and ray-1.13.0.

[0053] As Figure 2 shown, step S2: Based on the list of Python libraries to be crawled obtained in step S1, use a script to parse the requires_python field in the meta-information of ray-1.12.1 and ray-1.13.0 based on regular expressions to obtain the Python version environment required for installing the library version. Further, use the preprocess function module of the AexPy tool to crawl the source code packages of ray-1.12.1 and ray-1.13.0 that meet the corresponding Python versions in turn. Further, use the extract function module of the AexPy tool to extract the APIs in the library version, including modules, classes, functions, and attributes, to form the internal API knowledge base of each library version. Further, based on the extracted library version APIs, use a script to compare two consecutive library versions to find the evolved APIs and form the API evolution knowledge graph between library versions. The evolution types refer to those mentioned in the reference "How Do Python Framework APIs Evolve?An Exploratory Study".

[0054] As Figure 3 As shown, Step S3: Based on the API evolution knowledge graph between library versions obtained in Step S2, use an LLM-based agent to extract code usage examples from the API docstring. For example, the "range_arrow" function in ray-1.12.1 was renamed to "range_table" in ray-1.13.0. Use the LLM agent to process their docstrings and extract code usage examples as follows:

[0055] {"liabrary_version":"ray==1.12.1","API":"range_arrow",

[0056] "code_example":"import ray\n ds=ray.data.range_arrow(1000)\n ds.map(lambda r:{"v2":r[\"value\"]*2}).show()”}

[0057] Among them, the code_example field is from the docstring of the range_arrow function. The extraction method is to write the possible code example patterns in the docstring into the prompt as context to inform the LLM agent, so as to guide it to identify and extract valid code snippets.

[0058] Next, use the LLM-based agent to generate a functionality description for the successfully extracted code snippets, for example:

[0059] {"liabrary_version":"ray==1.12.1","API":"range_arrow",

[0060] "code_example":"import ray\n ds=ray.data.range_arrow(1000)\n ds.map(lambda r:{"v2":r[\"value\"]*2}).show()”,

[0061] "functionality_description":"The code creates a dataset,maps eachelement to double its value,and displays the results.”}

[0062] Among them, the functionality_description field is the function description generated by the LLM agent for the code_example. The role of the function description is to be written into the prompt as part of the dataset, together with relevant information such as the library version, as context to inform the LLM agent, so as to guide it to generate code corresponding to the version and function. Thus, metadata in the form of a tuple and containing at least three core elements, namely the library version, the function description, and the code snippet, is obtained for subsequent task form processing.

[0063] As Figure 4 shown, step S4: Based on the metadata obtained in step S3, perform dataset task form processing. Taking the code completion task (token level) as an example, as follows: Replace "range_arrow" in the "code_example" field with " <mask>", that is, perform masking processing to obtain "masked_code": "import ray\n ds=ray.data. <mask>(1000)\n ds.map(lambda r:{"v2":r["value"]*2}).show()”. Inform the LLM of "library_verison", "functionality_description", and "masked_code" to let the LLM make inference predictions" <mask>"Part, i.e., the version-sensitive part, is used to detect the performance of the LLM for version-sensitive code completion. Taking the code migration task as an example, for the evolution types involving cross-versions, such as API renaming, relocation, etc. Based on the API evolution knowledge graph between library versions, each set of data of two-version APIs is used as a test instance. The library version, function description, and code snippet of ray == 1.12.1 (source version), and the library version and function description of ray == 1.13.0 (target version) are informed to the LLM, allowing the LLM to infer the code snippet of the evolved API "range_arrow" in the target version, that is, to implement the code migration task of ray == 1.12.1 to ray == 1.13.0 for the API "range_arrow" (range_arrow is renamed to range_table in ray == 1.13.0). So far, the final form of data has been obtained, which is the updated data incorporated into the VersiCode dataset.

[0064] The technical means disclosed in the solution of the present invention are not limited to the technical means disclosed in the above embodiments, but also include technical solutions composed of any combination of the above technical features.< / mask> < / mask> < / mask> < / mask>

Claims

1. An automatically evolving method for a version - controllable dataset based on a large - language model, characterized in that: It includes a data crawling module, a data parsing module, a data annotation module, and a task form processing module. Specifically, it includes the following steps: S1. The data crawling module collects metadata of Python library versions from multiple data sources through automated means, ensuring the integrity and timeliness of the collected data, and performing preliminary screening on the data; S2. The data parsing module performs structured processing and preliminary cleaning on the crawled data through automated means, extracts key fields, and standardizes the data format to ensure the consistency and accuracy of the data, providing reliable input for subsequent annotation and processing; S3. The data annotation module automatically annotates the parsed data through an LLM-based proxy and preset rules, and adjusts the annotation strategy according to different task scenarios; S4. The task form processing module converts the annotated data into the form required for specific tasks through an LLM-based proxy and automated script tools, making it more suitable for the evaluation process, and ensuring that the data can meet the LLM performance evaluation scenario.

2. The automatic evolution method of a version-controllable dataset based on a large language model according to claim 1, wherein: The specific steps of S1 are as follows: S101. Use a crawler script to retrieve the PyPI website and obtain a list of legal Python libraries; S102. Retrieve the Python library projects in the list on the Github website and obtain meta-information, including the number of stars of the project and the release time; S103. Screen the target Python libraries based on the Github star threshold; S104. Combine the local VersiCode dataset situation. For libraries that already exist in the dataset, only crawl the source code packages of the new versions, and for new libraries, crawl all historical versions within the specified time range; S105. Organize the relevant meta-information of the two types of Python libraries and hand it over to the data parsing module for subsequent data processing.

3. The automatic evolution method of a version-controlled data set based on a large language model according to claim 1, characterized in that: The specific steps of S2 are as follows: S201. Optimize the AexPy tool to make it meet the requirements of dataset construction; S202. Based on the Python library meta-information, parse the requires_python field corresponding to each library version to obtain the Python version environment required to install this library version; S203. Use the preprocess function module of the AexPy tool to crawl the source code packages of the library versions corresponding to the Python version; S204. Use the extract function module of the AexPy tool to perform API extraction and obtain the internal knowledge base of each library version; S205. According to the Python API evolution pattern, based on the internal knowledge base of each library version, find the APIs that have evolved between two consecutive library versions, and obtain the API evolution knowledge graph between library versions.

4. The automatic evolution method of a version - controllable data set based on a large - language model according to claim 1, characterized in that: The specific steps of S3 are as follows: S301. Use the LLM proxy to extract code usage examples for the evolved APIs based on the docsting in the meta-information; S302. Use the LLM proxy to generate function descriptions for the successfully extracted code snippets; S303. Perform data cleaning according to heuristic rules to obtain the successfully processed data, denoted as meta-data, which contains at least three core elements: library version, function description, and code snippet.

5. The automatic evolution method of a version - controllable data set based on a large - language model according to claim 1, characterized in that: The specific steps of S4 are as follows: S401. Further process, based on metadata, into two task forms of code completion and code migration included in the VersiCode dataset; For the code completion task, locate version-sensitive APIs based on AST parsing and replace the target APIs with <mask>Mark, combined with context hints containing library version constraints, to construct code completion data; < / mask> For the code migration task, based on the API evolution knowledge graph between library versions, combined with context hints containing library version constraints, construct code migration data pairs, including source version data and target version data.

Citation Information

Patent Citations

  • Software project and third-party library knowledge graph construction method for software system

    CN111241307A

  • Python domain knowledge graph construction method

    CN115291944A

  • Automatic API case updating method based on double-side change information of library and customer project

    CN116028115A

  • Zero-sample large model generation code detection method and system

    CN117608648A

  • Question and answer method, system and device and medium

    CN118093828A

Cited By

  • Optimization method, device and equipment for large language model oriented to extensive scene, storage medium and product

    CN121166656A