Python dependency library API migration method based on large language model
Through the multi-proxy strategy of the large language model, the programmer process is simulated to carry out Python dependency library API migration, solving the high labor cost and inefficiency problems caused by relying on structured data in the existing technology, and achieving automated and high-accuracy API migration, which is suitable for tasks such as artificial intelligence software upgrades, legacy code transformations, etc.
Patent Information
- Application Number
- CN202510532974.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2045-04-25
AI Technical Summary
The existing technology relies on structured data during the Python dependency library API migration process, resulting in high labor costs and low migration efficiency, making it difficult to achieve automated and efficient API migration.
Using a multi-proxy strategy based on a large language model, we simulate the programmer process, through four modules: user intent identification, network resource discovery and collection, migration and testing verification, and use unstructured web page information to perform API migration, including syntax tree information and differential testing, reducing labor costs and improving migration accuracy.
It realizes automated migration of Python dependent library APIs that do not rely on structured data, reduces labor costs, improves migration efficiency and accuracy, and is suitable for code updates, code reconstruction and legacy code processing.
Smart Images

Figure CN120429008A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer software technology, specifically, to the technical field of API migration between large language models and Python programming language, and in particular to a Python dependency library API migration method based on a large language model. Background Art
[0002] Python libraries are an important part of building software systems, especially artificial intelligence systems. Here we take the deep learning framework TensorFlow and the data processing library Pandas as examples. TensorFlow is a widely used deep learning framework that provides a rich API (Application Programming Interface, API) Python libraries and their APIs play an essential role in software engineering. To add new features, fix security vulnerabilities, and upgrade the technology stack, Python libraries often undergo frequent version updates, resulting in frequent migrations of existing code, such as from older versions to newer versions.
[0003] API migration can currently be performed using two methods: manual migration and automatic migration. Manual migration means that programmers, based on a full understanding of the existing code, the pre-migration API, and the post-migration API, manually search for equivalent APIs for all old versions of the API in the new version of the dependency library and verify their correctness. This manual migration process is a typical code debugging process, which is time-consuming and labor-intensive and requires a lot of manpower. A series of automatic API migration methods can reduce the labor cost of manual API migration, but most of these methods rely on pre-constructed structured data, such as dependency migration mappings and code pairs. This structured data is often difficult to obtain. Therefore, API migration remains a time-consuming and labor-intensive problem. It is crucial to establish an automated API migration method that does not rely on pre-constructed structured data. Summary of the Invention
[0004] To address the above-mentioned shortcomings of the existing technology, the present invention proposes an automated, structured data-independent, large language model-based Python dependency library API migration method. Based on large language models (LLMs), the present invention adopts a multi-agent strategy to simulate and automate the migration of Python dependency library APIs by programmers, reducing the labor cost of API migration and improving the efficiency of Python dependency library API migration. The multi-agent strategy enables large language models to achieve good performance in complex tasks such as API migration. The present invention can be applied to software engineering-related tasks such as code updates, code refactoring, and legacy code processing, helping to reduce the labor cost of software development and maintenance.
[0005] The system of the present invention consists of four modules: user intent recognition, network resource discovery and acquisition, migration, and testing and verification. It involves several agents based on a large language model. Using unstructured web data readily available online, the system completes the migration of user-specified dependency library APIs within the user-provided code through five steps: user intent recognition, network resource discovery, random walk-based information acquisition, migration, and testing and verification. The intent recognition and testing and verification steps effectively mitigate the unreliable output caused by hallucinations in large language models. The technical solution of the present invention is described in detail below.
[0006] The present invention provides a Python dependency library API migration method based on a large language model, which uses a multi-agent strategy based on a large language model to migrate and automate the Python dependency library API, including the following steps: Step 1: User Intent Identification Identify specific APIs that need to be migrated in the code to be migrated based on the code to be migrated and the migration requirements input by the user; Step 2: Network resource discovery and collection Based on the migration requirements input by users, relevant network resources are discovered and collected to obtain the necessary knowledge for API migration; Step 3: Migrate Perform API migration based on user migration requirements and refer to the acquired knowledge about API migration tasks. Step 4: Test and Verify Verify the multiple generated migrated codes to determine the correctness of the migration.
[0007] In the present invention, in step 1, a proxy based on a large language model is used to determine the user's true migration intention.
[0008] In the present invention, in step 2, two agents based on a large language model are used to discover network resources from trusted sources and determine whether the collected network resources are sufficient to support the API migration task; wherein, one agent is first used to search and record the URL in the Python document source according to the specific migration requirements input by the user, and then another agent is used to determine whether all the recorded web page resources meet the requirements of the migration task.
[0009] In this invention, Python documentation sources include PyPI, DuckDuckGo, and GitHub.
[0010] In the present invention, in step 2, three agents based on a large language model are used to collect information on the discovered network resources; specifically, the method includes: first, using one agent to extract keywords from all discovered network resources; then, using one agent to score and rank the importance of each network resource to the migration task based on the extracted keywords; and finally, using one agent to collect API-related information contained in the discovered network resources based on the importance ranking in a random walk manner.
[0011] In the present invention, when scoring and ranking the importance of the migration task, the text matching degree between the keywords and the migration requirements and the semantic matching degree measured by cosine similarity are used for measurement.
[0012] In the present invention, in step 3, an agent based on a large language model is used to perform API migration; specifically, the method includes: using an agent to generate multiple copies of migrated code based on the code to be migrated and the migration requirements input by the user, while referring to the AST syntax tree of the code to be migrated.
[0013] In the present invention, in step 4, two agents based on a large language model are used to test and verify the migrated code; specifically, the process includes: first using an agent based on a large language model to generate test cases for the pre-migration and post-migration codes, and then using an agent based on a large language model to build a test environment and perform differential testing on the pre-migration and post-migration codes, verifying the correctness of the generated post-migration codes one by one.
[0014] The present invention also provides a Python dependency library API migration system based on a large language model, which is used to execute the above method, and includes a user intention recognition module, a network resource discovery and acquisition module, a migration module and a test verification module; wherein: The user intent recognition module identifies the specific APIs that need to be migrated in the code to be migrated based on the code to be migrated and the migration requirements input by the user; The network resource discovery and acquisition module discovers and acquires relevant network resources based on the migration requirements entered by the user, thereby acquiring the necessary knowledge for API migration; The migration module performs API migration based on user migration requirements and references the acquired knowledge about API migration tasks. The test verification module verifies the multiple post-migration codes generated by the migration module to determine the correctness of the migration.
[0015] Compared with the prior art, the present invention has the following beneficial effects: Python dependency library API automated migration method using multi-agent design pattern: This method imitates the API migration workflow of human programmers and achieves automation using a large language model. Existing automated API migration methods often require difficult-to-obtain, pre-built structured information. This method overcomes this shortcoming and can achieve accurate and low-cost automated migration of Python dependency library APIs. Here, the traditional API automated migration method that relies on structured information is used as a comparison method for illustration. The execution process of this method and the comparison method is compared. Figure 3 As shown in the figure, this method eliminates the two steps of structured information construction and model training, achieving end-to-end input and output. In preliminary evaluation experiments, this method saved a lot of manpower and economic costs associated with structured information construction and model training.
[0016] Information collection method based on random walk: This method discovers relevant knowledge related to API in trusted sources according to the migration requirements of user input, and collects it according to importance, providing reliable and high-quality input for large language models. Random walk is mostly used for web page ranking in search engines, and traditional methods mostly use traversal for information collection. The introduction of random walk can improve the relevance and quality of the collected information. Here, the traditional information collection method using traversal is used as a comparison method for explanation, and the execution process of this method and the comparison method is compared. Figure 4 As shown in the figure, this method only collects information with an importance score exceeding a threshold, and the random walk strategy ensures that more important and relevant information is collected first. The comparison method does not judge importance and relevance. Therefore, the information collected by this method is of higher quality and more relevant to the transfer task. Because large language models are sensitive to input information, higher-quality input information will lead to more accurate output. In preliminary evaluation experiments, the random walk-based information collection method effectively improved the accuracy of automatic transfer compared to the comparison method.
[0017] Introducing syntax trees in the migration step: Methods that use large language models to handle code-related tasks usually treat code and natural language equally, relying on the knowledge acquired by the large language model through extensive training to complete the task. In the migration step, the present invention provides the syntax tree information of the code to be migrated to the large language model, and the migration agent decides whether to use it. Here, the traditional automation method that treats code and natural language equally is used as a comparison method for illustration. The execution process of this method and the comparison method is compared. Figure 5 As shown in the figure, this method not only considers the semantic and syntactic information of the code, but also the structural information of the code as a semi-structured language, conveyed by the syntax tree. The comparison method does not consider this structural information. In preliminary evaluation experiments, the structural information of the semi-structured language conveyed by the syntax tree effectively improved the accuracy of automatic transfer compared to the comparison method.
[0018] Test verification method based on differential testing: This method is used to verify whether the automatically generated API migration plan is accurate and to remove inaccurate migration plans. Differential testing is often used to discover defects in software systems. Traditional automated API migration methods rely on structured data (such as code pairs) to provide verification standards. The introduction of differential testing means that verification no longer requires structured data represented by code pairs, and the accuracy of automatic API migration is also improved. Here, the traditional automated method that relies on code pairs to provide inspection standards is used as a comparison method for illustration. The execution process of this method and the comparison method is compared. Figure 6 As shown in the figure, since the migration plan is not unique, the migrated code is not unique. The comparison method uses textual matching between the migrated code and the verification criteria to determine whether the migration is successful. This judgment method cannot handle cases where the migrated code is not unique, and is prone to misjudgment. This method does not focus on textual matching, but instead focuses on the input and output of the code. As long as the pre- and post-migration code produces the same output for the same input, the migration is considered successful. This method effectively improves the accuracy of determining whether the migration is successful, thereby improving migration quality. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 This is a schematic diagram of the Python dependency library API migration method framework based on a large language model.
[0020] Figure 2 It is the execution process of the Python dependency library API migration method based on the large language model.
[0021] Figure 3 This is a schematic diagram comparing the Python dependency library API migration method based on a large language model and the traditional API automatic migration method that relies on structured information.
[0022] Figure 4This is a schematic diagram comparing the information collection method of the Python dependency library API migration method based on a large language model with the traditional information collection method using traversal.
[0023] Figure 5 This diagram compares the Python dependency library API migration method based on a large language model with the traditional automated method that treats code and natural language equally during migration.
[0024] Figure 6 This diagram shows a comparison between the Python dependency library API migration method based on a large language model and the traditional automated method that relies on code to provide inspection standards in verification. DETAILED DESCRIPTION
[0025] The technical solution of the present invention is described in detail below with reference to the accompanying drawings and embodiments.
[0026] In the API migration task, it can be observed that human programmers can accurately migrate APIs, and a large amount of structured data is not required during the migration process. During the API migration process, programmers obtain information more from information sources such as documents and web pages, and the content provided by these information sources is unstructured. After the information is acquired, the programmer uses the professional knowledge and experience accumulated in the work to process this information and finally comes up with a migration plan. At the same time, the large language model is trained on a large amount of predictions including code, so it is reasonable to assume that professional knowledge related to the code is already available. The above observations inspired the present invention to use a large language model to simulate the process of human programmers performing API migration, thereby obtaining an automated API migration method that does not rely on pre-constructed structured data.
[0027] However, achieving the above approach presents two pressing challenges. First, large language models are essentially more complex deep neural networks. The interpretability issues and black-box nature of deep neural networks make large language models susceptible to hallucinations, potentially producing unreliable outputs (Challenge 1). Consequently, automated API migration may produce unreliable migration solutions. Second, the effectiveness of large language models is sensitive to input information. If the input information is of low quality, it is difficult to generate an accurate migration solution, and the effectiveness of automated API migration cannot be guaranteed (Challenge 2).
[0028] The present invention uses a multi-agent design pattern to propose an automated Python dependency library API migration method based on a large language model. To address the first challenge, the invention designs a user intent recognition module and a test verification module. These two modules can effectively avoid the inaccurate API migration solution caused by the illusion of a large language model, which leads to invalid reconstruction of the code and inaccurate output. To address the second challenge, the present invention designs a network resource discovery and acquisition module. This module can discover relevant knowledge related to the API from a trusted source based on the migration requirements input by the user, and collect it based on the random walk strategy according to its importance, providing reliable and high-quality input for the large language model.
[0029] The framework of this method is as follows Figure 1 This method takes the code to be migrated and the migration requirements as input and the migrated code as output. It mainly consists of four parts: user intention recognition module, network resource discovery and acquisition module, migration module and test verification module. The relationship between the four parts is as follows: Figure 1 shown.
[0030] The User Intent Recognition module identifies the specific APIs within the code to be migrated based on the user's input of the code to be migrated and the migration requirements. A piece of code to be migrated often calls multiple APIs, many of which may not require migration. This module involves a proxy based on a large language model to prevent unnecessary refactoring of the code to be migrated due to hallucinations.
[0031] The network resource discovery and acquisition module discovers and acquires relevant network resources based on the migration requirements entered by the user, thereby acquiring the necessary knowledge for API migration. This module first searches for network resources in three trusted sources, then determines the importance of the search results, and finally acquires information based on this importance. This module involves a total of five agents based on the large language model. This module provides reliable and task-related knowledge for the migration task, effectively helping the large language model complete the migration task.
[0032] The migration module performs API migration based on user migration requirements and references existing knowledge about API migration tasks. This module generates multiple copies of the migrated code. This module involves a single agent based on a large language model, which independently decides whether to reference the syntax tree of the migrated code during the migration process. The inclusion of a syntax tree helps improve the accuracy of API migrations based on the large model.
[0033] The test verification module verifies the post-migration code generated by the migration module to determine the correctness of the migration. This module uses differential testing technology for verification, involving two agents based on a large language model. The migration is considered successful if the pre- and post-migration code produce the same output for all test cases generated by the module.
[0034] The execution process of the present invention is as follows Figure 2 As shown; the specific implementation method of the present invention is as follows: Step 1: User Intent Identification. This step involves an agent (Agent 1) based on the large language model and consists of one substep. User intent identification effectively avoids unnecessary refactoring of the code to be migrated due to hallucinations in the large language model.
[0035] Step 1.1: Agent 1 identifies the specific APIs that need to be migrated in the code to be migrated based on the code to be migrated and the migration requirements input by the user.
[0036] Step 2: Network Resource Discovery. This step consists of two substeps, involving two agents (Agent 2 and Agent 3) based on the large language model. Network resource discovery provides the large language model with reliable information for the migration task, ensuring its quality.
[0037] Step 2.1: Agent 2 searches for network resources from three trusted sources, PyPI, DuckDuckGo, and GitHub, based on the user's migration requirements. This creates a directory of network resources related to the migration task. This directory records the URLs of the network resources.
[0038] Step 2.2: Agent 3 determines whether the obtained network resource directory related to the migration task is sufficient to support the migration task based on the user's migration requirements. If the information is sufficient, proceed to the next step. If the information is insufficient to support the migration task, repeat step 2.1.
[0039] Step 3: Network Resource Collection. This step consists of three substeps, involving three agents (Agents 4-6) based on the large language model. Network resource collection uses a random walk strategy, prioritizing network resources based on importance. This step provides the large language model with more task-relevant information, ensuring the quality of the transfer task.
[0040] Step 3.1: Agent 4 performs a network resource search on each URL in the network resource directory related to the migration task. Perform keyword extraction.
[0041] Step 3.2: Agent 5 scores the importance of the network resources based on the keywords extracted in step 2.3 according to the user's migration requirements. The importance is measured by the textual match between the keywords and the migration requirements and the semantic match using the cosine similarity metric.
[0042] Step 3.3: Based on the scoring results of step 3.2, agent 6 uses a random walk strategy to extract the information contained in the network resources in an importance-prioritized manner, thereby building the knowledge base required for the migration task.
[0043] Step 4: Migration. This step consists of one sub-step and involves one agent (agent 7) based on the large language model.
[0044] Step 4.1: Agent 7 uses the knowledge base required for the migration task built in Step 2 and decides whether to refer to the syntax tree of the code to be migrated to generate multiple copies of the migrated code.
[0045] Step 5: Testing and Verification. This step consists of two substeps and involves two agents (Agent 8 and Agent 9) based on the large language model. This testing and verification uses differential testing to determine the success of the migration task. The migration is considered successful if the pre- and post-migration code produce the same output for a series of test cases.
[0046] Step 5.1: Agent 8 generates a set of test cases for the code to be migrated and a copy of the migrated code.
[0047] Step 5.2: Agent 9 sets up a test environment and performs differential testing based on the test cases generated in Step 4.1. If the differential test passes, the migration is successful and the migrated code is output to the user. If the differential test fails, the process returns to Step 5.1 to verify the correctness of the next migrated code.
[0048] The present invention can be applied to the following main application scenarios: 1) AI software upgrades and maintenance. During the development and maintenance of AI software systems, Python dependency libraries frequently update, requiring constant migration of APIs within the code. This invention automates the migration of dependency library APIs, eliminating the need for manual code modification and significantly reducing labor costs.
[0049] 2) Legacy Code Transformation. Many companies have accumulated a large amount of legacy code that relies on older versions of Python libraries. However, due to historical reasons, this code cannot be directly run in modern software environments. This invention can be applied to the modernization of legacy code, automatically migrating the legacy Python library APIs that the legacy code relies on to current mainstream library versions. This invention avoids the huge costs of redevelopment for companies.
Claims
1. A Python dependency library API migration method based on a large language model, characterized in that: A multi-agent strategy based on a large language model is used to migrate and automate the Python dependency library API, including the following steps: Step 1: User Intent Identification Identify specific APIs that need to be migrated in the code to be migrated based on the code to be migrated and the migration requirements input by the user; Step 2: Network resource discovery and collection Based on the migration requirements input by users, relevant network resources are discovered and collected to obtain the necessary knowledge for API migration; Step 3: Migrate Perform API migration based on user migration requirements and refer to the acquired knowledge about API migration tasks. Step 4: Test and Verify Verify the multiple generated migrated codes to determine the correctness of the migration.
2. The Python dependency library API migration method based on a large language model according to claim 1, characterized in that: In step 1, a proxy based on a large language model is used to determine the user's true transfer intention.
3. The Python dependency library API migration method based on a large language model according to claim 1, characterized in that: In step 2, two agents based on large language models are used to discover network resources from trusted sources and determine whether the collected network resources are sufficient to support the API migration task. One agent is first used to search and record the URLs in the Python document source according to the specific migration requirements entered by the user, and then another agent is used to determine whether all recorded web resources meet the requirements of the migration task.
4. The Python dependency library API migration method based on a large language model according to claim 3, characterized in that: Sources of Python documentation include PyPI, DuckDuckGo, and GitHub.
5. The Python dependency library API migration method based on a large language model according to claim 1, characterized in that: In step 2, three agents based on a large language model are used to collect information about the discovered network resources. Specifically, one agent is used to extract keywords from all discovered network resources. Another agent is used to score and rank the importance of each network resource to the migration task based on the extracted keywords. Finally, another agent is used to collect API-related information contained in the discovered network resources based on the importance ranking using a random walk method.
6. The Python dependency library API migration method based on a large language model according to claim 5, characterized in that: When scoring and ranking the importance of the migration task, the textual matching degree between the keywords and the migration requirements and the semantic matching degree using the cosine similarity measurement are used for measurement.
7. The Python dependency library API migration method based on a large language model according to claim 1, characterized in that: In step 3, an agent based on a large language model is used to perform API migration. Specifically, the agent generates multiple copies of migrated code based on the code to be migrated and the migration requirements entered by the user, while referring to the AST syntax tree of the code to be migrated.
8. The Python dependency library API migration method based on a large language model according to claim 1, characterized in that: In step 4, two agents based on the large language model are used to test and verify the migrated code. Specifically, one agent based on the large language model is used to generate test cases for the pre- and post-migration code. Another agent based on the large language model is then used to build a test environment and perform differential testing on the pre- and post-migration code, verifying the correctness of the generated post-migration code one by one.
9. A Python dependency library API migration system based on a large language model, which is used to execute the method according to any one of claims 1 to 8, characterized in that: It includes a user intention recognition module, a network resource discovery and acquisition module, a migration module, and a test and verification module; among which: The user intent recognition module identifies the specific APIs that need to be migrated in the code to be migrated based on the code to be migrated and the migration requirements input by the user; The network resource discovery and acquisition module discovers and acquires relevant network resources based on the migration requirements entered by the user, thereby acquiring the necessary knowledge for API migration; The migration module performs API migration based on user migration requirements and references the acquired knowledge about API migration tasks. The test verification module verifies the multiple post-migration codes generated by the migration module to determine the correctness of the migration.
Citation Information
Patent Citations
Code abstract generation method based on code knowledge graph and knowledge migration
CN111797242A
Source code credential migration system and method
CN116450213A
PyPI warehouse system facing autonomous architecture and rapid construction method
CN116661858A
Software automatic migration and optimization method based on domestic software and hardware environment
CN116737232A
Hybrid front-end framework migration method based on AST and LLM
CN117608656A