API code optimization method and device based on large language model and electronic equipment

By building knowledge building modules and large language models, extracting and storing API architecture features and generating operational knowledge, the performance and compatibility problems caused by API misuse are solved, and the diversity and security improvement of code optimization is achieved.

CN120256281AInactive Publication Date: 2025-07-04JIANGXI NORMAL UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510676043.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-24
Publication Date
2025-07-04
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the prior art, the misuse of APIs has caused the impact of program performance and resource efficiency, the risk of data leakage increases, the compatibility problem during system upgrades has not been effectively solved, and the rules-based methods are not comprehensive and labor-intensive.

Method used

Build a knowledge building module including text collection unit, extraction unit and knowledge graph unit, extract and store API architecture features through large language models, generate operational knowledge, and guide code optimization.

Benefits of technology

It improves the understanding of APIs by the large language model, provides diversified code optimization results, enhances program execution efficiency, resource utilization and security, reduces the risk of data leakage, and improves system upgrade compatibility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120256281A_ABST
    Figure CN120256281A_ABST
Patent Text Reader

Abstract

The invention provides an API code optimization method and device based on a large language model and electronic equipment, and belongs to the field of software code optimization. The method comprises the following steps: constructing a knowledge construction module comprising a text collection unit, an extraction unit and a knowledge graph unit; a code sample is obtained through a text collection unit, and API architecture features of the code sample are extracted through an extraction unit; the API architecture features are stored in a triple form through a knowledge graph unit to form a knowledge base; extracting a to-be-retrieved API of the to-be-optimized code by using a knowledge retriever; obtaining a triple matched with the to-be-retrieved API and a corresponding explanatory text from a knowledge base to generate retrieval information; the retrieval information is converted into operable knowledge through a knowledge injector; and transferring the operable knowledge to the large language model, and generating and outputting a plurality of code optimization results. According to the method, the understanding capability of the large model on the API is enhanced, and the code optimization efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of software code optimization, and particularly relates to an API code optimization method, device, and electronic device based on a large language model. Background Art

[0002] In modern software development, non-functional errors, although not immediately causing system failures, can significantly affect the quality and efficiency of software systems. These errors may lead to poor performance, low resource utilization efficiency, increased latency, and may even introduce security vulnerabilities. Therefore, identifying and mitigating non-functional errors is crucial for maintaining high-quality software.

[0003] Code optimization has become a key practice for addressing non-functional errors in software. In recent years, various techniques have been proposed to improve code performance and resource management from different aspects. For example, to address the problem of low execution performance, methods for optimizing loop structures have been proposed. These methods mainly focus on improving the efficiency of repetitive operations, which are usually performance bottlenecks. In addition, methods for optimizing data structures have been proposed to improve resource utilization. Efficient data structures can significantly reduce memory usage and processing time, making the application more responsive and scalable. Finally, problems such as data leakage and poor maintainability have been addressed by optimizing database search algorithms and implementing secure coding practices. These methods can ensure the protection of sensitive information and improve the maintainability and scalability of software.

[0004] However, existing work has overlooked a common and important cause of non-functional errors, namely the misuse of APIs. Specifically, overusing complex and memory-consuming APIs can directly affect program performance and resource efficiency, calling unauthorized APIs may pose a risk of data leakage, and relying on outdated or unstable APIs may cause compatibility issues during system upgrades.

[0005] This application particularly focuses on the use of APIs in code optimization, aiming to improve the execution efficiency, resource utilization, security, and maintainability of programs by improving the use of APIs in code. To effectively adopt the best API practices, it is crucial to deeply understand various API usage scenarios and constraints. Previous studies have attempted to detect API misuse in code snippets using API usage patterns collected from API specifications or large-scale projects. However, these studies have not focused on the non-functional errors introduced by API misuse, and rule-based methods are not only labor-intensive but also not comprehensive enough.

[0006] With the emergence of large language models (LLMs), the field of code intelligence has witnessed a remarkable revolution. LLMs have demonstrated excellent capabilities in code understanding and have excelled in multiple tasks, including code generation and vulnerability detection. Leveraging their advanced code understanding capabilities and the extensive knowledge obtained from various codebases, LLMs can master the knowledge of using various APIs and assist in code optimization. However, relying solely on the general knowledge of LLMs is insufficient to meet the requirements of real-world tasks and is still prone to the limitation of a single optimization result. Summary of the Invention

[0007] The purpose of the embodiments of this application is to provide an API code optimization method, device, and electronic device based on a large language model to solve the problems in the prior art that the misuse of APIs affects program performance and resource efficiency, there is a risk of data leakage, and incompatibility occurs during system upgrades.

[0008] To solve the above technical problems, this application is implemented as follows: In a first aspect, the embodiments of this application provide an API code optimization method based on a large language model. The method includes: Construct a knowledge construction module including a text collection unit, an extraction unit, and a knowledge graph unit; Obtain code samples through the text collection unit and extract the API architecture features of the code samples using the extraction unit; among them, the API architecture features include entity samples and non-functional relationship samples; Store the API architecture features in the form of triples through the knowledge graph unit to form a knowledge base; Use a knowledge retriever to extract the APIs to be retrieved in the code to be optimized; Obtain the triples and corresponding explanatory texts matching the APIs to be retrieved from the knowledge base to generate retrieval information; Convert the retrieval information into operationalizable knowledge through a knowledge injector; among them, the operationalizable knowledge includes entities, relationships, instructions, and rules corresponding to the retrieval information; Transmit the operationalizable knowledge to the large language model to generate and output multiple code optimization results.

[0009] Preferably, the specific steps of obtaining code samples through the text collection unit and extracting the API architecture features of the code samples using the extraction unit include: Use the text collection unit to collect initial texts according to screening criteria; among them, the screening criteria include at least: there is at least one feasible code solution in the initial text, there is at least one API in the feasible code solution, and the scoring value of the feasible code solution reaches a preset score; Split the initial text into multiple paragraphs to form code samples.

[0010] Preferably, the specific steps of obtaining the code sample through the text collection unit and extracting the API architecture features of the code sample by the extraction unit further include: Extract the API samples in the code sample; Extract the fully qualified name samples of the API samples; Input the fully qualified name samples into the large language model to generate relevant knowledge samples of the API samples; Based on the relevant knowledge samples, determine the relationship samples of the API samples.

[0011] Preferably, the specific steps of extracting the fully qualified name samples of the API samples include: Extract the simple name pairs of the API samples; wherein, the simple name pairs are composed of class names and method names or individual method names; Complete the simple name pairs to obtain the fully qualified name samples.

[0012] Preferably, the specific steps of inputting the fully qualified name samples into the large language model to generate relevant knowledge samples of the API samples include: Preset multiple API relationship types; wherein, the multiple API relationship types include: efficiency comparison, function similarity, function collaboration, function replacement, logical constraint, and behavior difference; Generate corresponding knowledge mining prompts for the multiple API relationship types; Based on the knowledge mining prompts, input the fully qualified name samples into the large language model to generate knowledge samples.

[0013] Preferably, the specific steps of using the knowledge retriever to extract the to-be-retrieved APIs of the code to be optimized include: According to the declarations and method instances in the code to be optimized, respectively extract the to-be-retrieved class APIs and to-be-retrieved method APIs; Use the extraction unit to extract the to-be-retrieved fully qualified name of the to-be-retrieved method API; wherein, both the to-be-retrieved class API and the to-be-retrieved method API are presented in the form of fully qualified names.

[0014] Preferably, the specific steps of obtaining the triples and corresponding explanatory texts matching the to-be-retrieved APIs from the knowledge base to generate retrieval information include: According to the entity names of the to-be-retrieved APIs, obtain the corresponding paired API information from the knowledge base; wherein, the paired API information includes triples and corresponding explanatory texts; Use the large language model to calculate the similarity between the to-be-retrieved APIs and the paired API information; According to the similarity, screen the paired API information to obtain the retrieval information.

[0015] Preferably, the specific steps of transmitting operationalizable knowledge to the large language model to generate and output multiple code optimization results include: Convert the operationalizable knowledge into structured prompts; Input the structured prompts into the large language model to output multiple code optimization results.

[0016] Compared with the prior art, the above technical solutions provided by this application at least include the following beneficial effects: 1. This application defines a form of operationalizable API knowledge, including four API elements: API entities, relationships, usages, and rules. This knowledge enhances the large language model's understanding of APIs, enabling it to more effectively guide code optimization. Exploring and defining this knowledge format for specific elements in domain knowledge serves as a valuable reference for the large language model when applying knowledge across different domains; 2. This application constructs a method for using operationalizable API knowledge to guide the large language model in code optimization. This method establishes a comprehensive pipeline from building a knowledge base to retrieving and applying this knowledge. This method effectively provides diverse results for the optimization of target code; 3. This application designs a strategy for the large language model to effectively utilize API knowledge. This strategy uses structured prompts as a bridge to dynamically inject API knowledge to guide the large language model in optimizing code. In addition to code optimization, this strategy also has the potential to provide technical guidance for large language models based on domain knowledge and can be applied to other generation tasks.

[0017] In a second aspect, an API code optimization device based on a large language model provided by an embodiment of this application includes: A unit module for constructing a knowledge construction module including a text collection unit, an extraction unit, and a knowledge graph unit; A sample module for obtaining code samples through the text collection unit and extracting API architecture features of the code samples using the extraction unit; among them, the API architecture features include entity samples and non-functional relationship samples; A knowledge graph module for storing the API architecture features in the form of triples through the knowledge graph unit to form a knowledge base; A retrieval and extraction module for using a knowledge retriever to extract the APIs to be retrieved in the code to be optimized; A knowledge retrieval module for obtaining triples and corresponding explanatory texts matching the APIs to be retrieved from the knowledge base to generate retrieval information; A knowledge injection module for converting the retrieval information into operationalizable knowledge through a knowledge injector; among them, the operationalizable knowledge includes entities, relationships, instructions, and rules corresponding to the retrieval information; An optimized code output module for transmitting operationalizable knowledge to a large language model to generate and output multiple code optimization results.

[0018] In a third aspect, an embodiment of the present application provides an electronic device, which includes a processor, a memory, and a program or instruction stored in the memory and executable on the processor. When the program or instruction is executed by the processor, the steps of the method described in the first aspect are implemented.

[0019] It can be understood that the beneficial effects of the technical solutions provided in the above second and third aspects can refer to the relevant descriptions in the first aspect above, and will not be elaborated here.

[0020] Additional aspects and advantages of the present application will be given in part in the following description, will become apparent in part from the following description, or will be learned through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The above and / or additional aspects and advantages of the present application will become apparent and be readily understood from the description of the embodiments in conjunction with the following drawings, where: Figure 1 is a schematic flowchart of an API code optimization method based on a large language model provided by some embodiments of the present application; Figure 2 is a block diagram of an API code optimization device based on a large language model shown by some embodiments of the present application; Figure 3 is a block diagram of an electronic device shown by some embodiments of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0022] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments of the present application belong to the scope of protection of the present application.

[0023] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein. In addition, "and / or" in the specification and claims means at least one of the connected objects, and the character " / " generally means an "or" relationship between the associated objects before and after.

[0024] The following will combine the accompanying drawings and, through specific embodiments and their application scenarios, elaborate in detail on an API code optimization method provided by the embodiments of the present application based on a large language model.

[0025] Figure 1 FIG. is a schematic flowchart of an API code optimization method based on a large language model shown in the first embodiment of the present application. Please refer to Figure 1 and this method includes: Step S101: Construct a knowledge construction module including a text collection unit, an extraction unit, and a knowledge graph unit; Step S102: Obtain code samples through the text collection unit and extract the API architecture features of the code samples by using the extraction unit; wherein, the API architecture features include entity samples and non-functional relationship samples; Specifically, use the text collection unit to collect initial texts according to the screening criteria; wherein, the screening criteria at least include: there is at least one feasible code solution in the initial text, there is at least one API in the feasible code solution, and the scoring value of the feasible code solution reaches a preset score; split the initial text into multiple paragraphs to form code samples.

[0026] Specifically, it further includes: extracting API samples from the code samples; extracting the fully qualified name samples of the API samples; inputting the fully qualified name samples into the large language model to generate relevant knowledge samples of the API samples; determining the relationship samples of the API samples based on the relevant knowledge samples.

[0027] Among them, extracting the fully qualified name samples of the API samples specifically includes: extracting the simple name pairs of the API samples; wherein, the simple name pairs are composed of a class name and a method name or a single method name; complete the simple name pairs to obtain the fully qualified name samples.

[0028] Among them, inputting the fully qualified name samples into the large language model to generate relevant knowledge samples of the API samples specifically includes: presetting multiple API relationship types; wherein, the multiple API relationship types include: efficiency comparison, function similarity, function collaboration, function replacement, logical constraint, and behavior difference; generating corresponding knowledge mining prompts for the multiple API relationship types; based on the knowledge mining prompts, input the fully qualified name samples into the large language model to generate knowledge samples.

[0029] In a possible implementation, the API text collection strategy of the text collection unit is as follows: collect natural language texts that potentially contain API knowledge from StackOverflow, and extract API architecture features including API entities and API non-functional relationships from the texts. Among them, the natural language texts need to meet the following conditions: texts from posts on the Stack Overflow forum; the posts contain the Java tag and have a score of no less than 10; the texts contain Java APIs.

[0030] For example, in this embodiment, 4,253 posts marked with "Java" are collected from Stack Overflow, and the highest-scoring answer in each post is selected as the data source. The screening criteria include: the post must have an accepted answer; the answer must contain at least one Java API; the answer score is at least 10 points; the answer text must involve at least two different Java APIs. The collected texts are split into multiple paragraphs, each paragraph containing rich API context information, providing a basis for subsequent API relationship extraction.

[0031] In a possible implementation, the process of using the extraction unit to extract the API architecture features of the code sample is sequentially and collaboratively completed by the following four sub-units: API non-fully qualified name extraction sub-unit: This unit takes the natural language text containing the API as input, and identifies and extracts the simple name pairs of Java APIs (API1 simple name, API2 simple name) in the text. The simple name consists of the class name combined with the method name or a single method name.

[0032] API fully qualified name completion sub-unit: This unit takes the API simple name pairs obtained from the previous unit as input, and automatically completes the simple name of each API to the fully qualified name, and updates the API simple name pairs to API pairs.

[0033] API knowledge extraction sub-unit: This unit takes the API fully qualified name as input, and uses the GPT-4o model to expand the corresponding functional knowledge for each API, including descriptions of the API's functions, features, performance, I / O streams, and API usage methods.

[0034] API relationship reasoning sub-unit: This unit integrates the results of the above three sub-units as the original input of the API relationship reasoning task, designs the API non-functional relationship definition as the reasoning basis, and uses the GPT-4o model as the reasoning model. The inferred API non-functional relationship is added to the API triple, and all triples are stored using the knowledge graph as the carrier to form the API knowledge base.

[0035] For example, first, the fully qualified names of APIs are extracted. Since the API names in the code may be ambiguous, especially simple names may correspond to multiple APIs. For example, "A List is a sequence of ordered elements, while a Set is a collection of unordered elements", which contains the behavioral difference relationship between List and Set. However, both List and Set are simple names and may refer to different APIs with fully qualified names. For example, List may be java.util.List or java.awt.List, and Set may be java.util.Set or org.hibernate.mapping.Set. This name ambiguity makes it difficult to accurately determine the relationship between two APIs. To solve this problem, before accurately establishing the API relationship, the fully qualified names of APIs need to be extracted from natural language text.

[0036] To this end, two units are designed in this embodiment: the API non-fully qualified name extraction unit and the API fully qualified name inference unit. The API non-fully qualified name extraction unit is used to extract simple names or partially qualified names in the text. The API fully qualified name inference unit is responsible for inferring the corresponding fully qualified name according to the context, such as java.util.List or java.awt.List. This step is the key to ensuring the accuracy of API relationship extraction. In this way, the name ambiguity can be eliminated and the relationship between APIs can be accurately inferred. And when the text contains non-fully qualified name references, the API non-fully qualified name extraction unit is deployed to extract these references. Subsequently, the API fully qualified name inference unit is used to infer the corresponding fully qualified name. These units are carefully designed and equipped with prompts to enhance their effects.

[0037] Among them, the prompt of the API non-fully qualified name extraction unit helps to extract the non-fully qualified names of APIs from natural language text, that is, simple names and partially qualified names; in this task, the task description is accompanied by five examples and provides a space for inputting the text to be processed and obtaining its non-fully qualified name. The prompt of the API fully qualified name inference unit converts the non-fully qualified name into a fully qualified name; in this task, the task description is "Parse the non-fully qualified name...", and then five fully qualified names and their corresponding fully qualified name examples are given. The non-fully qualified names generated by the API non-fully qualified name extraction unit will be appended to the end of the input text to generate the corresponding fully qualified names. The generated fully qualified names will be paired up and output as API pairs. It should be noted that this embodiment assumes that if two APIs are mentioned in a piece of text, there may be a certain relationship between these two APIs.

[0038] Secondly, API knowledge mining is carried out. The API knowledge mining unit enhances the knowledge of API pairs through the LLM, extracts relevant knowledge for each pair of APIs, and generates comprehensive paragraphs. This knowledge covers six API relationship types: efficiency comparison, function similarity, function collaboration, function replacement, logical constraint, and behavior difference. For each relationship type, the API knowledge mining unit designs specific prompts to guide the LLM to obtain the required knowledge. For example, for the function similarity relationship, the prompt asks about the main use of each API; for the efficiency comparison relationship, the prompt asks about the performance of the API. This knowledge will be integrated and provided to the subsequent code optimization process.

[0039] For example, when performing API knowledge extraction to infer the function similarity relationship between APIs, in the prompt of the API knowledge mining unit, the task is described as "answering questions about API knowledge". The prompt for the function similarity relationship emphasizes the usage knowledge of each API in the API pair by asking "What is the main use of {{API}}?". Other relationships are also guided by relevant knowledge, such as behavior knowledge for behavior differences ("What are the characteristics of {{API}}?"), performance knowledge for efficiency comparison ("How is the performance of {{API}}?"), input / output (I / O) knowledge for function collaboration ("What are the input and output of {{API}}?"), conditional knowledge for logical constraints ("What should be done before and after using {{API}}?"), and the prompt for the function replacement relationship focuses on usage scenario knowledge, asking "When to use {{API1}}?" and "When not to use {{API2}}?".

[0040] Next, API relationship inference is carried out. After extracting and enhancing API knowledge, the next step is to infer the relationship between API pairs. The API relationship inference unit uses the LLM to infer the most appropriate API relationship based on the input API pairs and their related knowledge. The system presets six relationship types and also supports the "unknown" option, which is selected when the relationship cannot be determined. Each inference is based on the API pair and its corresponding knowledge, and the LLM will select the relationship type that best fits the context.

[0041] In the prompt design of the API relationship inference unit, the prompts used in this embodiment guide the LLM to select the most accurate relationship between two APIs from six options and also provide the "unknown" option if the relationship cannot be determined. The prompt includes five examples of API pairs and their related knowledge as the input for the LLM to make inferences. The output is the inferred relationship, which processes the API pair and its related knowledge and outputs in a given format. Any relationship not in the provided options or marked as unknown is considered irrelevant.

[0042] Step S103: Store the API architecture features in the form of triples through the knowledge graph unit to form a knowledge base; In a possible implementation, use the knowledge graph as a carrier to store the diverse API architecture features extracted. The knowledge organization form in which the API architecture features are stored is an API triple composed of API entities and 6 types of API non-functional relationships. Its specific form is: (API entity 1, API non-functional relationship, API entity 2); where the API entity name is the API fully qualified name, and the fully qualified name includes the package, class, method, and parameters to which the API belongs, in the form of: "package name.class name.method name (parameters)". And the API non-functional relationships are divided into six types, and the definition of each type is as follows: Function Similarity: Two different API entities have similar usage and behavior; Function Replace: Under special circumstances, API1 can be replaced by API2; Function Collaboration: Two API entities collaborate to execute the same event; Logic Constraint: One API entity must logically depend on another API entity; Behavior Difference: Two different API entities have similar usage, but there are differences in behavior; Efficiency Comparison: Two API entities with similar functions have differences in execution efficiency.

[0043] Step S104: Use the knowledge retriever to extract the APIs to be retrieved in the code to be optimized; Specifically, according to the declarations and method instances in the code to be optimized, extract the class APIs to be retrieved and the method APIs to be retrieved respectively; use the extraction unit to extract the fully qualified names to be retrieved of the method APIs to be retrieved; among them, both the class APIs to be retrieved and the method APIs to be retrieved are presented in the form of fully qualified names.

[0044] In a possible implementation, based on the knowledge base, use the knowledge retriever to accurately retrieve relevant API practices for the code to be optimized, including two strategies: design an API extraction strategy to automatically extract the APIs in the code according to the problem code input by the user; design a triple retrieval strategy to search for API triples related to the API from the API knowledge base based on the extracted API.

[0045] In the extraction strategy, the input is a piece of problem code, and the output is an API list containing multiple Java APIs. The input problem code is specifically composed of four parts: Usage statements: Include all import statements in the code. This part observes whether there are any additions or deletions to the import statements when introducing or replacing API usage.

[0046] Key methods: Represent the core optimized parts compared to the target code, i.e., the places where API usage has changed.

[0047] Class attributes: These are the member attributes of the class. Introducing new API usage may require declaring a new class, thus leading to changes in class attributes.

[0048] Caller - callee methods: These are the methods that directly call the key methods or are called by the key methods.

[0049] Method signatures: This part includes other methods that have no direct call relationship with the key methods.

[0050] For example, the knowledge retriever dynamically retrieves API knowledge that matches the code from the database based on the user's input question code, providing a knowledge reserve for optimizing the code, including two processes: API extraction and API knowledge retrieval.

[0051] In the API extraction process, the code contains custom APIs and basic APIs. Custom APIs refer to newly created classes or methods in the code project, while basic APIs come from various API documents such as Java JDK, Apache Common, and Joda - Time. In the optimization process, the focus of this embodiment is to optimize the use of basic APIs. This choice is made because custom APIs lack clear API relationships, and optimizing basic APIs also indirectly optimizes custom APIs since the latter essentially includes the use of basic APIs. APIs are divided into class - level and method - level, and the extraction methods for each type are as follows: Class - level: All class - level APIs are declared in the usage statements of the code. Those APIs declared in the format of java.{package}.{class} are considered basic APIs, while the others belong to custom APIs.

[0052] Method - level: Method - level APIs usually appear in the code in the form of {class}.{method}({parameters}), indicating the call to the {method} method of the API. Therefore, this embodiment first extracts all instances of {class}.{method}({parameters}) in the code. Subsequently, it further filters out those cases where {class} belongs to a class - level custom API, because in this case, {method} is also considered a custom API. The remaining {method} instances represent method - level basic APIs.

[0053] At this stage, the class-level APIs already extracted in this embodiment are presented in fully qualified names, while the method-level APIs are currently in simple name format. To solve this problem, this embodiment uses the fully qualified name inference method detailed in step S102 to infer the fully qualified names of method-level APIs.

[0054] Step S105: Obtain triples and corresponding explanatory texts that match the API to be retrieved from the knowledge base to generate retrieval information; Specifically, according to the entity name of the API to be retrieved, obtain the corresponding paired API information from the knowledge base; among them, the paired API information includes triples and corresponding explanatory texts; use a large language model to calculate the similarity between the API to be retrieved and the paired API information; according to the similarity, screen the paired API information to obtain retrieval information.

[0055] In a possible implementation manner, the second strategy for accurately retrieving relevant APIs in the practice of optimizing code by the knowledge retriever is the triple retrieval strategy.

[0056] In the triple retrieval strategy, the input is the fully qualified name of a single Java API, and the output is an API triple, which consists of three parts and is in the form of: <API entity 1, one of the 6 types of API non-functional relationships, API entity 2>. API entity 1 and API entity 2 represent two Java APIs with different fully qualified names.

[0057] For example, during the API knowledge retrieval process, after API extraction, the naming format consistency between the APIs in the code and the API entities in the knowledge base is maintained. Through this consistency, triples and their related explanatory texts can be directly retrieved from the API knowledge base by matching API names. Then, use the similarity calculation mode of the large model GPT-4o to calculate the similarity between the two matched APIs, and retain the part with a similarity above 80%.

[0058] Each retrieved triple contains at least one API entity that matches the API in the code. These triples will be used as structured API knowledge for subsequent code optimization processes.

[0059] Step S106: Convert the retrieval information into actionable knowledge through a knowledge injector; where the actionable knowledge includes entities, relationships, instructions, and rules corresponding to the retrieval information; In a possible implementation manner, design a knowledge extension strategy based on API triples. In the knowledge extension strategy, the input is the API triple, and the output is actionable knowledge; The actionable knowledge consists of four API elements: API Entity: This element is the fully qualified name of the Java API included in the API triple; API Relationship: This element is one of the six types of API non-functional relationships; Instruction: This element specifies how to operationalize the use of API knowledge; Rule: This element specifies the usage scenarios for operationalizing API knowledge.

[0060] Among them, the rule element consists of static rules and dynamic rules, a total of nine. The specific rule content is as follows: Static Rules: Rule 1: The optimized API cannot be limited to one.

[0061] Rule 2: Not all API relationships are useful; only consider optimizations relevant to the current scenario.

[0062] Dynamic Rules: Rule 3: In efficiency comparison, function similarity, function replacement, or behavior difference relationships, consider using the API to replace the one in the code.

[0063] Rule 4: In function collaboration and logical constraint relationships, APIs can complement each other.

[0064] Rule 5: When there are multiple efficiency comparison relationships, prefer the API with the highest efficiency.

[0065] Rule 6: When an API has a function collaboration relationship with the API in the code, only consider the most efficient one.

[0066] Rule 7: When using function collaboration relationships, ensure that the original intention of the code is not changed.

[0067] Rule 8: When using function similarity and function replacement relationships, maintain the original intention of the code.

[0068] Rule 9: When both behavior difference relationships and function similarity relationships exist, prefer the API in the behavior difference relationship.

[0069] Step S107: Transmit the operationalizable knowledge to the large language model to generate and output multiple code optimization results.

[0070] Specifically, transform the operationalizable knowledge into structured prompts; input the structured prompts into the large language model to output multiple code optimization results.

[0071] In a possible implementation, a structured prompt-guided knowledge injection strategy is designed to effectively utilize operationalizable knowledge to enhance the generation ability of large models and achieve the diversity and accuracy of optimized code generated by large models.

[0072] The structured prompts used in the knowledge expansion strategy consist of a total of seven tags: @persona, @context-control{}, @terminology{}, @instruction, @command{}, @rule{}, and @format.

[0073] Among them, the @terminology{} and @context-control{} tags are located at the outermost layer of the structured prompt, providing a wide range of knowledge references across multiple functional areas. The @command{} and @rule{} tags are nested within the @instruction{} tag, representing an independent functional area. @context-control{}, @terminology{}, @command{}, and @rule{} are used to store the content of API knowledge. The API triples and elements for operating API knowledge are divided into four components: API triples, API entities and relationships, usage, and rules. These components are incorporated into the @context-control{}, @terminology{}, @command{}, and @rule{} tags respectively.

[0074] Furthermore, the optimization directions for the optimized code include: Usability: human factors, aesthetics, consistency, documentation; Reliability: frequency / severity of failures, recoverability, predictability, accuracy, mean time to failure; Performance: speed, efficiency, resource consumption, throughput, response time; Supportability: testability, scalability, adaptability, maintainability, compatibility, configurability, serviceability, installability, localization, portability.

[0075] For example, through the API knowledge expansion strategy and the knowledge injection strategy, API knowledge is applied to guide the LLM for code optimization. The API knowledge expansion strategy provides operational API knowledge to the LLM, while the knowledge injection strategy connects the operational API knowledge with the LLM through structured prompt design. The key to this module is to help the LLM understand and effectively utilize API knowledge through systematic API knowledge, including API triples, API entities and relationships, usage, and rules, so as to generate diverse code optimization suggestions.

[0076] To maximize the utility of operational API knowledge, this embodiment designs an injection strategy based on program structure, and effectively injects API knowledge into the LLM through seven tags of structured prompts (such as @persona, @context-control, @terminology, etc.). These tags correspond to different types of API knowledge, such as API triples, entities and relationships, usage and rules, where the @command and @rule tags are nested within the @instruction tag to form a clear functional area. This design ensures that the LLM can accurately execute relevant commands and improve the quality of optimization results through a hierarchical and clear prompt structure.

[0077] The API code optimization method based on the large language model provided by the above embodiment defines an operationalizable API knowledge form, including four API elements: API entities, relationships, usage, and rules. This knowledge enhances the large language model's understanding of the API and enables it to more effectively guide code optimization. Exploring and defining this knowledge format for specific elements in domain knowledge serves as a valuable reference for the large language model when applying knowledge across different domains; it also constructs a method for using operationalizable API knowledge to guide the large language model in code optimization. This method establishes a comprehensive pipeline from building a knowledge base to retrieving and applying this knowledge. This method effectively provides diverse results for the optimization of target code; it also designs a strategy for the large language model to effectively utilize API knowledge. This strategy uses structured prompts as a bridge to dynamically inject API knowledge to guide the large language model in optimizing code. In addition to code optimization, this strategy also has the potential to provide technical guidance for large language models based on domain knowledge and can be applied to other generation tasks.

[0078] It should be noted that for the API code optimization method based on the large language model provided in the embodiments of the present application, the execution subject can be an API code optimization device based on the large language model, or a control module in the API code optimization device based on the large language model for executing the API code optimization method of loading the large language model. In the embodiments of the present application, taking the API code optimization device based on the large language model executing the API code optimization method of loading the large language model as an example, the method of the API code optimization device based on the large language model provided in the embodiments of the present application is described.

[0079] Figure 2 is a schematic diagram of an API code optimization device based on the large language model shown in the second embodiment of the present application. Please refer to Figure 2 , the API code optimization device 200 based on the large language model includes: Unit module 201: Construct a knowledge construction module including a text collection unit, an extraction unit, and a knowledge graph unit; Sample module 202: Obtain code samples through a text collection unit, and extract API architecture features of the code samples using an extraction unit; among them, the API architecture features include entity samples and non-functional relationship samples; Specifically, use the text collection unit to collect initial texts according to screening criteria; among them, the screening criteria at least include: the initial text includes at least one feasible code solution, the feasible code solution includes at least one API, and the scoring value of the feasible code solution reaches a preset score; split the initial text into multiple paragraphs to form code samples.

[0080] Specifically, it further includes: extracting API samples from the code samples; extracting fully qualified name samples of the API samples; inputting the fully qualified name samples into a large language model to generate related knowledge samples of the API samples; determining relationship samples of the API samples based on the related knowledge samples.

[0081] Among them, extracting the fully qualified name samples of the API samples specifically includes: extracting simple name pairs of the API samples; among them, the simple name pairs are composed of a class name and a method name or a single method name; complete the simple name pairs to obtain fully qualified name samples.

[0082] Among them, inputting the fully qualified name samples into a large language model to generate related knowledge samples of the API samples specifically includes: presetting multiple API relationship types; among them, the multiple API relationship types include: efficiency comparison, function similarity, function collaboration, function replacement, logical constraint, and behavior difference; generating corresponding knowledge mining prompts for the multiple API relationship types; based on the knowledge mining prompts, inputting the fully qualified name samples into the large language model to generate knowledge samples.

[0083] In a possible implementation manner, start with a piece of code containing four Java APIs: stream, filter, forEach, and add. After extracting these APIs, their fully qualified names are obtained: java.util.Arrays.stream, java.util.stream.Stream.filter, java.util.stream.Stream.forEach, and java.util.ArrayList.add.

[0084] Knowledge graph module 203: Store the API architecture features in the form of triples through a knowledge graph unit to form a knowledge base; Retrieval and extraction module 204: Use a knowledge retriever to extract the APIs to be retrieved in the code to be optimized; Specifically, according to the declarations and method instances in the code to be optimized, the API of the class to be retrieved and the API of the method to be retrieved are extracted respectively; using the extraction unit, the fully qualified name to be retrieved of the API of the method to be retrieved is extracted; wherein, both the API of the class to be retrieved and the API of the method to be retrieved are presented in the form of fully qualified names.

[0085] Knowledge retrieval module 205: Obtain triples and corresponding explanatory texts that match the API to be retrieved from the knowledge base to generate retrieval information; Specifically, according to the entity name of the API to be retrieved, obtain the corresponding paired API information from the knowledge base; wherein, the paired API information includes triples and corresponding explanatory texts; use the large language model to calculate the similarity between the API to be retrieved and the paired API information; according to the similarity, screen the paired API information to obtain retrieval information.

[0086] Retrieve API knowledge from the knowledge base and find four triples related to java.util.stream.Stream.forEach and java.util.Arrays.stream. These triples are marked by colors: yellow, green, purple, and blue.

[0087] Knowledge injection module 206: Convert the retrieval information into operationalizable knowledge through a knowledge injector; wherein, the operationalizable knowledge includes entities, relationships, instructions, and rules corresponding to the retrieval information; In a possible implementation manner, after having API knowledge, this embodiment clarifies it through terms, commands, and rules. Two API relationships and four different APIs are extracted from the four triples for the API entities and relationships, providing two optimization perspectives and generating four suggestions. The usage is used to call terms, API knowledge, and rules. The rules include four API knowledge usage rules. These rules include two static rules, marked as Rule 1 and Rule 4 respectively, and two dynamic rules, represented by Rule 2 and Rule 3 respectively. These dynamic rules are derived based on the "efficiency comparison" and "behavior difference" relationships.

[0088] Optimized code output module 207: Transmit the operationalizable knowledge to the large language model to generate and output multiple code optimization results.

[0089] Specifically, convert the operationalizable knowledge into structured prompts; input the structured prompts into the large language model to output multiple code optimization results.

[0090] In a possible implementation, the integration of API entities, relationships, usages, and rules creates operational API knowledge. Through the knowledge application component, a prompt containing this operational API knowledge is created. When this prompt is input into the large model, it generates four code optimization suggestions that perfectly fit four triples and contribute to diverse code optimization.

[0091] The API code optimization device based on the large language model in the embodiments of the present application can be a device, or a component, an integrated circuit, or a chip in a terminal. The device can be a mobile electronic device or a non-mobile electronic device. Exemplarily, the mobile electronic device can be a mobile phone, a tablet computer, a laptop computer, a handheld computer, a vehicle-mounted electronic device, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc., and the non-mobile electronic device can be a server, a Network Attached Storage (NAS), a personal computer (PC), a television (TV), a teller machine, or a self-service machine, etc. The embodiments of the present application do not make specific limitations.

[0092] The API code optimization device based on the large language model in the embodiments of the present application can be a device with an operating system. The operating system can be the Android operating system, the iOS operating system, or other possible operating systems. The embodiments of the present application do not make specific limitations.

[0093] The API code optimization device provided in the embodiments of the present application can implement Figure 1 each process implemented by the API code optimization device in the method embodiments. To avoid repetition, it will not be elaborated here.

[0094] Optionally, please refer to Figure 3 , the embodiments of the present application further provide an electronic device 300, including a processor 301, a memory 302, and a computer program 303 stored on the memory 302 and executable on the processor 301. When the computer program 303 is executed by the processor 301, it implements each process of the above-mentioned API code optimization method embodiments based on the large language model and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.

[0095] An embodiment of the present application also provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, it implements each process of the above embodiment of the API code optimization method based on a large language model and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0096] Wherein, the processor is the processor in the electronic device described in the above embodiment. The readable storage medium includes a computer-readable storage medium, such as a computer read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, etc.

[0097] Another embodiment of the present application provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run a program or instruction to implement each process of the above embodiment of the API code optimization method based on a large language model and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0098] It should be understood that the chip mentioned in the embodiment of the present application may also be referred to as a system-on-chip, a system chip, a chip system, or a system-on-a-chip, etc.

[0099] It should be noted that in this article, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or device. Without more limitations, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in a reverse order according to the functions involved. For example, the described methods may be performed in an order different from that described, and various steps may be added, omitted, or combined. In addition, the features described with reference to certain examples may be combined in other examples.

[0100] Through the description of the above embodiments, those skilled in the art can clearly understand that the above method of the embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in various embodiments of the present application.

[0101] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims, and all of them belong to the protection scope of the present application.

Claims

1. An API code optimization method based on large language models, characterized in that, Including: Construct a knowledge construction module including a text collection unit, an extraction unit, and a knowledge graph unit; Obtain code samples through the text collection unit, and extract API architecture features of the code samples by using the extraction unit; wherein, the API architecture features include entity samples and non-functional relationship samples; Store the API architecture features in the form of triples through the knowledge graph unit to form a knowledge base; Use a knowledge retriever to extract the API to be retrieved of the code to be optimized; Obtain triples and corresponding explanatory texts matching the API to be retrieved from the knowledge base to generate retrieval information; Convert the retrieval information into operationalizable knowledge through a knowledge injector; wherein, the operationalizable knowledge includes entities, relationships, instructions, and rules corresponding to the retrieval information; Transmit the operationalizable knowledge to a large language model to generate and output multiple code optimization results.

2. The API code optimization method based on a large language model according to claim 1, characterized in that The specific steps of obtaining code samples through the text collection unit and extracting API architecture features of the code samples by using the extraction unit include: Use the text collection unit to collect initial texts according to screening criteria; wherein, the screening criteria at least include: there is at least one feasible code solution in the initial text, there is at least one API in the feasible code solution, and the scoring value of the feasible code solution reaches a preset score; Split the initial text into multiple paragraphs to form the code samples.

3. The API code optimization method based on the large language model according to claim 1, characterized in that The specific steps of obtaining code samples through the text collection unit and extracting API architecture features of the code samples by using the extraction unit further include: Extract API samples from the code samples; Extract fully qualified name samples of the API samples; Input the fully qualified name samples into a large language model to generate relevant knowledge samples of the API samples; Determine relationship samples of the API samples based on the relevant knowledge samples.

4. The API code optimization method based on the large language model according to claim 3, wherein The specific steps of extracting fully qualified name samples of the API samples include: Extract simple name pairs of the API samples; wherein, the simple name pairs are composed of a class name and a method name or a single method name; Complete the simple name pairs to obtain the fully qualified name samples.

5. The API code optimization method based on the large language model according to claim 3, wherein The specific steps of inputting the fully qualified name samples into a large language model to generate relevant knowledge samples of the API samples include: Preset multiple API relationship types; wherein, the multiple API relationship types include: efficiency comparison, function similarity, function collaboration, function replacement, logical constraint, and behavior difference; Generate corresponding knowledge mining prompts for the multiple API relationship types; Based on the knowledge mining prompts, input the fully qualified name samples into the large language model to generate the knowledge samples.

6. The API code optimization method based on the large language model according to claim 1, characterized in that The specific steps of using a knowledge retriever to extract the API to be retrieved of the code to be optimized include: Extract the API to be retrieved of the class and the API to be retrieved of the method respectively according to the declarations and method instances in the code to be optimized; Using the extraction unit, extract the fully qualified name to be retrieved of the method API to be retrieved; wherein, both the class API to be retrieved and the method API to be retrieved are presented in the form of fully qualified names.

7. The API code optimization method based on a large language model according to claim 1, wherein The specific steps of obtaining the triples and corresponding explanatory texts matching the API to be retrieved from the knowledge base to generate retrieval information include: According to the entity name of the API to be retrieved, obtain the corresponding paired API information from the knowledge base; wherein, the paired API information includes triples and corresponding explanatory texts; Use a large language model to calculate the similarity between the API to be retrieved and the paired API information; According to the similarity, filter the paired API information to obtain the retrieval information.

8. The API code optimization method based on a large language model according to claim 1, wherein The specific steps of transmitting the operationalizable knowledge to the large language model to generate and output multiple code optimization results include: Convert the operationalizable knowledge into structured prompts; Input the structured prompts into the large language model to output the multiple code optimization results.

9. An API code optimization device based on a large language model, which is used to execute the API code optimization method based on a large language model according to any one of claims 1-8, characterized in that Include: A unit module for constructing a knowledge construction module including a text collection unit, an extraction unit, and a knowledge graph unit; A sample module for obtaining code samples through the text collection unit and using the extraction unit to extract the API architecture features of the code samples; wherein, the API architecture features include entity samples and non-functional relationship samples; A knowledge graph module for storing the API architecture features in the form of triples through the knowledge graph unit to form a knowledge base; A retrieval and extraction module for using a knowledge retriever to extract the API to be retrieved of the code to be optimized; A knowledge retrieval module for obtaining the triples and corresponding explanatory texts matching the API to be retrieved from the knowledge base to generate retrieval information; A knowledge injection module for converting the retrieval information into operationalizable knowledge through a knowledge injector; wherein, the operationalizable knowledge includes entities, relationships, instructions, and rules corresponding to the retrieval information; An optimized code output module for transmitting the operationalizable knowledge to the large language model to generate and output multiple code optimization results.

10. An electronic device, characterized in that, Include: A memory, a processor, and a program or instruction stored on the memory and executable on the processor, and when the program or instruction is executed by the processor, it performs the steps of the API code optimization method based on a large language model according to any one of claims 1-8.

Citation Information

Patent Citations

  • Code abstract generation method based on code knowledge graph and knowledge migration

    CN111797242A

  • API relation reasoning method and system based on large pre-training language model

    CN116776981A

  • API (Application Program Interface) task demand processing method, browser access method and related device

    CN119045817A

  • Process industry safety knowledge graph error detection method and system based on large language model

    CN119740644A

  • Systems and methods for question answering with diverse knowledge sources

    US20250103592A1