Code snippet type inference method and device integrating constraints and statistics
By integrating constraints and statistics to infer code snippet types, the problem of high code quality requirements and low accuracy in the existing technology is solved, and code type inference with high applicability and high accuracy is achieved.
Patent Information
- Application Number
- CN202410070463.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-17
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2044-01-17
AI Technical Summary
Existing constraint-based code snippet type inference methods have high requirements on code quality and limited adaptability, while statistical-based methods lack specificity and have low accuracy.
A code snippet type inference method that integrates constraints and statistics performs initial type analysis and context enhancement through a constraint-based method, combines it with a statistics-based method for type prediction, and obtains optimized constraints and statistical inference types by iteratively optimizing the API knowledge base.
The accuracy and applicability of code snippet type inference are improved, the dependence on code quality is reduced, and the reliability and efficiency of inference results are improved.
Smart Images

Figure CN117891951B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of text recognition, and in particular to a method, apparatus, and device for inferring the type of code snippets by integrating constraints and statistics. Background Art
[0002] An Application Programming Interface (API) is a programming interface that encapsulates certain application functions. An API library (such as the Java JDK) contains different types of API elements, such as classes, interfaces, methods, and fields, as well as relationships between API elements, such as inheritance relationships between classes, implementation relationships between classes and interfaces, and inclusion relationships between classes and methods / fields. The library that stores different types of API elements and their relationships is called an API knowledge base. The fully qualified name (FQN) of an API element is used to uniquely identify the type of an API element. The FQN consists of the element name and its structural hierarchy in the library. Since elements with the same name may exist in different API knowledge bases, the structural hierarchy is also required to uniquely identify them, so that the FQN can uniquely identify the information and source of the current element.
[0003] Type inference in code snippets involves identifying the types of API elements within a code snippet. Among these element types, class and interface inference is most important, as other element knowledge can be further inferred from these classes and interfaces based on the relationships between API elements. Existing type inference methods can be broadly divided into two categories: constraint-based and statistically-based.
[0004] Constraint-based type inference methods have high accuracy, but to extract type constraints from code snippets, they require the input code snippet to be successfully compiled, meaning it must be free of lexical, syntactic, or type errors. Therefore, this method has limited applicability and struggles with uncompilable code snippets, resulting in low recall. Statistical type inference methods do not require the extraction of type constraints from code snippets. While this method has good applicability and can handle all code snippets, it lacks specificity, resulting in lower accuracy. Summary of the Invention
[0005] The present application provides a code snippet type inference method and device that integrates constraints and statistics, which are used to solve the technical problems that existing constraint-based methods have high requirements on code quality and limited adaptability, and statistics-based methods lack specificity and have low accuracy.
[0006] In view of this, the first aspect of the present application provides a code snippet type inference method that integrates constraints and statistics, including:
[0007] The first constraint-based method is used to perform type analysis on the current code snippet according to the preset type constraints and the initial API knowledge base to obtain the initial constraint inference type;
[0008] Performing enhancement processing on the context of the current code snippet using the initial constraint inference type to obtain an enhanced code sequence;
[0009] Performing type analysis on the enhanced code sequence using a second statistically-based method to obtain an element candidate type list, wherein each API element corresponds to one element candidate type list, and the element candidate type list includes multiple element candidate types;
[0010] Replacing the initial API knowledge base with a reduced API knowledge base, returning to the step of performing type analysis on the current code snippet using the first constraint-based method according to preset type constraints and the initial API knowledge base, until an iteration termination condition is satisfied, thereby obtaining an optimized constraint inference type and an optimized statistical inference type, wherein the reduced API knowledge base is obtained by performing an API knowledge base reduction operation on the element candidate type list;
[0011] The type of the code snippet is determined according to the optimization constraint inference type and the optimization statistical inference type to obtain a target inference type.
[0012] Preferably, the enhancing the context of the current code snippet by using the initial constraint inference type to obtain an enhanced code sequence includes:
[0013] Extracting an API element set from the current code snippet;
[0014] Dividing the API element set into an inferred sub-set and a non-inferred sub-set according to the initial constraint inference type;
[0015] The context enhancement of the current code snippet is achieved by replacing the inferred subset with a preset inference type, and an enhanced code sequence is generated by combining the non-inferred subset.
[0016] Preferably, the step of replacing the initial API knowledge base with the simplified API knowledge base and returning to the step of performing type analysis on the current code snippet according to the preset type constraints and the initial API knowledge base using the first constraint-based method until an iteration termination condition is satisfied further comprises:
[0017] Collecting a constraint candidate type set and a statistical candidate type set in the element candidate type list based on classes and interfaces;
[0018] API knowledge is screened in the initial API knowledge base according to the constraint candidate type set and the statistical candidate type set, so as to simplify the API knowledge base and obtain a simplified API knowledge base.
[0019] Preferably, the method of collecting the constraint candidate type set and the statistical candidate type set in the element candidate type list based on the class and the interface further includes:
[0020] Filtering the element candidate types in the element candidate type list using type probability to obtain a filtered candidate type list;
[0021] Then, the collecting of the constraint candidate type set and the statistical candidate type set in the element candidate type list based on the class and the interface includes:
[0022] A constraint candidate type set and a statistical candidate type set are collected in the filtering candidate type list based on classes and interfaces.
[0023] Preferably, determining the type of the code snippet according to the optimization constraint inference type and the optimization statistical inference type to obtain a target inference type includes:
[0024] If the optimization constraint inference type is a null value, determining the optimization statistical inference type as the target inference type;
[0025] If the optimization statistical inference type is a null value, determining the optimization constraint inference type as the target inference type;
[0026] If both the optimization constraint inference type and the optimization statistical inference type are non-null values, the optimization constraint inference type is selected as the target inference type.
[0027] A second aspect of the present application provides a code snippet type inference device that integrates constraints and statistics, including:
[0028] A constraint inference unit, configured to perform type analysis on a current code snippet according to preset type constraints and an initial API knowledge base using a first constraint-based method to obtain an initial constraint inference type;
[0029] A code enhancement unit, configured to enhance the context of the current code snippet using the initial constraint inference type to obtain an enhanced code sequence;
[0030] a statistical inference unit, configured to perform type analysis on the enhanced code sequence using a second statistically-based method to obtain an element candidate type list, wherein each API element corresponds to one element candidate type list, and the element candidate type list includes a plurality of element candidate types;
[0031] a reduction iteration unit, configured to replace the initial API knowledge base with a reduced API knowledge base, trigger the constraint inference unit, and obtain an optimized constraint inference type and an optimized statistical inference type until an iteration termination condition is satisfied, wherein the reduced API knowledge base is obtained by performing an API knowledge base reduction operation based on the element candidate type list;
[0032] A type determination unit is used to determine the type of the code snippet according to the optimization constraint inference type and the optimization statistical inference type to obtain a target inference type.
[0033] Preferably, the code enhancement unit is specifically used to:
[0034] Extracting an API element set from the current code snippet;
[0035] Dividing the API element set into an inferred sub-set and a non-inferred sub-set according to the initial constraint inference type;
[0036] The context enhancement of the current code snippet is achieved by replacing the inferred subset with a preset inference type, and an enhanced code sequence is generated by combining the non-inferred subset.
[0037] Preferably, it also includes:
[0038] A candidate collection unit, configured to collect a constraint candidate type set and a statistical candidate type set in the element candidate type list based on classes and interfaces;
[0039] The reduction and screening unit is used to screen API knowledge in the initial API knowledge base according to the constraint candidate type set and the statistical candidate type set, thereby reducing the API knowledge base and obtaining a reduced API knowledge base.
[0040] Preferably, it also includes:
[0041] a type filtering unit, configured to filter the element candidate types in the element candidate type list using type probabilities to obtain a filtered candidate type list;
[0042] Then, the candidate collection unit is specifically used to:
[0043] A constraint candidate type set and a statistical candidate type set are collected in the filtering candidate type list based on classes and interfaces.
[0044] Preferably, the type determination unit is specifically configured to:
[0045] If the optimization constraint inference type is a null value, determining the optimization statistical inference type as the target inference type;
[0046] If the optimization statistical inference type is a null value, determining the optimization constraint inference type as the target inference type;
[0047] If both the optimization constraint inference type and the optimization statistical inference type are non-null values, the optimization constraint inference type is selected as the target inference type.
[0048] It can be seen from the above technical solutions that the embodiments of the present application have the following advantages:
[0049] In the present application, a code snippet type inference method that integrates constraints and statistics is provided, including: using a first constraint-based method to perform type analysis on the current code snippet according to preset type constraints and an initial API knowledge base to obtain an initial constraint inference type; enhancing the context of the current code snippet through the initial constraint inference type to obtain an enhanced code sequence; using a second statistics-based method to perform type analysis on the enhanced code sequence to obtain an element candidate type list, wherein each API element corresponds to an element candidate type list, and the element candidate type list includes multiple element candidate types; replacing the initial API knowledge base with a simplified API knowledge base, returning to the step of performing type analysis on the current code snippet according to the preset type constraints and the initial API knowledge base using the first constraint-based method until the iteration termination condition is met, thereby obtaining an optimized constraint inference type and an optimized statistical inference type, wherein the simplified API knowledge base is obtained by performing an API knowledge base simplification operation based on the element candidate type list; determining the type of the code snippet according to the optimized constraint inference type and the optimized statistical inference type to obtain a target inference type.
[0050] The code snippet type inference method that integrates constraints and statistics provided by this application combines the constraint-based inference method and the statistics-based inference method. It does not require the code to have high-quality grammatical characteristics and is basically applicable without restrictions; it can also take into account type constraint knowledge to ensure the accuracy of the inference results; and it adds a code snippet context enhancement operation to strengthen the context relevance of the code snippet and improve the reliability of the inference results; in addition, the operation of simplifying the API knowledge base can remove redundant knowledge in the knowledge base, thereby improving the efficiency of code snippet type inference. Therefore, this application can solve the technical problems that the existing constraint-based method has high requirements for code quality and has limited adaptability, and the statistical-based method lacks pertinence and has a low accuracy rate. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 A flowchart of a code snippet type inference method that integrates constraints and statistics provided in an embodiment of the present application;
[0052] Figure 2 A schematic diagram of the structure of a code snippet type inference device integrating constraints and statistics provided in an embodiment of the present application;
[0053] Figure 3 An example diagram of API element type inference for a code snippet provided in an embodiment of the present application;
[0054] Figure 4 Schematic diagram of a code snippet type inference framework that integrates constraints and statistics provided in an embodiment of the present application;
[0055] Figure 5 An example graph of the average recall rate data of the four methods provided for the application example of this application on two datasets;
[0056] Figure 6 An example graph of the average accuracy of the four methods provided in this application example on two data sets;
[0057] Figure 7 An example graph of the standard deviation of recall rates for all code snippets in two datasets for the four methods provided in this application example. DETAILED DESCRIPTION
[0058] In order to help those skilled in the art better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0059] Explanation of terms:
[0060] Fully Qualified Name (FQN) of an API element: The fully qualified name of an API element consists of its name and its hierarchy within the API library, such as org.joda.time.DateTime, where DateTime is the name of a class and org.joda.time is the package hierarchy within the API library. Note that different API libraries may contain API elements with the same name. The FQN of an API element uniquely identifies its type, indicating which API library it comes from.
[0061] Code Snippet (CS): A piece of program code written in a specific programming language (such as Java) that is often used to demonstrate the use of an API element. Code snippets are usually incomplete, with only the name of the API element displayed, not the Full Qualification Number (FQN).
[0062] Type Constraint: In a code snippet, although the API element does not specify the FQN, it will contain constraint information (referred to as "type constraint") that can be used to infer the type of the API element (i.e., FQN), such as the method called by a variable of a class, the number and type of parameters accepted by a method, etc.
[0063] Type Inference: Inferring the types of API elements based on the type constraints in a code snippet. Within a code snippet, there are five types of API elements that require type inference: classes, interfaces, methods, fields, and variables. Of these elements, type inference for classes and interfaces is the most important, as the types of other elements can be inferred from the class / interface types based on the relationships between API elements. For example, after determining the FQN of a class, the FQN of its variables is the same as the class's FQN. The FQN of the methods called by the variables or the FQN of the fields accessed by the variables can be determined from the methods or fields contained in the class.
[0064] Abstract Syntax Tree (AST): A tree representation of program code. Each node in the tree represents a code element, such as a statement (such as a variable declaration), an inner statement (such as a method call), or a token (such as a modifier or value).
[0065] Prompt Learning: A natural language processing technique whose core idea is to improve the output of a task by adding additional "prompt information" to the input of the task without significantly changing the structure and parameters of the pre-trained language model.
[0066] Hallucination: A common problem of generative language models (such as masked language models (MLMs)) is that the content generated by the model does not match the input content or facts.
[0067] For easier understanding, see Figure 1 The present application provides an embodiment of a code snippet type inference method that integrates constraints and statistics, including:
[0068] Step 101: Use a first constraint-based method to perform type analysis on the current code snippet according to preset type constraints and an initial API knowledge base to obtain an initial constraint inference type.
[0069] The first constraint-based method in this embodiment can be the currently optimal constraint-based type inference method, such as SnR. Baker, DepRes and other methods can also be selected according to actual conditions, and the specific ones are not limited here. The first constraint-based method uses the type constraint knowledge of API elements extracted from the code snippet, that is, the preset type constraints, and combines it with the initial API knowledge base built from the API library to perform type inference, and can obtain a preliminary inference result, that is, the initial constraint inference type. It can be understood that the current code snippet is the code text to be identified, which can be the text content obtained in real time, or it can be the existing code text, and the specific ones are not limited.
[0070] See also Figure 3 For a given current code snippet, after performing type inference using the first constraint-based method of this embodiment, a type inference result of the API element can be obtained, where the part indicated by the box is the inferred type.
[0071] Step 102: Enhance the context of the current code snippet using the initial constraint inference type to obtain an enhanced code sequence.
[0072] Furthermore, step 102 includes:
[0073] Extract the API element set in the current code snippet;
[0074] Divide the API element set into an inferred sub-set and a non-inferred sub-set according to the initial constraint inference type;
[0075] The context of the current code snippet is enhanced by replacing the inferred subset with the preset inference type, and the enhanced code sequence is generated by combining the non-inferred subset.
[0076] This implementation also requires a second statistically-based method for type prediction. This method uses a statistical language model to model and predict code snippets, capturing the correlation characteristics between code contexts. To improve the prediction accuracy of the second statistically-based method, this embodiment performs context enhancement on the current code snippet before prediction.
[0077] Each current code snippet can be regarded as a token sequence, denoted as CS=t1,t1,......,t m , where there are multiple API elements. The set of these API elements can be recorded as APIs(CS), that is, the API element set. According to the initial constraint inference type, the API elements in the API element set APIs(CS) can be divided into two categories. The API elements with inferred types constitute the inferred sub-set Typed_APIs C(CS), API elements that do not have inferred types constitute the non-inferred subset NonTyped_APIs C (CS). The code snippet context enhancement is an operation performed on the inferred sub-collection, which will be inferred sub-collection Typed_APIs C Each API element in (CS) i Replaced with the default inferred type type C (t i ), non-inferred subcollection NonTyped_APIs C The API elements in (CS) remain unchanged, and the enhanced word sequence is combined with the non-inferred subset to generate the enhanced code sequence ACS=y1,y1,......,y m ,in:
[0078]
[0079] See also Figure 3 , using the enhanced code sequence obtained after contextual enhancement for statistical language model-based type prediction results in more accurate results. Before the code snippet was contextually enhanced, the type inference result was incorrect, and the type inferred was not real, but was caused by the generative language model's hallucination problem. However, after the enhancement, the model paid more attention to and learned the relationship between the context of the code snippet, resulting in the correct type inference result, namely com.google.gwt.dom.client.Document.
[0080] Step 103: Use the second statistical-based method to perform type analysis on the enhanced code sequence to obtain an element candidate type list. Each API element corresponds to an element candidate type list, and the element candidate type list includes multiple element candidate types.
[0081] Statistical methods are based on the naturalness of software code. Human-written code is often simple and repetitive. Therefore, statistical language models can be used to capture the intertextual relevance of software code, enabling modeling and prediction. The second statistical method in this embodiment uses the MLMTyper model, which offers superior performance compared to other statistical methods. Other statistical methods, such as STATTYPE, COSTER, and RESICO, can also be used depending on the specific situation, and are not limited to these in this embodiment.
[0082] The second statistical-based method in this embodiment, namely, the MLMTyper model, can generate a candidate type list for each API element in the enhanced code sequence, namely, an element candidate type list. This list includes multiple element candidate types. Due to the existence of the hallucination problem, many of these element candidate types do not actually exist. Therefore, this embodiment needs to screen the types in the element candidate type list and retain the candidate types in the API knowledge base.
[0083] Step 104: Replace the initial API knowledge base with the simplified API knowledge base, and return to step 101 until the iteration termination condition is met to obtain the optimized constraint inference type and the optimized statistical inference type. The simplified API knowledge base is obtained by performing an API knowledge base simplification operation based on the element candidate type list.
[0084] Furthermore, before step 104, the following steps are also included:
[0085] Collect constraint candidate type sets and statistical candidate type sets in element candidate type lists based on classes and interfaces;
[0086] API knowledge is screened in the initial API knowledge base according to the constraint candidate type set and the statistical candidate type set, the API knowledge base is reduced, and a reduced API knowledge base is obtained.
[0087] Furthermore, based on classes and interfaces, constraint candidate type sets and statistical candidate type sets are collected in the element candidate type list, which also includes:
[0088] Filtering the element candidate types in the element candidate type list using the type probability to obtain a filtered candidate type list;
[0089] Then, based on the class and interface, the constraint candidate type set and the statistical candidate type set are collected in the element candidate type list, including:
[0090] The constraint candidate type set and the statistical candidate type set are collected in a filtered candidate type list based on classes and interfaces.
[0091] It should be noted that the constraint-based type inference method only needs to search for suitable types of API elements from the pre-built initial API knowledge base according to the type constraints in the code snippet. The API knowledge base needs to cover as many API libraries as possible to ensure the performance of the method; therefore, there are usually multiple candidate types of a certain API element in the API knowledge base; if the input code snippet does not contain enough type constraints or the method cannot extract enough type constraints, making it difficult to determine the type of the API element from the candidate elements, then the method will not be able to give an inference result. To address this existing problem, this embodiment uses a statistical-based type inference method to alleviate it; the purpose of the API knowledge base simplification operation is to remove redundant / irrelevant types in the knowledge base, thereby reducing the candidate type space for the constraint-based method and improving the inference performance.
[0092] In this embodiment, before the API knowledge base is simplified, the element candidate types in the element candidate type list can be filtered using type probability to obtain a filtered candidate type list. The type probability is the probability value corresponding to each element candidate type when the element candidate type list is generated. The element candidate types can be arranged in order from high to low according to the type probability. The higher the type probability, the more likely the element candidate type is to be the correct type. The element candidate types with high probability are selected and compared with the types in the API knowledge base. The existing ones are retained, and the types that do not exist in the library are discarded. In this way, the element candidate types can be screened to obtain a filtered candidate type list.
[0093] Since the most important thing in the type inference process of a code snippet is the inference of classes and interfaces, this embodiment collects the constraint candidate type set and the statistical candidate type set in the filtered candidate type list based on classes and interfaces. The obtained overall candidate type set can be expressed as:
[0094]
[0095] Among them, CIs(CS) is the class and interface collection in the current code snippet CS, Represents the top-k candidate types generated by the API element e that filters the candidate type list, Typed_APIs C (CS) is the inferred subset, type C (e) indicates the type inferred by the first constraint-based method for API element e. CanCITypes is the overall candidate type set consisting of the constraint candidate type set and the statistical candidate type set, that is, the type set inferred based on the constraint method and the type set inferred based on the statistical method.
[0096] Based on the overall candidate type set CanCITypes consisting of the constraint candidate type set and the statistical candidate type set, we can collect candidate class and interface types ctype∈CanCITypes in the initial API knowledge base AKB, as well as ctype's parent class or parent interface SupCITypes(ctype,AKB) in the AKB, and also include API knowledge such as its corresponding methods and fields. Based on this process, we can obtain the reduced API knowledge base, which is expressed as:
[0097]
[0098]
[0099] Among them, Syn_Knowl(ctype, AKB) represents the API knowledge related to type ctype in AKB, e.Methods and e.Fields respectively represent the method and field sets defined corresponding to class / interface e. It can be understood that the API elements in the generated reduced API knowledge base Reduced_AKB still maintain an association relationship.
[0100] After replacing the initial API knowledge base with the reduced API knowledge base, the process returns to step 101 and re-uses the first constraint-based method to perform type analysis on the current code snippet based on the preset type constraints and the reduced API knowledge base to obtain an updated constraint inference type. By continuously iterating, an optimized constraint inference type and an optimized statistical inference type can be obtained when the iteration stops. The optimized constraint inference type refers to the type result inferred by the first constraint-based method when the iteration stops; the optimized statistical inference type refers to the type result inferred by the second statistical-based method when the iteration stops.
[0101] This embodiment has two iteration termination conditions. The first is that the type inference result reaches a stable state, that is, the updated constraint inference type and updated statistical inference type obtained in the current iteration are within a threshold range compared to the results obtained in the previous iteration. The threshold range can be set according to actual conditions. In other words, if the type inference result does not change much and basically tends to a stable state, the iteration can be stopped. The second is to set the maximum number of iterations δ. In this embodiment, δ = 10, that is, if the number of iterations reaches 10, the iteration is stopped. If either of the two iteration termination conditions is met, the iteration can be stopped and the current optimized constraint inference type and optimized statistical inference type can be obtained.
[0102] Step 105: Determine the type of the code snippet according to the optimization constraint inference type and the optimization statistical inference type to obtain a target inference type.
[0103] Furthermore, step 105 includes:
[0104] If the optimization constraint inference type is null, the optimization statistical inference type is determined as the target inference type;
[0105] If the optimization statistical inference type is null, the optimization constraint inference type is determined to be the target inference type;
[0106] If both the optimization constraint inference type and the optimization statistical inference type are non-null values, the optimization constraint inference type is selected as the target inference type.
[0107] It is understandable that the optimization constraint inference type and the optimization statistical inference type are both the results obtained at the last iteration, which can be recorded as comb_type respectively. C , Taking an API element e as an example, the optimization constraint inference type and optimization statistical inference type are expressed as comb_type C (e),
[0108] Since the first constraint-based method and the second statistics-based method cannot always infer the type results of each element in the code snippet, this embodiment combines the two methods to determine the target inference type. If only one of the two methods obtains an inference result, this inference result is used as the target inference type; if both methods obtain corresponding inference results, the optimized constraint inference type is used as the target inference type, because the accuracy of the constraint-based method is generally higher than that of the statistics-based method. For the overall process of code snippet type inference that combines constraints and statistics in this embodiment, please refer to Figure 4 .
[0109] To evaluate the inference performance of iCSTyper, a code snippet type inference method proposed in this application that combines constraints and statistics, we selected two open datasets commonly used in type inference research: StatType-SO and Short-SO, containing 268 and 120 code snippets, respectively. Short-SO snippets are relatively short, each containing no more than three lines of code. In contrast, StatType-SO snippets are longer, with an average of 28 lines of code per snippet. Furthermore, we selected three methods, including the currently best constraint-based type inference method SnR and the statistics-based type inference method MLMTyper, as well as the mainstream large language model ChatGPT, as baselines for experimental comparison. In this comparison, precision and recall were used as evaluation metrics. Precision refers to the proportion of correct types inferred by a method; recall refers to the ratio of the number of correctly inferred types to the total number of types inferred. We also used the Wilcoxon rank sum test to test whether the performance differences between iCSTyper and each baseline method were statistically significant.
[0110] See also Figure 5 The average recall of the four methods across all code snippets in the two datasets was calculated. Our method achieved the highest average recall on both datasets. Compared to SnR and MLMTyper, our method, iCSTyper, achieved statistically significant improvements of at least 6.9% and 27.72%, respectively. Compared to ChatGPT, iCSTyper significantly improved StatType-SO by 3.68% and Short-SO by 0.84%.
[0111] See also Figure 6 , and evaluated the final average accuracy of the constraint-based type inference method (iCSTyper-C) and the statistical-based type inference method (iCSTyper-S) of the present application method iCSTyper on two data sets. The average accuracy of iCSTyper-C is improved by 1.66% and 1.64% on StatType-SO and Short-SO respectively compared with the original SnR, and the improvement on StatType-SO is statistically significant. Compared with the original MLMTyper, the average accuracy of iCSTyper-S on the two data sets is significantly improved by 20.78% and 20.01% respectively. This shows that the code snippet context enhancement and API knowledge base simplification mechanism of the present application have a promoting effect on the performance improvement of the statistical-based and constraint-based type inference methods.
[0112] See also Figure 7 We calculated the standard deviation of recall for the four methods across all code snippets in both datasets. On StatType-SO, the standard deviations for MLMTyper, SnR, ChatGPT, and iCSTyper were 0.312, 0.208, 0.147, and 0.095, respectively; on Short-SO, their standard deviations were 0.367, 0.281, 0.212, and 0.201, respectively. As can be seen, iCSTyper, our method, achieved the lowest standard deviation on both datasets, demonstrating its most stable performance.
[0113] The code snippet type inference method that integrates constraints and statistics provided by the embodiment of the present application combines the constraint-based inference method and the statistics-based inference method. It does not require the code to have high-quality grammatical characteristics and is basically applicable without restrictions; it can also take into account type constraint knowledge to ensure the accuracy of the inference results; and it adds a code snippet context enhancement operation to strengthen the context relevance of the code snippet and improve the reliability of the inference results; in addition, the operation of simplifying the API knowledge base can remove redundant knowledge in the knowledge base, thereby improving the efficiency of code snippet type inference. Therefore, the embodiment of the present application can solve the technical problems that the existing constraint-based method has high requirements for code quality and has limited adaptability, and the statistical-based method lacks pertinence and has a low accuracy rate.
[0114] For easier understanding, see Figure 2 The present application provides an embodiment of a code snippet type inference device that integrates constraints and statistics, including:
[0115] The constraint inference unit 201 is configured to perform type analysis on the current code snippet according to preset type constraints and an initial API knowledge base using a first constraint-based method to obtain an initial constraint inference type;
[0116] A code enhancement unit 202 is configured to enhance the context of the current code snippet using an initial constraint inference type to obtain an enhanced code sequence;
[0117] A statistical inference unit 203 is configured to perform type analysis on the enhanced code sequence using a second statistical method to obtain an element candidate type list, wherein each API element corresponds to an element candidate type list, and the element candidate type list includes multiple element candidate types;
[0118] A reduction iteration unit 204 is configured to replace the initial API knowledge base with the reduced API knowledge base, trigger the constraint inference unit, and obtain an optimized constraint inference type and an optimized statistical inference type until an iteration termination condition is satisfied. The reduced API knowledge base is obtained by performing an API knowledge base reduction operation based on the element candidate type list.
[0119] The type determination unit 205 is configured to determine the type of the code snippet according to the optimization constraint inference type and the optimization statistic inference type to obtain a target inference type.
[0120] Furthermore, the code enhancement unit 202 is specifically configured to:
[0121] Extract the API element set in the current code snippet;
[0122] Divide the API element set into an inferred sub-set and a non-inferred sub-set according to the initial constraint inference type;
[0123] The context of the current code snippet is enhanced by replacing the inferred subset with the preset inference type, and the enhanced code sequence is generated by combining the non-inferred subset.
[0124] Furthermore, it also includes:
[0125] A candidate collection unit 206 is configured to collect a constraint candidate type set and a statistical candidate type set in the element candidate type list based on classes and interfaces;
[0126] The reduction and screening unit 207 is used to screen the API knowledge in the initial API knowledge base according to the constraint candidate type set and the statistical candidate type set, thereby reducing the API knowledge base and obtaining a reduced API knowledge base.
[0127] Furthermore, it also includes:
[0128] a type filtering unit 208 for filtering the element candidate types in the element candidate type list using type probabilities to obtain a filtered candidate type list;
[0129] Then, the candidate collection unit 206 is specifically configured to:
[0130] The constraint candidate type set and the statistical candidate type set are collected in a filtered candidate type list based on classes and interfaces.
[0131] Furthermore, the type determination unit 205 is specifically configured to:
[0132] If the optimization constraint inference type is null, the optimization statistical inference type is determined as the target inference type;
[0133] If the optimization statistical inference type is null, the optimization constraint inference type is determined to be the target inference type;
[0134] If both the optimization constraint inference type and the optimization statistical inference type are non-null values, the optimization constraint inference type is selected as the target inference type.
[0135] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0136] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0137] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0138] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for executing all or part of the steps of the method described in each embodiment of the present application through a computer device (which can be a personal computer, server, or network device, etc.). The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (English full name: Read-Only Memory, English abbreviation: ROM), random access memory (English full name: Random Access Memory, English abbreviation: RAM), disk or optical disk and other media that can store program code.
[0139] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A code snippet type inference method integrating constraints and statistics, characterized in that: include: The first constraint-based method is used to perform type analysis on the current code snippet according to the preset type constraints and the initial API knowledge base to obtain the initial constraint inference type; Performing enhancement processing on the context of the current code snippet using the initial constraint inference type to obtain an enhanced code sequence; Performing type analysis on the enhanced code sequence using a second statistically-based method to obtain an element candidate type list, wherein each API element corresponds to one element candidate type list, and the element candidate type list includes multiple element candidate types; Replacing the initial API knowledge base with a reduced API knowledge base, returning to the step of performing type analysis on the current code snippet using the first constraint-based method according to preset type constraints and the initial API knowledge base, until an iteration termination condition is satisfied, thereby obtaining an optimized constraint inference type and an optimized statistical inference type, wherein the reduced API knowledge base is obtained by performing an API knowledge base reduction operation on the element candidate type list; The type of the code snippet is determined according to the optimization constraint inference type and the optimization statistical inference type to obtain a target inference type.
2. The code snippet type inference method integrating constraints and statistics according to claim 1 is characterized in that: The enhancing the context of the current code snippet by using the initial constraint inference type to obtain an enhanced code sequence includes: Extracting an API element set from the current code snippet; Dividing the API element set into an inferred sub-set and a non-inferred sub-set according to the initial constraint inference type; The context enhancement of the current code snippet is achieved by replacing the inferred subset with a preset inference type, and an enhanced code sequence is generated by combining the non-inferred subset.
3. The code snippet type inference method integrating constraints and statistics according to claim 1 is characterized in that: The step of replacing the initial API knowledge base with the simplified API knowledge base and returning to the step of performing type analysis on the current code snippet according to the preset type constraints and the initial API knowledge base using the first constraint-based method until an iteration termination condition is satisfied, further includes: Collecting a constraint candidate type set and a statistical candidate type set in the element candidate type list based on classes and interfaces; API knowledge is screened in the initial API knowledge base according to the constraint candidate type set and the statistical candidate type set, so as to simplify the API knowledge base and obtain a simplified API knowledge base.
4. The code snippet type inference method integrating constraints and statistics according to claim 3 is characterized in that: The method of collecting the constraint candidate type set and the statistical candidate type set in the element candidate type list based on the class and the interface further includes: Filtering the element candidate types in the element candidate type list using type probability to obtain a filtered candidate type list; Then, the collecting of the constraint candidate type set and the statistical candidate type set in the element candidate type list based on the class and the interface includes: A constraint candidate type set and a statistical candidate type set are collected in the filtering candidate type list based on classes and interfaces.
5. The code snippet type inference method integrating constraints and statistics according to claim 1 is characterized in that: The determining the type of the code snippet according to the optimization constraint inference type and the optimization statistical inference type to obtain a target inference type includes: If the optimization constraint inference type is a null value, determining the optimization statistical inference type as the target inference type; If the optimization statistical inference type is a null value, determining the optimization constraint inference type as the target inference type; If both the optimization constraint inference type and the optimization statistical inference type are non-null values, the optimization constraint inference type is selected as the target inference type.
6. A code snippet type inference device integrating constraints and statistics, characterized in that: include: A constraint inference unit, configured to perform type analysis on a current code snippet according to preset type constraints and an initial API knowledge base using a first constraint-based method to obtain an initial constraint inference type; A code enhancement unit, configured to enhance the context of the current code snippet using the initial constraint inference type to obtain an enhanced code sequence; a statistical inference unit, configured to perform type analysis on the enhanced code sequence using a second statistically-based method to obtain an element candidate type list, wherein each API element corresponds to one element candidate type list, and the element candidate type list includes a plurality of element candidate types; a reduction iteration unit, configured to replace the initial API knowledge base with a reduced API knowledge base, trigger the constraint inference unit, and obtain an optimized constraint inference type and an optimized statistical inference type until an iteration termination condition is satisfied, wherein the reduced API knowledge base is obtained by performing an API knowledge base reduction operation based on the element candidate type list; The type determination unit is used to determine the type of the code snippet according to the optimization constraint inference type and the optimization statistical inference type to obtain a target inference type.
7. The device for inferring code snippet type by integrating constraints and statistics according to claim 6, characterized in that: The code enhancement unit is specifically used to: Extracting an API element set from the current code snippet; Dividing the API element set into an inferred sub-set and a non-inferred sub-set according to the initial constraint inference type; The context enhancement of the current code snippet is achieved by replacing the inferred subset with a preset inference type, and an enhanced code sequence is generated by combining the non-inferred subset.
8. The device for inferring code snippet type by integrating constraints and statistics according to claim 6, characterized in that: Also includes: A candidate collection unit, configured to collect a constraint candidate type set and a statistical candidate type set in the element candidate type list based on classes and interfaces; The reduction and screening unit is used to screen API knowledge in the initial API knowledge base according to the constraint candidate type set and the statistical candidate type set, thereby reducing the API knowledge base and obtaining a reduced API knowledge base.
9. The device for inferring code snippet type by integrating constraints and statistics according to claim 8, characterized in that: Also includes: a type filtering unit, configured to filter the element candidate types in the element candidate type list using type probabilities to obtain a filtered candidate type list; Then, the candidate collection unit is specifically used to: A constraint candidate type set and a statistical candidate type set are collected in the filtering candidate type list based on classes and interfaces.
10. The device for inferring code snippet type by integrating constraints and statistics according to claim 6, characterized in that: The type determination unit is specifically configured to: If the optimization constraint inference type is a null value, determining the optimization statistical inference type as the target inference type; If the optimization statistical inference type is a null value, determining the optimization constraint inference type as the target inference type; If both the optimization constraint inference type and the optimization statistical inference type are non-null values, the optimization constraint inference type is selected as the target inference type.
Citation Information
Patent Citations
C program code standard checking device based on PRDL rule description language
CN106970819A
Providing question and answers with deferred type evaluation using text with limited structure
WO2012040356A1