Method and apparatus for type inference of programming language
By obtaining the nodes to be inferred in the abstract syntax tree and determining their types based on related constraints, the high consumption problem of type inference in the prior art is solved, real-time and incremental type inference are realized, and efficiency and user experience are improved.
Patent Information
- Application Number
- PCT/CN2024/127570
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-13
- Filing Date
- 2024-10-28
- Publication Date
- 2025-06-19
AI Technical Summary
The prior art requires inference of the entire project when performing type inference of programming languages, resulting in severe memory consumption, especially when dealing with large projects or large batches of files.
During the process of program code development, the nodes to be inferred in the abstract syntax tree are obtained and their types are determined based on constraints related to the node to be inferred, and type inference of the entire project is avoided, thereby reducing memory consumption.
Real-time refresh and incremental inference of types are implemented, which reduces memory consumption and improves the efficiency and user experience of type inference.
Smart Images

Figure CN2024127570_19062025_PF_FP_ABST
Abstract
Description
Method and device for type inference of programming language
[0001] This application claims priority to Russian Federal Patent Application No. 2023132940 filed with the State Intellectual Property Office of the Russian Federation on December 13, 2023, and priority to Russian Federal Patent Application entitled “Method and Device for Type Inference in Programming Languages”, the entire contents of which are incorporated herein by reference. Technical Field
[0002] The present application relates to the field of information technology, and more specifically, to a method and device for type inference in a programming language. Background Art
[0003] Type inference in programming languages is a crucial tool for improving program reliability and plays a crucial role in program understanding and error prevention. For example, in integrated development environments (IDEs), type inference of programming language expressions underlies many IDE features, such as source code constraint resolution, code completion, and program semantic verification. Therefore, type inference is a key task of the IDE's code model component.
[0004] These solutions typically require type inference across the entire project. If the source code changes, types must be re-inferred for the entire project. This memory usage is particularly significant when working with large projects or large batches of files.
[0005] Summary of the Invention
[0006] The present application provides a method and apparatus for type inference in a programming language, which is conducive to reducing the memory occupied by type inference.
[0007] In a first aspect, a method for type inference in a programming language is provided, which can be applied to an IDE. The method includes: during program code development, obtaining a node to be inferred in an abstract syntax tree, where the node to be inferred indicates a field to be inferred in the program code, and the field to be inferred is one of the following: a variable, a class, a function, or a function; obtaining constraints related to the node to be inferred, where the constraints are used to indicate the relationship involved in the node to be inferred in the abstract syntax tree; determining the type of the node to be inferred according to the constraints, where the type of the node to be inferred includes one of the following: a numeric type, a sequence type, a mapping type, a collection type, an iterator type, a context manager type, a module type, a metatype, a null value type, or a callable type.
[0008] In the solution of the embodiment of the present application, the type of the node to be inferred is inferred based on the constraints related to the node to be inferred, that is, the type of the node to be inferred is determined based on the relationship involved in the node to be inferred in the AST. There is no need to infer the type of the entire project, which is beneficial to reducing memory consumption and improving the efficiency of type inference, thereby realizing real-time refresh of types and improving user experience.
[0009] At the same time, in the embodiments of the present application, the inference result of the type of the node to be inferred is obtained based on the constraints related to the node to be inferred, which is conducive to the implementation of incremental type inference, that is, the implementation of incremental type modification. For example, if the source code is changed, the type can be inferred only for the changed source code file and the corresponding dependencies, without having to re-infer the type of the entire project.
[0010] Taking Python as an example, the node to be inferred may include any of the following: a variable node, a class node, a function node, or a parameter node.
[0011] Exemplarily, a node constraint may include at least one of the following: a node's base class type constraint, a node's resolved reference constraint, a node's attribute reading constraint, a node's default value constraint, a node's call constraint, a node's assignment constraint, a node's binary constraint, a node's parameter passing constraint, or a node's return constraint.
[0012] In combination with the first aspect, in certain implementations of the first aspect, the type of the node to be inferred is used for at least one of the following: displaying the type of the node to be inferred, code completion, code refactoring, code inspection, or quick fix.
[0013] In combination with the first aspect, in certain implementations of the first aspect, obtaining a node to be inferred in an abstract syntax tree includes: taking out a first inference item from an inference queue, the first inference item being used to determine the type of the node to be inferred, the inference queue including one or more inference items; and obtaining the node to be inferred.
[0014] In an embodiment of the present application, an inference queue can be used to process type inference tasks, allowing type inference in a delayed mode, which is beneficial to improving inference performance.
[0015] In combination with the first aspect, in certain implementations of the first aspect, the constraint includes a parsed reference constraint of the first node in the abstract syntax tree, and the parsed reference constraint of the first node is used to indicate the reference relationship of the node to be inferred to the first node. The method also includes: obtaining the type of the first node, and determining the type of the node to be inferred based on the constraint, including: determining the type of the node to be inferred based on the type of the first node and the constraint.
[0016] In the solution of the embodiment of the present application, if the node to be inferred involves references to other nodes, the type of the referenced node is first obtained, and then the type of the node to be inferred is inferred based on the type of the referenced node, without inferring the type of the entire project.
[0017] In combination with the first aspect, in certain implementations of the first aspect, the constraint includes an assignment constraint of the second node, the assignment constraint of the second node is used to indicate the assignment relationship involving the second node, and the type of the assigned child node in the second node is determined based on the type to the right of the assignment number.
[0018] In combination with the first aspect, in certain implementations of the first aspect, the constraint includes an attribute read constraint of the third node, the attribute read constraint of the third node is used to indicate the subordinate relationship involved in the third node, and the type of the owner of the attribute in the third node includes the type of the attribute.
[0019] In combination with the first aspect, in certain implementations of the first aspect, the method also includes: creating a type variable corresponding to the node to be inferred; and determining the type of the node to be inferred based on the constraints, including: determining the value of the type variable corresponding to the node to be inferred based on the constraints, the value of the type variable corresponding to the node to be inferred being used to indicate the type of the node to be inferred.
[0020] The newly created type variable can be called a unique free type variable. A unique free type variable is a variable that is not currently associated with any other specific type.
[0021] Taking Python as an example, the fields to be inferred can include any of the following: class, static variable, instance class variable, function, parameter, or return value. New type variables can be created for each class, static variable, instance class variable, function, parameter, or return value.
[0022] In combination with the first aspect, in some implementations of the first aspect, the method further includes: storing the type of the node to be inferred in a type environment, where the type environment is used to store the type of the node in the abstract syntax tree.
[0023] In combination with the first aspect, in some implementations of the first aspect, the method further includes: outputting an inference result, where the inference result indicates the type of the node to be inferred.
[0024] In a second aspect, a device for type inference of a programming language is provided, which is applied to an integrated development environment (IDE), and the device includes: a first acquisition module, used to obtain a node to be inferred in an abstract syntax tree during program development, where the node to be inferred indicates a field to be inferred in the program code, and the field to be inferred includes one of the following: a variable, a class, a function, or a parameter; a second acquisition module, used to obtain constraints related to the node to be inferred, where the constraints are used to indicate the relationship involved in the node to be inferred in the abstract syntax tree; a processing module, used to determine the type of the node to be inferred based on the constraints, where the type of the node to be inferred includes at least one of the following: a numeric type, a sequence type, a mapping type, a collection type, an iterator type, a context manager type, a module type, a metatype, a null value type, or a callable type.
[0025] In combination with the second aspect, in certain implementations of the second aspect, the type of the node to be inferred is used for at least one of the following: displaying the type of the node to be inferred, code completion, code refactoring, code inspection, or quick fix.
[0026] In combination with the second aspect, in certain implementations of the second aspect, the first acquisition module is specifically used to: take out a first inference item from the inference queue, the first inference item is used to determine the type of the node to be inferred, and the inference queue includes one or more inference items; and obtain the node to be inferred.
[0027] In combination with the second aspect, in certain implementations of the second aspect, the constraint includes a parsed reference constraint of the first node in the abstract syntax tree, and the parsed reference constraint of the first node is used to indicate the reference relationship of the node to be inferred to the first node. The method also includes a third acquisition module for obtaining the type of the first node, and the processing module is specifically used to determine the type of the node to be inferred based on the type of the first node and the constraint.
[0028] In combination with the second aspect, in certain implementations of the second aspect, the constraint includes an assignment constraint of the second node, and the assignment constraint of the second node is used to indicate the assignment relationship involving the second node, and the type of the assigned child node in the second node is determined based on the type to the right of the assignment number.
[0029] In combination with the second aspect, in certain implementations of the second aspect, the constraint includes an attribute read constraint of the third node, the attribute read constraint of the third node is used to indicate the subordinate relationship involved in the third node, and the type of the owner of the attribute in the third node includes the type of the attribute.
[0030] In combination with the second aspect, in certain implementations of the second aspect, the device also includes a creation module for creating a type variable corresponding to the node to be inferred; and the processing module is specifically used to: determine the value of the type variable corresponding to the node to be inferred based on the constraint, and the value of the type variable corresponding to the node to be inferred is used to indicate the type of the node to be inferred.
[0031] In combination with the second aspect, in some implementations of the second aspect, the apparatus further includes a storage module for storing the type of the node to be inferred in a type environment, where the type environment is used to store the type of the node in the abstract syntax tree.
[0032] It should be understood that the expansion, limitation, explanation and description of the relevant content in the above-mentioned first aspect also apply to the same content in the second aspect.
[0033] In a third aspect, a computing device cluster is provided, comprising at least one computing device, each computing device including a processor and a memory. The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster performs the method of the first aspect and any implementation of the first aspect.
[0034] In a fourth aspect, a computer-readable medium is provided, comprising computer program instructions. When the computer program instructions are executed by a computing device cluster, the computing device cluster executes the method in the first aspect and any one of the implementations of the first aspect.
[0035] In a fifth aspect, a computer program product comprising instructions is provided. When the instructions are executed by a computing device cluster, the computing device cluster executes the method in the above-mentioned first aspect and any one of the implementations of the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] FIG1 is a schematic flow chart of a method for type inference in a programming language according to an embodiment of the present application;
[0037] FIG2 is a schematic diagram of an assignment constraint according to an embodiment of the present application;
[0038] FIG3 is a schematic flow chart of a method for type inference in another programming language according to an embodiment of the present application;
[0039] FIG4 is a schematic block diagram of an apparatus for type inference in a programming language according to an embodiment of the present application;
[0040] FIG5 is a schematic block diagram of a computing device according to an embodiment of the present application;
[0041] FIG6 is a schematic block diagram of a computing device cluster according to an embodiment of the present application;
[0042] FIG7 is a schematic block diagram of another computing device cluster according to an embodiment of the present application. DETAILED DESCRIPTION
[0043] The technical solution in this application will be described below with reference to the accompanying drawings.
[0044] The terms used in the following embodiments are for the purpose of describing specific embodiments only and are not intended to limit the present application. As used in the specification and appended claims of this application, the singular expressions "a," "an," and "the" are intended to include expressions such as "one or more," unless the context clearly indicates otherwise. It should also be understood that in the following embodiments of this application, "at least one," "at least one," and "one or more" refer to one, two, or more. "First," "second," and various numerical designations are merely distinctions made for ease of description and are not intended to limit the scope of the embodiments of this application. "And / or" is used to describe the corresponding relationship between corresponding objects, indicating that three relationships can exist. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist, where A and B can be singular or plural. The character " / " generally indicates that the objects associated with each other are in an "or" relationship. The order of the sequence numbers of the processes below does not imply a sequence of execution. The execution order of each process should be determined by its function and inherent logic and should not constitute any limitation on the implementation process of the embodiments of this application. For example, in the embodiments of the present application, words such as "301", "401", and "501" are merely identifiers for the convenience of description and do not limit the order of executing the steps.
[0045] References to "one embodiment" or "some embodiments" in this specification mean that one or more embodiments of the present application include specific features, structures or characteristics described in conjunction with the embodiment. In this application, words such as "exemplary" or "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design described in this application as "exemplary" or "for example" should not be interpreted as being more preferred or more advantageous than other embodiments or design. Specifically, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete way. The terms "including", "comprising", "having" and their variations all mean "including but not limited to", unless otherwise specifically emphasized. In the embodiments of the present application, descriptions such as "when...", "in the case of...", "if" and "if" all mean that the device will perform corresponding processing under certain objective circumstances, and do not limit the time, nor do they require the device to perform judgment actions when implemented, nor do they mean that there are other limitations.
[0046] In this application, "used to indicate" can include being used for direct indication and being used for indirect indication. When describing that a certain indication information is used to indicate A, it can include that the indication information directly indicates A or indirectly indicates A, and it does not mean that the indication information must carry A.
[0047] In order to help those skilled in the art better understand the technical solutions of the present application, some terms that may be involved in the embodiments of the present application are explained below.
[0048] 1. Integrated development environment (IDE):
[0049] IDE is an application used to provide a program development environment, generally including tools such as code editors, compilers, debuggers, and graphical user interfaces. It is an integrated development software service suite that integrates code writing, analysis, compilation, debugging, and other functions.
[0050] IDEs may include local IDEs and web IDEs.
[0051] A local IDE is also called a desktop IDE. A local IDE is installed in a user's local development operating environment. For example, the local development operating environment can be a terminal device, such as a desktop computer, laptop, or mobile phone.
[0052] Web IDE refers to an online IDE service, consisting of an IDE front-end and an IDE back-end. The IDE back-end runs in a remote environment. For example, the remote environment can be a cloud server, providing the IDE to users as a cloud service. Users can purchase the cloud service by visiting a website, and the cloud service provider will create an IDE instance. Users can then develop programs based on this IDE instance.
[0053] 2. Abstract syntax tree (AST):
[0054] An AST, also known as a syntax tree, is an abstract representation of the grammatical structure of source code. It represents the grammatical structure of the source code in a tree-like format, with each node representing a structure in the source code, such as a variable declaration, expression, function call, or control structure. The root node of the tree typically represents the entire source code file, while the child nodes represent specific grammatical elements and their relationships.
[0055] The grammar is "abstract" because it doesn't represent every detail that would be present in a real grammar. For example, nested parentheses are implicit in the tree structure and not represented as nodes; conditional jump statements like if-condition-then can be represented using a node with three branches.
[0056] AST has a wide range of functions, such as code syntax checking, code highlighting, code error checking, code completion, code optimization, code repair, code generation, and code refactoring.
[0057] Type inference in static languages (such as Java or C#) is primarily performed by the compiler. Code models can obtain the required type information through type binding, which is typically stored in the program's metadata.
[0058] Type inference in dynamic languages is primarily performed by IDEs. Due to a lack of support for compilation metadata, type inference in dynamic languages remains an industry challenge. For example, Python has a fully dynamic and strict type system. Types are dynamically bound to variables at runtime, and both variables and types can be dynamically changed at runtime, making reliable type inference difficult.
[0059] Some related solutions can implement type inference for dynamic languages. For example, the type inference method of static languages is used to implement type inference for dynamic languages, that is, by parsing the metadata (.pyi file) of the generated program, and then obtaining the required type information based on the metadata. In the current solution, type inference needs to be performed based on the entire project. If the source code changes, the type needs to be re-inferred for the entire project. Type inference requires memory, especially when processing large projects or large batches of files, type inference will take up a lot of memory, resulting in relatively serious memory consumption.
[0060] In view of this, an embodiment of the present application provides a method for type inference of a programming language, which is conducive to reducing memory consumption.
[0061] The method of the embodiment of the present application can be applied to scenarios where type inference of a programming language is required.
[0062] For example, the solution of the present application example can be applied to a scenario where the type of a field needs to be displayed. For example, when the mouse hovers over a variable, the type of the variable can be displayed. The type of the variable can be determined by the method of the embodiment of the present application.
[0063] For example, the solution of the embodiment of the present application can be applied to scenarios that require subsequent processing based on the type of programming language. For example, the solution of the embodiment of the present application can be applied to scenarios such as code completion, code refactoring, code inspection or quick repair.
[0064] FIG1 shows a method for type inference of a programming language according to an embodiment of the present application. The method 100 shown in FIG1 can be executed by a device for type inference of a programming language.
[0065] Exemplarily, the device for inferring the type of a programming language may be a development tool or a functional module of a development tool. For example, the development tool may be an IDE. That is, method 100 may be applied to an IDE.
[0066] Exemplarily, the device for inferring the type of a programming language may also be a device used in conjunction with a development tool.
[0067] IDEs can include desktop IDEs or web IDEs. The backend of a web IDE runs in a remote environment. For example, the remote environment can be a cloud server.
[0068] The above is only an example. The solution of the embodiment of the present application can also be executed by other devices, and the embodiment of the present application is not limited to this.
[0069] As shown in FIG1 , method 100 includes the following steps.
[0070] 110, obtaining the node to be inferred in the AST.
[0071] 120. Obtain constraints related to the node to be inferred. The constraints are used to indicate the relationships involved in the node to be inferred in the AST.
[0072] 130. Determine the type of the node to be inferred according to the constraint.
[0073] In the solution of the embodiment of the present application, the type of the node to be inferred is inferred based on the constraints related to the node to be inferred, that is, the type of the node to be inferred is determined based on the relationship involved in the node to be inferred in the AST. There is no need to infer the type of the entire project, which is beneficial to reducing memory consumption and improving the efficiency of type inference, thereby realizing real-time refresh of types and improving user experience.
[0074] At the same time, in the embodiments of the present application, the inference result of the type of the node to be inferred is obtained based on the constraints related to the node to be inferred, which is conducive to the implementation of incremental type inference, that is, the implementation of incremental type modification. For example, if the source code is changed, the type can be inferred only for the changed source code file and the corresponding dependencies, without having to re-infer the type of the entire project.
[0075] In step 110, the node to be inferred may indicate a field to be inferred in the code segment. The type of the node to be inferred is the type of the field to be inferred.
[0076] Exemplarily, taking Python as an example, the node to be inferred may include any of the following: a variable node, a class node, a function node, or a parameter node, etc. The field to be inferred may include any of the following: a variable, a class, a function, or a parameter.
[0077] Accordingly, the type of the node to be inferred may include any of the following: the type of a variable, the type of a parameter, the type of a class, and the type of a function.
[0078] Exemplarily, taking Python as an example, the type of the node to be inferred includes at least one of the following: number type, sequence type, mapping type, set type, iterator type, context manager type, module type, metaclass type, none type or callable type.
[0079] For example, the number type may include any of the following: integer type, floating-point type, complex type, Boolean (bool) type, long integer type, etc.
[0080] For example, the sequence type may include any of the following: a string type, a list type, a tuple type, or a bytes type.
[0081] For example, a mapping type may include a dictionary type.
[0082] For example, collection types may include: Set and Immutable Collection.
[0083] For example, callable types can include: function, method, or class.
[0084] For example, iterator types can include any of the following: a generator, a generator function, or a built-in iterator.
[0085] For example, a context manager type might include objects that implement enter() and exit() methods.
[0086] For example, the module type may include any of the following: module or package.
[0087] Exemplarily, the type of the variable or the type of the parameter may include any of the following: a numeric type, a sequence type, a mapping type, a collection type, or a null value type.
[0088] Exemplarily, the type of a function may include a callable type. Further, the type of a function may also include the types of the function's parameters and the type of its return value.
[0089] Furthermore, the type of the node to be inferred may also include: any type or unknown type.
[0090] It should be understood that the above is only an example and does not constitute a limitation on the types of fields in the embodiments of the present application.
[0091] In the embodiments of the present application, the "function" in the code can also be called a "method", and the embodiments of the present application do not make this distinction.
[0092] The number of nodes to be inferred may be one or more, and the embodiment of the present application does not limit the number of nodes to be inferred.
[0093] The number of constraints associated with the node to be inferred may be one or more, and the embodiment of the present application does not limit the number of constraints.
[0094] It should be noted that if there are multiple nodes to be inferred, the nodes to be inferred can be obtained once or multiple times, and this embodiment of the present application does not limit this.
[0095] The method 100 may be applied in a process of developing program codes.
[0096] Method 100 may be triggered in a variety of ways.
[0097] In a possible implementation, method 100 may be triggered by an opening operation of a source code file.
[0098] For example, the source code file has not previously been subjected to a type inference task, and when a user opens the source code file, the method 100 is triggered for execution.
[0099] The AST may be an AST corresponding to a source code file, and the nodes to be inferred may be some or all of the nodes in the AST.
[0100] If the source code file is large, performing type inference on the entire source code file at once may consume a large amount of memory and affect processing efficiency. In the solution of the embodiment of the present application, type inference can be performed on part of the code in the source code file, or in other words, the code in the source code file can be processed in blocks, without having to process the entire source code file at once, which helps reduce memory consumption and ensures processing efficiency.
[0101] Exemplarily, the node to be inferred may be from a node corresponding to a field in the code segment displayed in the target area.
[0102] This allows type inference to be performed on code sections of a source code file that are displayed in the target area.
[0103] If the source code file is large, the code in the source code file may not be displayed in one interface at the same time. In this case, type inference can be performed on the code segments displayed in the same area (for example, the target area).
[0104] For example, method 100 can be executed by an IDE, and the target area can be a code editing box in the IDE. After opening a source code file, method 100 can be executed on the code segment currently displayed in the code editing box. For example, the node corresponding to the field in the code segment currently displayed in the code editing box is used as the node to be inferred, and method 100 is executed based on the node to be inferred to obtain an inference result of the type of the code segment currently displayed in the code editing box. This inference result can be displayed to the user.
[0105] In one possible implementation, method 100 may be triggered by a mouse event.
[0106] Exemplarily, the method 100 may be triggered by any one or more of the following: a hover operation of a mouse, a single click operation of a mouse, or a double click operation of a mouse.
[0107] For example, if method 100 is triggered by a mouse hover operation, for example, a user who wishes to obtain the type of a variable may hover the mouse over the variable. In response to the mouse hover operation, method 100 is executed. For example, the node corresponding to the variable is used as the node to be inferred, and method 100 is executed based on the node to be inferred to obtain an inferred result of the variable's type, which is then displayed to the user.
[0108] In one possible implementation, method 100 may be triggered by a keyboard event.
[0109] Exemplarily, the method 100 may be triggered by a pressing operation of a preset key.
[0110] For example, if a user needs to obtain the type of a code segment on the current interface, they press a preset key. In response to this operation, method 100 is executed. For example, a node of a field in the code segment on the current interface is used as a node to be inferred. Method 100 is executed based on the node to be inferred to obtain an inference result of the type of the code segment on the current interface, and the inference result is displayed to the user.
[0111] For another example, if a user needs to obtain the type of a particular code segment, they can select the segment and press a preset key. In response to this operation, method 100 is executed. For example, the node of the field in the selected segment is used as the node to be inferred. Method 100 is executed based on the node to be inferred, and an inference result of the type of the selected segment is obtained, which is then displayed to the user.
[0112] It should be understood that the above triggering methods are only examples and do not limit the solutions of the embodiments of the present application.
[0113] Optionally, the type of the node to be inferred is indicated by the value of a type variable corresponding to the node to be inferred.
[0114] Inferring the type of the node to be inferred, that is, inferring the value of the type variable corresponding to the node to be inferred.
[0115] Type variables are used to describe types. A type variable can be understood as a variable that represents a type. In other words, a type can be used as a variable; that variable is a type variable. The value of a type variable corresponding to a node indicates the type of that node. For example, the value of a type variable corresponding to a function node indicates the type of the function. Similarly, the value of a type variable corresponding to a parameter node indicates the type of the parameter.
[0116] Optionally, the method 100 may further include: creating corresponding type variables for the nodes to be inferred respectively.
[0117] This step can also be called initialization of type variables. Type variables can be initialized based on the class declaration, variable declaration, method declaration, parameter declaration, etc. in the code segment.
[0118] Specifically, a new type variable is created for each node to be inferred. For example, the node to be inferred may be a class declaration, a variable declaration, a method declaration, a parameter declaration, or other node in the code segment.
[0119] The newly created type variables are called unique free type variables. Unique free type variables are variables that are not currently associated with any specific type. The type associated with these variables, i.e., the value of the unique free variable, must be inferred. This means that the type variables of the node to be inferred are not associated with any specific type.
[0120] Taking Python as an example, the fields to be inferred can include any of the following: class, static variable, instance class variable, function, parameter, or return value. New type variables can be created for each class, static variable, instance class variable, function, parameter, or return value.
[0121] It should be understood that the above is only an example, and the node type can also be expressed in other forms, which is not limited in the embodiments of the present application.
[0122] Constraints can be used to describe relationships within a code snippet. For example, the relationships within a Python program snippet might include assignments, function calls, class instance creation, or binary operations.
[0123] Taking Python as an example, illustratively, the constraint may include at least one of the following: a base class type constraint, a resolved reference constraint, an attribute read constraint, an assignment constraint, a binary constraint, a default value constraint, a parameter passing constraint, a call constraint, or a return constraint.
[0124] For example, the constraint includes a base class type constraint of node #1, where the base class type constraint of node #1 is used to indicate that node #1 is a base class.
[0125] For another example, the constraint includes a resolved reference constraint of node #2 (an example of the first node), and the resolved reference constraint of node #2 is used to indicate a reference relationship to node #2.
[0126] For another example, the constraint includes an attribute read constraint of node #3 (an example of the third node), and the attribute read constraint of node #3 is used to indicate a subordinate relationship in node #3.
[0127] For another example, the constraint includes a default value constraint of node #4, where the default value constraint of node #4 is used to indicate that node #4 is assigned a default value.
[0128] For another example, the constraint includes a call constraint of node #5, and the call constraint of node #5 is used to indicate that node #5 is a callable function.
[0129] For another example, the constraint includes an assignment constraint of node #6 (an example of the second node), and the assignment constraint of node #6 is used to indicate the assignment relationship corresponding to node #6.
[0130] For another example, the constraint includes a binary constraint of node #7, and the binary constraint of node #7 is used to indicate a binary operation relationship corresponding to node #7.
[0131] For another example, the constraint includes a parameter transfer constraint of node #8, and the parameter transfer constraint of node #8 is used to indicate a parameter transfer relationship corresponding to node #8.
[0132] For another example, the constraint includes a return constraint of node #9, and the return constraint of node #9 is used to indicate a return relationship corresponding to node #9.
[0133] This constraint can be obtained in a number of ways.
[0134] Exemplarily, the constraint may be created in advance. In this case, step 120 may include: reading the constraint.
[0135] Exemplarily, the constraint may be created during the execution of method 100. In this case, step 120 may include: creating the constraint. The constraint may be created based on the relationship between syntax elements indicated by the AST.
[0136] Exemplarily, creating the constraint may include at least one of the following:
[0137] Create base class type constraints for nodes belonging to the base class;
[0138] Create resolved referential constraints for nodes involved in the reference;
[0139] Create attribute read constraints for nodes involved in dependency relationships;
[0140] Create default value constraints for nodes that are set to default values;
[0141] Create call constraints for nodes involved in function calls;
[0142] Create assignment constraints for nodes involved in assignment relationships;
[0143] Create binary constraints for nodes involved in binary operations;
[0144] Create parameter transfer constraints for nodes involved in parameter transfer relationships; or
[0145] Create a return constraint for the node that serves as the return value.
[0146] For example, in the case where node #1 is a base class, a base class type constraint is created for node #1.
[0147] For example, if there is a reference to node #2 in the AST, a resolution reference constraint is created for node #2.
[0148] Create attribute read constraints for nodes involved in dependency relationships, that is, create attribute read constraints for nodes involved in attribute reads.
[0149] For example, if node #3 in the AST involves a dependency relationship, a property read constraint is created for node #3.
[0150] For example, node #4 in the AST is assigned a default value, creating a default value constraint for node #4.
[0151] For example, if node #5 in the AST is a callable function, create a call constraint for node #5.
[0152] For example, node #6 in the AST corresponds to an assignment statement, and an assignment constraint is created for node #6.
[0153] For example, node #7 in the AST corresponds to an expression of a binary operation, and a binary constraint is created for node #7.
[0154] For example, if node #8 in the AST involves parameter passing of a function, a parameter passing constraint is created for node #8.
[0155] For example, node #9 in the AST corresponds to a return statement, and a return constraint is created for node #9.
[0156] The constraints related to the node to be inferred may include constraints directly related to the node to be inferred. The constraints directly related to the node to be inferred may be understood as constraints used to indicate the relationship between the node to be inferred and other syntax elements.
[0157] Furthermore, the constraints related to the node to be inferred may also include constraints indirectly related to the node to be inferred.
[0158] The constraints indirectly related to the node to be inferred may include a plurality of nested constraints obtained by decomposing the constraints directly related to the node to be inferred.
[0159] Figure 2 shows a schematic diagram of an assignment constraint in an embodiment of the present application. The constraints related to the variable "a" are exemplarily explained below in conjunction with Figure 2. In order to describe this solution more intuitively, the following mainly uses code snippets as an example for explanation. In the actual execution process, the relationship between the various syntax elements can be obtained through AST, and various constraints can be created based on nodes. For example, creating an assignment constraint for "a=b+c", that is, creating an assignment constraint for the assignment node corresponding to "a=b+c". For another example, creating a parsing reference constraint for the variable "b", that is, creating a parsing reference constraint for the node of the variable "b" (that is, the variable node "b"). For the sake of ease of description, the embodiment of the present application does not make a distinction between these.
[0160] For example, the code snippet "a=b+c" shown in Figure 2 is an assignment statement, involving an assignment relationship, and an assignment constraint is created for "a=b+c". This assignment constraint can be considered a constraint directly related to the variable "a".
[0161] As shown in Figure 2, this assignment constraint can be decomposed into the following nested constraints: "b" and "c" are referenced variables, so resolved reference constraints are created for "b" and "c"; "b + c" involves a binary operation, so a binary constraint is created for "b + c." The resolved reference constraint and the binary constraint can be considered as constraints indirectly related to variable "a." Furthermore, "a" can also be considered a referenced variable, so a resolved reference constraint is created for "a."
[0162] It should be understood that the above description only uses Python as an example, and does not limit the solution of the embodiment of the present application. The solution of the embodiment of the present application can also be applied to type inference of other dynamic languages.
[0163] The constraint may also be obtained through other means, which is not limited in the embodiments of the present application.
[0164] Optionally, if the constraint includes a resolved reference constraint, obtain the type of the dependency indicated by the resolved reference constraint.
[0165] In other words, if the node to be inferred involves references to other nodes, and the referenced nodes are dependencies, the types of the dependencies are first obtained, and then the type of the node to be inferred can be determined based on the types of the dependencies.
[0166] For example, the resolved reference constraint of node #2 indicates a reference to node #2. Node #2 is a dependency. Method 100 may include obtaining a node type of node #2. Step 130 may include determining the type of the node to be inferred based on the type of node #2 and the constraint.
[0167] If the resolved reference constraint indicates multiple dependencies, the types of the multiple dependencies may be obtained in the same or different ways.
[0168] Further, optionally, method 100 may include: determining whether the type of the dependency has been inferred.
[0169] Optionally, you can determine whether the type of the dependency has been inferred based on the type environment. The type environment can be used to store the type of the node in the AST.
[0170] That is, the type environment can be used to store the types of fields in the source code file.
[0171] Exemplarily, the type of a node can be indicated by the value of a type variable corresponding to the node. In this case, a type environment can be used to store the value of the type variable.
[0172] For example, to determine whether the type of a node has been inferred, one can determine whether a value of a type variable corresponding to the node exists in the type environment. If so, the type of the node has been inferred. If not, the type of the node has not yet been inferred.
[0173] If the type of the dependency has been previously inferred, obtaining the type of the dependency may include reading the type of the dependency.
[0174] If the type of the dependency has not been inferred, obtaining the type of the dependency may include inferring the type of the dependency. For example, if the type of a dependency has not been inferred, the current inference process is stopped and the type of the dependency is inferred. The type of the dependency may be determined using method 100, i.e., the dependency is used as a node to be inferred to obtain the type of the dependency.
[0175] After inferring the types of all dependencies indicated by resolved reference constraints, the type of the node to be inferred is determined based on the types of all dependencies and any other constraints in the constraint, excluding the resolved reference constraints. If the type of the node to be inferred is indicated by the value of a type variable, determining the type of the node to be inferred is simply determining the value of the type variable of the node to be inferred.
[0176] In the solution of the embodiment of the present application, the type of the field to be inferred can be inferred based on the type of the dependency, without inferring the type of the entire project.
[0177] The following describes the processing of constraints other than resolve referential constraints.
[0178] Exemplarily, if the constraint includes a base class type constraint of node #1, the type of node #1 includes the base class.
[0179] For example, if the constraint includes a read constraint for the attribute of node #3, the type of the attribute in node #3 comes from the type of the owner of the attribute. In other words, the type of the owner of the attribute in node #3 includes the type of the attribute.
[0180] For example, the type of a property can be determined based on the type of the owner. If the type of the owner is known, the type of the property can be requested from the type of the owner, that is, the type of the property can be requested by the type of the qualifier.
[0181] For another example, the type of the owner can be determined based on the type of the attribute. If the type of the owner is unknown, the type of the owner of the attribute can be determined based on the type of the attribute, that is, the type of the owner includes the type of the attribute.
[0182] For example, if the constraint includes a default value constraint for node #4, the type of node #4 can be determined based on the type of the default value.
[0183] Exemplarily, if the constraint includes a call constraint for node #5, the type of node #5 is a callable function.
[0184] Exemplarily, if the constraint includes an assignment constraint for node #6, the type of the assigned child node in node #6 is determined according to the type on the right side of the assignment sign "=".
[0185] Furthermore, the type of the assigned child node is obtained by unification of the current type of the assigned child node and the type to the right of the assignment sign "=". The current type of the assigned child node can be called the type to the left of the assignment sign "=".
[0186] If the current type of the child node is a specific type, for example, the type of the child node has been previously defined as type A, then after combining the current type on the left side of the "=" (i.e., type A) and the type on the right side of the "=" (e.g., type B)) (A|B), the resulting type of the child node can include the type on the left side of the "=" or the type on the right side of the "="", that is, the type of the child node can include type A or type B. If the current type of the child node is unknown, the type of the child node can include the type on the right side of the "=".
[0187] Exemplarily, if the constraint includes a parameter transfer constraint of node #8, the type of the parameter child node in node #8 is determined according to the parameter transfer relationship in node #8.
[0188] Exemplarily, if the constraint includes a return constraint of node #9, the type of the return value involved in node #9 is the type of the return result of node #9.
[0189] The following is an exemplary description of the inference process of the type of the variable "a" with reference to FIG. 2.
[0190] As shown in FIG. 2 , the process of inferring the type of variable “a” may include the following steps.
[0191] (1) Infer the types of "b" and "c".
[0192] As mentioned previously, the constraints in Figure 2 include resolving reference constraints. The dependencies indicated by the resolved reference constraints include "b" and "c." First, the types of "b" and "c" are inferred.
[0193] (2) Infer the type of the add operation.
[0194] The constraints in Figure 2 also include a binary constraint. The binary operation indicates a binary addition operation between b and c. The type of the addition operation is inferred based on the binary operation.
[0195] (3) Get the current type of variable "a" or the type variable of "a".
[0196] The constraint in FIG2 includes a resolved reference constraint, and the dependency indicated by the resolved reference constraint may include "a", and the type of "a" is obtained. If the type of the variable "a" has been defined before, the current type of the variable "a" is obtained.
[0197] Gets the type variable of variable "a" if variable "a" has not been previously inferred.
[0198] (4) Determine the type of the variable "a" based on the type obtained in step (2) and the type obtained in step (3).
[0199] The constraints in Figure 2 also include assignment constraints. The type obtained in step (2) is the type on the right side of the "=". The type obtained in step (3) is the type on the left side of the "=".
[0200] The inferred type of the variable "a" is the union of the type obtained in step (2) and the type obtained in step (3).
[0201] For example, the type of the addition operation obtained in step (2) is t1, and the current type of the variable "a" obtained in step (3) is t2. The union {t2|t1} is performed on t1 and t2. The inferred type of the variable "a" includes t1 or t2, that is, the type of the variable "a" includes t1 or t2.
[0202] For example, the type of the addition operation obtained in step (2) is t1, and the type variable of "a" obtained in step (3) is t3. Union t1 and the type variable t3 to obtain t3 = {t1}, that is, the value of the type variable t3 includes t1, that is, the type of the variable "a" includes t1.
[0203] Furthermore, the method 100 may further include: determining whether the type of the node to be inferred has been inferred.
[0204] If the type of the node to be inferred has not been inferred, step 130 may be executed.
[0205] Determining whether the type of the node to be inferred is inferred may include determining whether the types of some nodes in the node to be inferred are inferred. Alternatively, determining whether the type of the node to be inferred is inferred may include determining whether the types of all nodes in the node to be inferred are inferred.
[0206] Exemplarily, the node to be inferred includes the root node of the AST. If the type of the root node has not been inferred, step 130 may be executed. If the type of the root node has been inferred, step 130 is not executed.
[0207] Whether the root node's type has been inferred can be determined based on the type environment. For example, the type environment is checked. If the root node's type exists within the type environment, the root node's type has been inferred. If the root node's type does not exist within the type environment, the root node's type has not been inferred. If the root node's type has already been inferred, the inference task has already been performed on the source code file, and the inference result can be directly output. The inference result can be determined based on the type environment.
[0208] This can avoid repeated inference of types, which helps to further reduce memory consumption and improve inference efficiency.
[0209] Furthermore, the method 100 may further include: creating a first inference item, and adding the first inference item to the inference queue. The first inference item is used to determine the type of the node to be inferred.
[0210] In other words, the first inference item indicates an inference task of the type of the node to be inferred.
[0211] Optionally, step 110 may include: extracting a first inference item from an inference queue to obtain the node to be inferred.
[0212] An inference queue is a queue used to store inference items. An inference queue contains one or more inference items. An inference item indicates the type of inference task.
[0213] The inference queue may adopt a first in first out (FIFO) data structure.
[0214] In the embodiment of the present application, an inference queue can be used to process type inference tasks, allowing type inference in a delayed mode, which is beneficial to improving inference performance. The delayed mode can also be called lazy mode.
[0215] For example, in step 130, if the type of the dependency item is not inferred, the inference process of the first inference item is stopped, a second inference item is created, and the second inference item is added to the inference queue. The second inference item indicates an inference task for the type of the dependency item.
[0216] The "first inference item" and "second inference item" are only used to distinguish inference tasks of different node types and have no other limiting effect.
[0217] If the types of multiple dependencies among the dependencies indicated by the resolved reference constraint are not inferred, inference items can be created for each of the multiple dependencies and added to the inference queue. The inference items corresponding to the multiple dependencies can all be considered as second inference items.
[0218] After all the inference items corresponding to the dependent items are taken out from the inference queue and the inference is completed, the first inference item will be processed.
[0219] Furthermore, the method 100 may further include: storing the type of the node to be inferred in a type environment.
[0220] Furthermore, the method 100 may further include: outputting an inference result, where the inference result indicates the type of the node to be inferred.
[0221] Exemplarily, outputting the inference result may include displaying the type of the node to be inferred.
[0222] Optionally, the type of the node to be inferred may also be used for at least one of the following: code completion, code refactoring, code inspection, or quick fix.
[0223] FIG3 illustrates a method for type inference in a programming language provided by an embodiment of the present application. The method 300 shown in FIG3 can be considered as a specific implementation of the method 100. For related descriptions, reference can be made to the method 100. To avoid repetition, some descriptions of the method 300 are omitted as appropriate.
[0224] To better illustrate the solution of the embodiment of the present application, method 300 is exemplified below, primarily with reference to the following Python code. For example, method 300 may be triggered by opening a source code file. Code segment #1 may be the code segment currently displayed in the IDE code edit box. Method 300 is primarily described using this scenario as an example and does not limit the solution of the embodiment of the present application.
[0225] Code Snippet #1
[0226] 310, obtaining the node to be inferred of the AST.
[0227] 320, determine whether the type of the root node of the AST has been inferred.
[0228] Get the root node from the AST and determine whether the type of the root node has been inferred based on the type environment.
[0229] For example, as shown in Figure 3, the value of the root node's type variable is read from the type environment. If the value of the root node's type variable does not exist in the type environment, the root node's type variable is not inferred. If the value of the root node's type variable exists in the type environment, the root node's type variable is inferred.
[0230] Taking code snippet #1 as an example, obtain the root node of code snippet #1, CustomConnection, from the AST. Determine whether the type of CustomConnection has been inferred based on the type context.
[0231] If the type of the root node has been inferred, the current inference process can be ended.
[0232] Furthermore, for example, as shown in FIG3 , if the type of the root node has been inferred, the inference result may be output, for example, the types of all fields in code segment #1.
[0233] If the type of the root node has not been inferred, step 330 is executed.
[0234] 330, creating an inference item.
[0235] If step 330 is triggered by step 320 , an inference item of the root node is created.
[0236] Taking code segment #1 as an example, for example, in step 330 , an inference item of CustomConnection may be created, and the inference item is used to indicate the inference task of code segment #1.
[0237] If step 330 is triggered by step 380 , an inference item of the dependency item is created.
[0238] 340, adding the inference item to the inference queue.
[0239] For example, as shown in FIG3 , the inference queue includes inference item 1 , inference item 2 , inference item 3 , and inference item 4 .
[0240] 350, extract the first inference item from the inference queue.
[0241] Taking the inference queue shown in FIG3 as an example, the current first inference item is inference item 1, and inference item 1 is extracted from the inference queue.
[0242] After extracting the inference item from the inference queue, the inference task indicated by the inference item is executed. The following uses the inference item of the root node as an example to illustrate.
[0243] 360, perform type variable initialization.
[0244] Specifically, a new type variable, ie, a unique free type variable, is created for each node to be inferred.
[0245] For example, the node to be inferred may be a node corresponding to a class, a static variable, an instance class variable, a function, a function parameter, or a return value, etc. In step 360, a unique free type variable may be created for each class, static variable, instance class variable, function, function parameter, or return value, etc.
[0246] Taking the inference item currently being processed as CustomConnection as an example, create unique free type variables for the class, static variable, instance class variable, function, function parameter or return value in code segment #1 respectively.
[0247] The code snippet of the CustomConnection class after initialization can be represented by the following code snippet #2.
[0248] Code Snippet #2
[0249] It should be understood that the above code segment #2 is only for visually demonstrating the creation of type variables, and does not actually modify the code segment of the CustomConnection class in code segment #1 to code segment #2.
[0250] Among them, α, β, γ, Φ, T, All are unique free type variables. The nodes to be inferred include: _ConnectionBase, _write, _read, method _send, method _recv, buf, write, size, and receiver. For example, create a unique free type variable for _ConnectionBase The relationship between the two is represented by _ConnectionBase: For example, create a unique free type variable for _write The relationship between the two is expressed as _write: Similarly, the corresponding relationships between other nodes to be inferred and the only free type variables are expressed as follows: _read: buf:Φ;write: method_send: T; size: α; receiver: β; method_recv: γ.
[0251] 370, Get constraints.
[0252] The following uses some of the nodes to be inferred in code segment #1 as an example for explanation.
[0253] 1)_ConnectionBase;
[0254] Create a base class type constraint for _ConnectionBase.
[0255] 2)_write and _read;
[0256] 2.1) _write depends on _multiprocessing and creates a resolve reference constraint for _multiprocessing. _read depends on _multiprocessing and creates a resolve reference constraint for _multiprocessing. The resolve reference constraint of _multiprocessing can be used to indicate a reference to _multiprocessing.
[0257] 2.2) Create an attribute read constraint for _multiprocessing.send. The attribute read constraint of _multiprocessing.send can be used to indicate the type of the return value of the send attribute of _multiprocessing. Create an attribute read constraint for _multiprocessing.recv. The attribute read constraint of _multiprocessing.recv can be used to indicate the type of the return value of the recv attribute of _multiprocessing.
[0258] 2.3) Create an assignment constraint for "_write = _multiprocessing.send". Create an assignment constraint for "_read = _multiprocessing.recv".
[0259] 3)method_send;
[0260] In the method _send, the parameter write is assigned a default value (_write), creating a default value constraint for "write = _write".
[0261] write depends on _write, and a resolution reference constraint is created for _write. The resolution reference constraint of _write is used to indicate the reference to _write.
[0262] Other constraints can be created in the _send method, which are not described here. For specific examples, please refer to the description of the _recv method later.
[0263] 4)method_recv;
[0264] 4.1) Create the following constraints based on receiver._init_receiver(CustomConnection._read, size). In other words, receiver._init_receiver(CustomConnection._read, size) is decomposed into the following constraints.
[0265] Create resolve reference constraints for receiver, CustomConnection, and size respectively.
[0266] Create attribute read constraints for receiver._init_receiver and CustomConnection._read respectively.
[0267] Create a parameter passing constraint for (CustomConnection._read, size).
[0268] Create a calling constraint for init_receiver(*).
[0269] 4.2) Create the following constraints based on return receiver.process_receive_operation(). In other words, return receiver.process_receive_operation is decomposed into the following constraints.
[0270] Create a resolve referential constraint for the receiver.
[0271] Create an attribute read constraint for receiver.process_receive_operation.
[0272] Create a calling constraint for process_receive_operation().
[0273] Create a return constraint for the _recv return value.
[0274] In step 370 , constraints related to the type variables generated in step 360 may be collected, and then processing may be performed based on these constraints to obtain inference results of these type variables.
[0275] 380, dependency check.
[0276] Dependency checking involves processing the resolved referential constraints of the current inference item. Specifically, it checks whether the types of all the dependencies of the current inference item have been inferred.
[0277] If there are dependencies that have not yet been inferred, the current process is stopped and the type of the dependency is inferred. The type of the dependency can be obtained using the method in the embodiment of the present application. For example, steps 330 to 390 are executed for the dependency to obtain the type of the dependency.
[0278] If there is no inference item that has not yet been inferred, for example, after the inference of the types of all dependent items of the current inference item is completed, step 390 is executed.
[0279] Taking the constraints in step 370 as an example, assume that the dependencies of CustomConnection are all the dependencies indicated by the resolved reference constraints in step 370. The dependencies are as follows:
[0280] (a1)_multiprocessing, from _multiprocessing.send and _multiprocessing.recv;
[0281] (a2)_write, from write=_write;
[0282] (a3)receiver, CustomConnection and size, from receiver.init_receiver(CustomConnection._read,size);
[0283] (a4)receiver, from receiver.process_receive_operation().
[0284] The receivers in (a3) and (a4) are the same dependencies.
[0285] Checks whether the types of _multiprocessing, _write, receiver, CustomConnection, and size have been inferred.
[0286] If there are dependencies in _multiprocessing, _write, receiver, CustomConnection, and size that have not yet been inferred, the current processing is stopped and the type of the dependency is inferred instead.
[0287] Assuming that there is a dependency item size that has not yet been inferred, stop the current processing (i.e., the processing of the inference item of CustomConnection) and infer the type of size instead. For example, create an inference item of size, which indicates the inference task of size. Add the inference item of size to the inference queue. After taking out the inference item of size, obtain the constraints, and obtain the inference result of size based on the constraints. For the specific inference process, please refer to the description of steps 330 to 390. To avoid repetition, it will not be repeated here.
[0288] The above is just an example. If there are multiple unfinished dependencies for _multiprocessing, _write, receiver, CustomConnection, and size, the types of each dependency are inferred separately. For example, if the types of _multiprocessing, _write, receiver, CustomConnection, and size are all unfinished, you can create inference items for _multiprocessing, _write, receiver, CustomConnection, and size, respectively, and add them to the inference queue. After the inference items are removed from the inference queue, the corresponding inference tasks are executed.
[0289] After all dependencies of the CustomConnection inference item are inferred, that is, after the types of _multiprocessing, _write, receiver, CustomConnection, and size are inferred, step 390 is executed.
[0290] 390, inference processing.
[0291] In step 390 , other constraints of the current inference item except for the resolved reference constraint may be processed to obtain an inference result.
[0292] The following uses the code snippet "_read=_multiprocessing.recv" and the code snippet of the method _recv as examples for explanation.
[0293] 1)_read=_multiprocessing.recv;
[0294] As mentioned above, the constraints associated with the above code snippet include: a resolve reference constraint on _multiprocessing , a read attribute constraint on _multiprocessing.recv , and an assignment constraint on _read = _multiprocessing.recv .
[0295] The resolved reference constraint of _multiprocessing has been processed in step 380, resulting in the type of _multiprocessing being the _multiprocessing package.
[0296] According to the attribute reading constraint of _multiprocessing.recv, the type of the attribute recv function is determined from the type of _multiprocessing (that is, the _multiprocessing package).
[0297] According to the assignment constraint of _read = _multiprocessing.recv, the type on the left side of the "=" is combined with the type on the right side of the "=". The type on the left side of the "=", that is, the type of _read, is the only free type variable. Assuming the type on the right side of "=" is type B, the result after union can be expressed as Represents the only free type variable Contains B type.
[0298] 2)method_recv;
[0299] 2.1)receiver._init_receiver(CustomConnection._read,size);
[0300] As mentioned above, the constraints related to the code snippet receiver._init_receiver(CustomConnection._read,size) include: the resolve reference constraint of receiver, the resolve reference constraint of CustomConnection, the resolve reference constraint of size, the attribute read constraint of receiver._init_receiver, the attribute read constraint of CustomConnection._read, the parameter passing constraint of (CustomConnection._read,size), and the call constraint of init_receiver(*).
[0301] The resolve reference constraints of receiver, the resolve reference constraints of CustomConnection, and the resolve reference constraints of size have been processed in step 380.
[0302] According to the attribute read constraint of CustomConnection._read, the type of the attribute _read is determined from the type of CustomConnection.
[0303] Based on the attribute read constraint of receiver._init_receiver, the type of the attribute _init_receiver is determined from the type of receiver. As previously mentioned, receiver is a unique free type variable β. When requesting the type of an attribute from a unique free type variable, it can be determined that the unique free type variable includes the type of the attribute. Accordingly, β includes the type of _init_receiver. Based on the call constraint of init_receiver(*), the type of _init_receiver is determined to be a callable function, resulting in β = {_init_receiver:Callable}, indicating that β includes the type of _init_receiver and that _init_receiver is a callable function. The parameter passing constraint of (CustomConnection._read,size) can be used to determine the type of the parameter of _init_receiver.
[0304] 2.2)return receiver.process_receive_operation();
[0305] As mentioned above, the constraints related to the code snippet return receiver.process_receive_operation() include: the resolved reference constraint of receiver, the attribute read constraint of receiver.process_receive_operation, the call constraint of process_receive_operation(), and the return constraint of _recv return value.
[0306] The receiver's resolved reference constraints have already been processed in step 380.
[0307] The type of the receiver obtained in 2.1) is expressed as β = {_init_receiver: Callable}.
[0308] According to the attribute read constraint of receiver.process_receive_operation, the type of attribute process_receive_operation is determined from the type of receiver. Accordingly, β includes the type of process_receive_operation, and β = {_init_receiver: Callable} is modified to β = {_init_receiver: Callable, process_receive_operation}.
[0309] According to the calling constraint of process_receive_operation(), process_receive_operation is determined to be a callable function, that is, process_receive_operation:Callable. Based on this, the type of receiver is modified to obtain β = {_init_receiver:Callable, process_receive_operation:Callable}.
[0310] The return constraint of _recv determines that the return type of method _recv is the return type of the method process_receive_operation when it is called at runtime. The return type of method _recv is the only free type variable γ, that is, γ is the return type of the method process_receive_operation when it is called at runtime.
[0311] For example, after the above constraints are processed, the inference result of method_recv is as follows:
[0312] T0 represents the type of size, which is an unknown type. T1 represents the return type of the process_receive_operation method, that is, the return type of the _recv method, which is an unknown type. Any represents an unknown type.
[0313] _init_receiver:Callable[[CustomConnection._read,T0],Any] indicates that init_receiver is a callable function, the parameter types are CustomConnection._read and T0, and the return type is an unknown type.
[0314] process_receive_operation:Callable[[],T1] means that process_receive_operation is a callable function with no parameters and a return type of T1.
[0315] The inference result obtained in step 390 may be stored in the type environment.
[0316] 400. Output an inference result. The inference result may indicate the type of the node to be inferred.
[0317] The apparatus of the embodiment of the present application is described below with reference to Figures 4 to 7. It should be understood that the apparatus described below can execute the method of the aforementioned embodiment of the present application. To avoid unnecessary repetition, repeated descriptions are appropriately omitted when introducing the apparatus of the embodiment of the present application.
[0318] FIG4 is a schematic block diagram of an apparatus for programming language type inference according to an embodiment of the present application. The apparatus 2000 shown in FIG4 can be used to execute the method shown in FIG1 or FIG3. The apparatus 2000 includes a first acquisition module 2010, a second acquisition module 2020, and a processing module 2030.
[0319] In one possible implementation, the apparatus 2000 may be used to execute the method shown in FIG. 1 .
[0320] The first acquisition module 2010 is used to acquire a node to be inferred in the abstract syntax tree, where the node to be inferred corresponds to a field to be inferred in the program code. The field to be inferred is one of the following: a variable, a class, a function, or a parameter.
[0321] The second acquisition module 2020 is configured to acquire constraints related to the node to be inferred, where the constraints are used to indicate the relationship between the node to be inferred in the abstract syntax tree.
[0322] The processing module 2030 is configured to determine the type of the node to be inferred based on the constraint. The type of the node to be inferred is one of the following: a numeric type, a sequence type, a mapping type, a collection type, an iterator type, a context manager type, a module type, a metatype, a null type, or a callable type.
[0323] The first acquisition module 2010 and the second acquisition module 2020 may be the same module or different modules.
[0324] Optionally, the type of the node to be inferred is used for at least one of the following: displaying the type of the node to be inferred, code completion, code refactoring, code inspection, or quick fix.
[0325] Optionally, the first acquisition module 2010 is specifically configured to take out a first inference item from an inference queue, where the first inference item is used to determine the type of the node to be inferred. The inference queue includes one or more inference items; and acquire the node to be inferred.
[0326] Optionally, the constraint includes a parsed reference constraint of the first node in the abstract syntax tree, and the parsed reference constraint of the first node is used to indicate the reference relationship of the node to be inferred to the first node. The device 2000 also includes a third acquisition module (not shown in the figure) for obtaining the type of the first node, and the processing module 2030 is specifically used to determine the type of the node to be inferred based on the type of the first node and the constraint.
[0327] Optionally, the constraint includes an assignment constraint of the second node, the assignment constraint of the second node is used to indicate the assignment relationship involved in the second node, and the type of the assigned child node in the second node is determined according to the type on the right side of the assignment number.
[0328] Optionally, the constraint includes an attribute read constraint of the third node, the attribute read constraint of the third node is used to indicate a subordinate relationship involved in the third node, and the type of the owner of the attribute in the third node includes the type of the attribute.
[0329] Optionally, the device 2000 also includes a creation module (not shown in the figure) for creating a type variable corresponding to the node to be inferred; and a processing module 2030 is specifically used to determine the value of the type variable corresponding to the node to be inferred based on the constraint, and the value of the type variable corresponding to the node to be inferred is used to indicate the type of the node to be inferred.
[0330] Optionally, the apparatus 2000 further includes a storage module (not shown in the figure) for storing the type of the node to be inferred in a type environment, where the type environment is used to store the types of nodes in the abstract syntax tree.
[0331] Each module in the apparatus 2000 may be implemented by software or hardware. For example, the implementation of the processing module 2030 will be described below using the processing module 2030 as an example. Similarly, the implementation of other modules may refer to the implementation of the processing module 2030.
[0332] As an example of a software functional unit, the processing module 2030 may include code running on a computing instance. The computing instance may include at least one of a physical host (computing device), a virtual machine, and a container. Furthermore, the computing instance may be one or more. For example, the processing module 2030 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code may be distributed in the same region or in different regions. Furthermore, the multiple hosts / virtual machines / containers used to run the code may be distributed in the same availability zone (AZ) or in different AZs, each AZ including one data center or multiple geographically close data centers. Typically, a region may include multiple AZs.
[0333] Similarly, multiple hosts / virtual machines / containers running the code can be distributed within the same virtual private cloud (VPC) or across multiple VPCs. Typically, a VPC is set up within a region. Cross-region communication between two VPCs within the same region, or between VPCs in different regions, requires a communication gateway within each VPC to interconnect the VPCs.
[0334] As an example of a hardware functional unit, processing module 2030 may include at least one computing device, such as a server. Alternatively, processing module 2030 may be implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.
[0335] The multiple computing devices included in processing module 2030 can be distributed in the same region or in different regions. Similarly, the multiple computing devices included in processing module 2030 can be distributed in the same virtual private computer (VPC) or in multiple VPCs. The multiple computing devices can be any combination of servers, ASICs, PLDs, CPLDs, FPGAs, GALs, and other computing devices.
[0336] It should be noted that, in other embodiments, the processing module 2030 can be used to execute any step in the method of type inference of a programming language, the first acquisition module 2010 can be used to execute any step in the method of type inference of a programming language, and the second acquisition module 2020 can be used to execute any step in the method of type inference of a programming language. The steps that each module is responsible for implementing can be specified as needed, and all functions of the device 2000 are realized by each module implementing different steps in the method of type inference of a programming language.
[0337] This application also provides a computing device 1000. As shown in FIG5 , computing device 1000 includes a bus 1002, a processor 1004, a memory 1006, and a communication interface 1008. Processor 1004, memory 1006, and communication interface 1008 communicate with each other via bus 1002. Computing device 1000 can be a server or a terminal device. It should be understood that this application does not limit the number of processors and memories in computing device 1000.
[0338] Bus 1002 may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, among others. Buses may be classified as address buses, data buses, control buses, and the like. For ease of illustration, FIG5 shows a single line, but this does not imply a single bus or type of bus. Bus 1002 may include a path for transmitting information between various components of computing device 1000 (e.g., memory 1006, processor 1004, and communication interface 1008).
[0339] The processor 1004 may include any one or more processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).
[0340] The memory 1006 may include volatile memory, such as random access memory (RAM). The processor 1004 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD).
[0341] Memory 1006 stores executable program code. Processor 1004 executes the executable program code to implement the functions of the aforementioned first acquisition module 2010, second acquisition module 2020, and processing module 2030, thereby implementing the method for type inference in a programming language. In other words, memory 1006 stores instructions for executing the method for type inference in a programming language.
[0342] The communication interface 1008 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 1000 and other devices or a communication network.
[0343] Embodiments of the present application also provide a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smartphone.
[0344] As shown in Figure 6, the computing device cluster includes at least one computing device 1000. The memory 1006 in one or more computing devices 1000 in the computing device cluster may store the same instructions for executing the method for type inference of a programming language.
[0345] In some possible implementations, the memory 1006 of one or more computing devices 1000 in the computing device cluster may also store partial instructions for executing the method for type inference in a programming language. In other words, the combination of one or more computing devices 1000 can jointly execute the instructions for executing the method for type inference in a programming language.
[0346] It should be noted that the memory 1006 in different computing devices 1000 in the computing device cluster may store different instructions, each for executing a portion of the functions of the apparatus for programming language type inference. In other words, the instructions stored in the memory 1006 in different computing devices 1000 may implement the functions of one or more of the first acquisition module 2010, the second acquisition module 2020, and the processing module 2030.
[0347] In some possible implementations, one or more computing devices in a computing device cluster may be connected via a network. The network may be a wide area network (WAN) or a local area network (LAN), etc. FIG. 7 illustrates a possible implementation. As shown in FIG. 7 , two computing devices 1000A and 1000B are connected via a network. Specifically, the network is connected via a communication interface in each computing device. In this type of possible implementation, the memory 1006 in the computing device 1000A stores instructions for executing the functions of the first acquisition module 2010. Simultaneously, the memory 1006 in the computing device 1000B stores instructions for executing the functions of the processing module 2030 and the second acquisition module 2020.
[0348] The connection method between the computing device clusters shown in Figure 7 can be that considering that the type inference method of the programming language provided in this application may need to store data, it is considered to hand over the functions implemented by the processing module 2030 and the second acquisition module 2020 to the computing device 1000B for execution.
[0349] It should be understood that the functions of the computing device 1000A shown in FIG7 may also be completed by multiple computing devices 1000. Similarly, the functions of the computing device 1000B may also be completed by multiple computing devices 1000.
[0350] Embodiments of the present application also provide a computer program product comprising instructions. The computer program product may be software or a program product comprising instructions that can be run on a computing device or stored in any available medium. When the computer program product is run on at least one computing device, the computer program product causes the at least one computing device to perform a method for type inference in a programming language.
[0351] The embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that can be stored by a computing device or a data storage device such as a data center that contains one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive). The computer-readable storage medium includes instructions that instruct the computing device to perform a method for type inference in a programming language.
[0352] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the protection scope of the technical solutions of the embodiments of the present application.
[0353] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0354] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0355] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0356] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0357] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0358] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0359] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
Claims
1. A method for type inference in a programming language, characterized in that: The method is applied to an integrated development environment IDE, and the method comprises: In a process of developing a program code, a node to be inferred in an abstract syntax tree is obtained, wherein the node to be inferred indicates a field to be inferred in the program code, and the field to be inferred is one of the following: a variable, a class, a function, or a parameter; Acquire a constraint related to the node to be inferred, wherein the constraint is used to indicate a relationship involved in the node to be inferred in the abstract syntax tree; The type of the node to be inferred is determined according to the constraint, and the type of the node to be inferred is one of the following: a numeric type, a sequence type, a mapping type, a collection type, an iterator type, a context manager type, a module type, a meta type, a null value type, or a callable type.
2. The method according to claim 1, characterized in that The type of the node to be inferred is used for at least one of the following: displaying the type of the node to be inferred, code completion, code refactoring, code inspection, or quick repair.
3. The method according to claim 1 or 2, characterized in that: The step of obtaining a node to be inferred in the abstract syntax tree includes: Taking out a first inference item from an inference queue, where the first inference item is used to determine the type of the node to be inferred, and the inference queue includes one or more inference items; The node to be inferred is obtained.
4. The method according to any one of claims 1 to 3, characterized in that The constraint includes a parsed reference constraint of a first node in the abstract syntax tree, the parsed reference constraint of the first node is used to indicate a reference relationship between the node to be inferred and the first node, and the method further includes: Acquiring the type of the first node, and determining the type of the node to be inferred according to the constraint, includes: The type of the node to be inferred is determined according to the type of the first node and the constraint.
5. The method according to any one of claims 1 to 4, characterized in that The constraint includes an assignment constraint of the second node, and the assignment constraint of the second node is used to indicate an assignment relationship involved in the second node, and the type of the assigned child node in the second node is determined according to the type on the right side of the assignment number.
6. The method according to any one of claims 1 to 5, characterized in that The constraint includes an attribute read constraint of a third node, the attribute read constraint of the third node is used to indicate a subordinate relationship involved in the third node, and the type of the owner of the attribute in the third node includes the type of the attribute.
7. The method according to any one of claims 1 to 6, characterized in that The method further comprises: Creating a type variable corresponding to the node to be inferred; and The determining the type of the node to be inferred according to the constraint includes: The value of the type variable corresponding to the node to be inferred is determined according to the constraint, and the value of the type variable corresponding to the node to be inferred is used to indicate the type of the node to be inferred.
8. The method according to any one of claims 1 to 7, characterized in that The method further comprises: The type of the node to be inferred is stored in a type environment, where the type environment is used to store the types of nodes in the abstract syntax tree.
9. A device for type inference of a programming language, characterized in that: The device is applied to an integrated development environment IDE, and the device comprises: A first acquisition module is used to acquire a node to be inferred in an abstract syntax tree during the development of a program code, wherein the node to be inferred indicates a field to be inferred in the program code, and the field to be inferred is one of the following: a variable, a class, a function, or a parameter; A second acquisition module, used for acquiring constraints related to the node to be inferred, wherein the constraints are used for indicating the relationship involved in the node to be inferred in the abstract syntax tree; A processing module is used to determine the type of the node to be inferred according to the constraint, and the type of the node to be inferred is one of the following: a numeric type, a sequence type, a mapping type, a collection type, an iterator type, a context manager type, a module type, a meta type, a null value type or a callable type.
10. The device according to claim 9, characterized in that The type of the node to be inferred is used for at least one of the following: displaying the type of the node to be inferred, code completion, code refactoring, code inspection, or quick repair.
11. The device according to claim 9 or 10, characterized in that The first acquisition module is specifically used for: Taking out a first inference item from an inference queue, where the first inference item is used to determine the type of the node to be inferred, and the inference queue includes one or more inference items; The node to be inferred is obtained.
12. The device according to any one of claims 9 to 11, characterized in that The constraint includes a parsed reference constraint of a first node in the abstract syntax tree, the parsed reference constraint of the first node is used to indicate a reference relationship between the node to be inferred and the first node, the apparatus further includes a third acquisition module, used to acquire a type of the first node, and the processing module is specifically used to: The type of the node to be inferred is determined according to the type of the first node and the constraint.
13. The device according to any one of claims 9 to 12, characterized in that The constraint includes an assignment constraint of the second node, and the assignment constraint of the second node is used to indicate an assignment relationship involved in the second node, and the type of the assigned child node in the second node is determined according to the type on the right side of the assignment number.
14. The device according to any one of claims 9 to 13, characterized in that The constraint includes an attribute read constraint of a third node, the attribute read constraint of the third node is used to indicate a subordinate relationship involved in the third node, and the type of the owner of the attribute in the third node includes the type of the attribute.
15. The device according to any one of claims 9 to 14, characterized in that The device also includes a creation module for Creating a type variable corresponding to the node to be inferred; and The processing module is specifically used for: The value of the type variable corresponding to the node to be inferred is determined according to the constraint, and the value of the type variable corresponding to the node to be inferred is used to indicate the type of the node to be inferred.
16. The device according to any one of claims 9 to 15, characterized in that The device also includes a storage module, which is used to store the type of the node to be inferred in a type environment, and the type environment is used to store the type of the node in the abstract syntax tree.
17. A computing device cluster, characterized in that: comprising at least one computing device, each computing device comprising a processor and a memory; The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster executes the method according to any one of claims 1 to 8.
18. A computer-readable storage medium, characterized in that: The method comprises computer program instructions, and when the computer program instructions are executed by a computing device cluster, the computing device cluster performs the method according to any one of claims 1 to 8.
19. A computer program product comprising instructions, characterized in that When the instructions are executed by a computing device cluster, the computing device cluster is caused to perform the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Python code memory leak detection method based on mode
CN113407442A
Type inference in dynamic language
CN115686467A
Python code static analysis method and device
CN116303053A
Type inference system and method
US20070234288A1
RU2023132940A