Modeling Java source code with symbolic description language

By modeling Java source code through the Symbolic Description Language (SDL), the problem of missing language constructs when the Java compiler generates bytecode is solved, extensive reflection operations and program transformations are achieved at runtime, and compatibility and flexibility between different compilers are supported.

CN120677457APending Publication Date: 2025-09-19ORACLE INT CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202480012203.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-02-13
Filing Date
2024-02-12
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

In the existing technology, the bytecode generated by the Java compiler loses many language constructs, such as lambda expressions, try/catch/finally blocks, loops, etc., resulting in the reflection operation being unable to provide sufficient information to generate new code that retains the language structure of the original code. In addition, the translation strategies of different compilers lead to inconsistent bytecodes, and the self-organizing solution cannot be generalized.

Method used

The Symbolic Description Language (SDL) is used to model Java source code and generate SDL representation, which retains the language constructs lost during compilation to bytecode and retrieves them at runtime. The JVM is modified to perform reflection operations on the SDL representation, avoiding direct reflection on the bytecode.

Benefits of technology

It realizes extensive reflection operations on Java programs at runtime, generates differentiated programs, optimized programs, etc., solves the problem of insufficient information in existing technologies, and supports compatibility and flexibility between different compilers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120677457A_ABST
    Figure CN120677457A_ABST
Patent Text Reader

Abstract

Techniques for modeling Java source codes in a symbol description language are disclosed, comprising: obtaining a set of Java source codes; determining that the group of Java source codes contains a user-defined type; determining that the group of Java source codes contains a loop; a Symbol Description Language (SDL) model is generated based on the set of Java source codes, the model containing a user defined type of SDL representation and a loop of SDL representation.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application includes subject matter related to that disclosed in U.S. Patent Application No. 18 / 168,161, entitled "TRANSFORMING JAVA SOURCE CODE USING A SYMBOLIC DESCRIPTION LANGUAGE MODEL," which is hereby incorporated by reference herein. Technical Field

[0003] The present disclosure relates to source code modeling. In particular, the present disclosure relates to modeling Java source code using a symbolic description language. Background Art

[0004] Generally speaking, a Java compiler takes Java source code and generates bytecode that is compiled according to the specifications of the Java Virtual Machine (JVM). During the process of compiling source code, much of the information contained in the source code may be lost. For example, Java bytecode does not preserve language constructs such as lambda expressions, try / catch / finally blocks, loops, patterns, and so on.

[0005] Code reflection is the process of examining code at runtime (i.e., while the JVM is executing the bytecode) to determine one or more characteristics of the code. For example, given a runtime object of unknown type, reflection can determine the type name and whether that type contains a particular method. Generally speaking, reflection typically only allows runtime querying to obtain "surface" details of a class, such as the type of the class itself, its fields, its method declarations, etc. The code of the method body itself is opaque and cannot be queried via reflection.

[0006] In some cases, due to the limitations of reflection, developers may resort to ad-hoc solutions to try to obtain more information. For example, developers can write code to obtain the bytecode of a method. Because the Java platform does not provide any standard way to access the bytecode of a method, such solutions must be ad-hoc and / or platform-dependent. Even so, bytecode is designed to be executed by the Java virtual machine, and the process of compiling source code into bytecode destroys information such as structure and types. In addition, different Java compiler implementations use different translation strategies, resulting in different bytecodes even if the program meaning is preserved between different compilers according to the Java specification. Because bytecode does not preserve all language constructs, solutions that rely on bytecode are limited to information that is not destroyed when compiled into bytecode. Moreover, such ad-hoc solutions cannot be generalized to other cases and may require modifications to the compiler itself.

[0007] Given the above, the information available to Java programs via reflection (including ad hoc solutions) is often insufficient to achieve the desired purpose. For example, reflection does not provide sufficient information to generate new code that preserves the language structure of the original code. Reflection also does not provide sufficient information to generate transformations of the original program, such as differentiated programs, optimized programs, functionally similar programs compiled according to different language specifications, and so on.

[0008] The approaches described in this section are approaches that could be pursued, but are not necessarily approaches that have been previously conceived or pursued. Therefore, unless otherwise indicated, it should not be assumed that any approach described in this section qualifies as prior art merely by virtue of its inclusion in this section. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] The embodiments are illustrated in the figures of the accompanying drawings by way of example and not limitation. Reference to "one" or "an" embodiment in this disclosure does not necessarily refer to the same embodiment and means at least one. In the drawings:

[0010] Figure 1 An example computing architecture is illustrated in which the techniques described herein may be practiced.

[0011] Figure 2 is a block diagram illustrating one embodiment of a computer system suitable for implementing the methods and features described herein.

[0012] Figure 3 Illustrated is an example virtual machine memory layout in block diagram form, according to an embodiment.

[0013] Figure 4 An example framework in block diagram form is illustrated according to an embodiment.

[0014] Figure 5 illustrates a block diagram of an SDL mode of operation according to one or more embodiments;

[0015] Figure 6 illustrates a set of example operations for modeling Java source code using a symbolic description language according to one or more embodiments;

[0016] Figures 7A-7B An example of modeling Java source code using a symbolic description language according to one or more embodiments is illustrated; and

[0017] Figure 8 A block diagram illustrating a computer system in accordance with one or more embodiments is shown. DETAILED DESCRIPTION

[0018] In the following description, numerous specific details are set forth for purposes of explanation and to provide a thorough understanding. One or more embodiments may be practiced without these specific details. Features described in one embodiment may be combined with features described in different embodiments. In some examples, well-known structures and devices are described in block diagram form to avoid unnecessarily obscuring the present invention.

[0019] The following table is provided for reference only and should not be construed as limiting the scope of one or more embodiments.

[0020] 1. General Overview…………………………………………6

[0021] 2. System Architecture Overview……………………………………6

[0022] 2.1. Example System Architecture……………………………………6

[0023] 2.2. Example class file structure………………………………10

[0024] 2.3. Example Virtual Machine Architecture…………………………13

[0025] 2.4. Loading, linking and initialization…………………………16

[0026] 3. Symbolic Description Language……………………………………19

[0027] 3.1. General SDL Features……………………………………19

[0028] 3.2. Example SDL Mode …………………………………………19

[0029] 4. Modeling Java source code using symbolic description language …………23

[0030] 4.1. Java Core Dialects 25

[0031] 4.1.1. Modeling Java Methods 26

[0032] 4.1.2. Static Methods 27

[0033] 4.1.3. Instance Methods 27

[0034] 4.1.4. Generic Methods 28

[0035] 4.1.5. Annotation Method……………………………………29

[0036] 4.1.6. Exceptions ……………………………………………… 29

[0037] 4.1.7. Modeling Field Access 29

[0038] 4.1.8. Modeling method calls 30

[0039] 4.1.9. Reflection Operation……………………………………30

[0040] 4.1.10. Parsing and Access Control ………………………… 31

[0041] 4.1.11. Static and instance methods and fields 32

[0042] 4.1.12. Coercion and Conversion ..................................................32

[0043] 4.1.13. Preventing Heap Pollution 32

[0044] 4.1.14. Translation into bytecode………………………………33

[0045] 4.1.15. Modeling Lambda Expressions 33

[0046] 4.1.16. Reference Operation……………………………………34

[0047] 4.1.17. Modeling Local Variables 35

[0048] 4.2. Advanced Java Dialects……………………………………36

[0049] 4.2.1. Modeling the loop .................................................................. 37

[0050] 4.2.2. Modeling the Enhanced FOR Loop 37

[0051] 4.2.3. Modeling a Counting FOR Loop 40

[0052] 4.2.4. Modeling the WHILE Loop 41

[0053] 4.2.5. Modeling IF-THEN and IF-THEN-ELSE Statements 43

[0054] 4.2.6. Modeling TRY / CATCH / FINALLY 46

[0055] 5. Example operation of modeling Java source code with SDL…………49

[0056] 6. Example Embodiments 53

[0057] 7. Application Examples…………………………………………54

[0058] 8. Additional Examples…………………………………………55

[0059] 9. Machine Learning ………………………………………………57

[0060] 10. Computer Networks and Cloud Networks…………………………57

[0061] 11. Hardware Overview……………………………………61

[0062] 12. Other matters; extension …………………………………… 65

[0063] 1. General Overview

[0064] One or more embodiments generate a Symbolic Description Language (SDL) representation of Java source code. For example, a Java compiler can generate an SDL representation from the source code of a given method or lambda expression's body. The SDL representation of Java source code preserves language constructs (e.g., structure and type information) that might otherwise be lost during compilation to bytecode and can be retrieved at runtime. Thus, the SDL representation allows for more reflective operations than would otherwise be possible.

[0065] Using SDL representations of Java programs, the system can generate transformations of these programs. For example, the system can generate differentiated programs, optimized programs, functionally similar programs using different programming languages, etc., without the need for a self-assembled solution. SDL generation can be added to the Java compiler and / or provided as a standalone tool. The JVM can be modified so that reflection operations are performed on the corresponding SDL representation rather than on the bytecode. In order to access an existing SDL representation, the system can first use standard reflection to obtain the SDL representation. For example, in order to access the SDL representation of a method body (if it exists), the system can first obtain a java.lang.reflect.Method instance. The system can then query the reflection object to obtain its SDL representation. Therefore, one or more embodiments allow a wide range of runtime functionality that was not originally a standard part of the Java development and runtime environment.

[0066] One or more embodiments described in this specification and / or claimed may not be included in this general overview.

[0067] 2. System Architecture Overview

[0068] 2.1 Example Architecture

[0069] Figure 1 The diagram illustrates an example architecture in which the techniques described herein may be practiced. The software and / or hardware components described with respect to the example architecture may be omitted or associated with a different set of functionality than that described herein. According to one or more embodiments, software and / or hardware components not described herein may be used within the environment. Therefore, the example environment should not be construed as limiting the scope of any claims.

[0070] like Figure 1 As shown, a computing architecture 100 includes source code files 101, which are compiled by a compiler 102 into class files 103 representing a program to be executed. The class files 103 are then loaded and executed by an execution platform 112, which includes a runtime environment 113, an operating system 111, and one or more application programming interfaces (APIs) 110 that enable communication between the runtime environment 113 and the operating system 111. The runtime environment 113 includes a virtual machine 104, which includes various components such as a memory manager 105 (which may include a garbage collector), a class file validator 106 that checks the validity of the class files 103, a class loader 107 that locates and builds in-memory representations of classes, an interpreter 108 for executing the code of the virtual machine 104, and a just-in-time (JIT) compiler 109 for producing optimized machine-level code.

[0071] In an embodiment, computing architecture 100 includes source code files 101 that contain code that has been written in a particular programming language (such as Java, C, C++, C#, Ruby, Perl, etc.). Thus, source code files 101 follow a particular set of grammatical and / or semantic rules for the associated language. For example, code written in Java follows the Java language specification. However, because specifications are updated and revised over time, source code files 101 may be associated with a version number that indicates the revision of the specification that source code files 101 follow. The exact programming language used to write source code files 101 is generally not critical.

[0072] In various embodiments, the compiler 102 converts source code written according to a specification for the convenience of the programmer into a machine or object code that can be directly executed by a specific machine environment, or an intermediate representation ("virtual machine code / instructions"), such as bytecode, that can be executed by a virtual machine 104 that can run on various specific machine environments. The virtual machine instructions can be executed by the virtual machine 104 in a more direct and efficient manner than the source code. Converting the source code to the virtual machine instructions includes mapping the source code functionality from the language to the virtual machine functionality that utilizes the underlying resources (such as data structures). Typically, the functionality presented in simple terms by the programmer via the source code is converted into more complex steps that map more directly to the instruction set supported by the underlying hardware on which the virtual machine 104 resides.

[0073] Generally speaking, programs are executed as compiled or interpreted programs. When a program is compiled, the code is globally transformed from a first language to a second language before execution. Because the work of transforming the code is performed in advance, compiled code often has excellent runtime performance. In addition, because the transformation occurs globally before execution, techniques such as constant folding, dead code elimination, and inlining can be used to analyze and optimize the code. However, depending on the program being executed, startup time can be significant. In addition, inserting new code requires taking the program offline, recompiling, and re-executing it. For many dynamic languages ​​(such as Java) that are designed to allow code to be inserted during program execution, a purely compiled solution may not be suitable. When a program is interpreted, the program's code is read line by line and converted into machine-level instructions while the program is executing. As a result, the program has a short startup time (can begin execution almost immediately), but runtime performance is reduced due to the transformation being performed on the fly. In addition, because each instruction is analyzed separately, many optimizations that rely on a more global analysis of the program cannot be performed.

[0074] In some embodiments, the virtual machine 104 includes an interpreter 108 and a JIT compiler 109 (or components that implement these two aspects), and uses a combination of interpretation and compilation techniques to execute the program. For example, the virtual machine 104 can initially start by interpreting the virtual machine instructions representing the program via the interpreter 108, while tracking statistical information related to program behavior, such as how frequently different code sections or code blocks are executed by the virtual machine 104. Once a code block exceeds a threshold (becomes "hot"), the virtual machine 104 calls the JIT compiler 109 to perform an analysis of the block and generate optimized machine-level instructions, which replace the "hot" code block for future execution. Since programs often spend most of their time executing a small part of the entire code, only compiling the "hot" part of the program can provide performance similar to that of fully compiled code, but without the startup penalty. In addition, although the optimization analysis is constrained to the "hot" block being replaced, there is still a much greater optimization potential than converting each instruction separately. There are multiple variants in the above example, such as layered compilation.

[0075] To provide a clear example, source code file 101 has been illustrated as a "top-level" representation of a program to be executed by execution platform 112. Although computing architecture 100 depicts source code file 101 as a "top-level" program representation, in other embodiments, source code file 101 may be an intermediate representation received via a "higher-level" compiler that processes a code file in a different language into the language of source code file 101. Some examples in the following disclosure assume that source code file 101 conforms to a class-based, object-oriented programming language. However, this is not a requirement for utilizing the features described herein.

[0076] In an embodiment, compiler 102 receives source code file 101 as input and converts source code file 101 into class file 103 in a format expected by virtual machine 104. For example, in the context of the JVM, the Java Virtual Machine Specification defines a specific class file format that class file 103 is expected to follow. In some embodiments, class file 103 contains virtual machine instructions that have been converted from source code file 101. However, in other embodiments, class file 103 may also contain other structures, such as tables identifying constant values ​​and / or metadata associated with various structures (classes, fields, methods, etc.).

[0077] The following discussion assumes that each of the class files 103 represents a corresponding "class" defined in the source code file 101 (or dynamically generated by the compiler 102 / virtual machine 104). However, the aforementioned assumption is not a strict requirement and will depend on the implementation of the virtual machine 104. Therefore, no matter how the exact format of the class file 103 is, the technology described herein can still be executed. In some embodiments, the class file 103 is divided into one or more "libraries" or "packages", each of which includes a collection of classes that provide related functionality. For example, a library can contain one or more class files that implement input / output (I / O) operations, mathematical tools, cryptographic techniques, graphics utilities, etc. In addition, some classes (or fields / methods within these classes) can include access restrictions that limit their use to specific classes / libraries / packages or to classes with appropriate permissions.

[0078] 2.2 Example class file structure

[0079] Figure 2 An example structure of a class file 200 is illustrated in block diagram form according to an embodiment. To provide a clear example, the remainder of this disclosure assumes that the class file 103 of the computing architecture 100 follows the structure of the example class file 200 described in this section. However, in an actual environment, the structure of the class file 200 will depend on the implementation of the virtual machine 104. In addition, one or more features discussed herein may modify the structure of the class file 200 to, for example, add additional structure types. Therefore, the exact structure of the class file 200 is not critical to the techniques described herein. For the purposes of Section 2.1, a "class" or "current class" refers to the class represented by the class file 200.

[0080] exist Figure 2 In the embodiment, class file 200 includes constant table 201, field structure 208, class metadata 207 and method structure 209. In an embodiment, constant table 201 is a data structure that also serves as a symbol table of a class in addition to other functions. For example, constant table 201 can store data related to the various identifiers used in source code file 101, such as type, range, content and / or position. Constant table 201 has entries for value structure 202 (constant values ​​representing types int, long, double, float, byte, string, etc.), class information structure 203, name and type information structure 204, field reference structure 205 and method reference structure 206 derived from source code file 101 by compiler 102. In an embodiment, constant table 201 is implemented as an array that maps index i to structure j. However, the exact implementation of constant table 201 is not critical.

[0081] In some embodiments, entries of constant table 201 include structures that index other entries of constant table 201. For example, an entry in one of value structures 202 used to represent a string may hold a tag identifying its "type" as a string, and an index to one or more other value structures 202 of constant table 201 that store char, byte, or int values ​​representing the ASCII characters of the string.

[0082] In an embodiment, a field reference structure 205 of the constant table 201 holds an index into one of the class information structures 203 in the constant table 201 that represents the class that defines the field, and an index into one of the name and type information structures 204 in the constant table 201 that provide the name and descriptor of the field. A method reference structure 206 of the constant table 201 holds an index into one of the class information structures 203 in the constant table 201 that represents the class that defines the method, and an index into one of the name and type information structures 204 in the constant table 201 that provide the name and descriptor of the method. A class information structure 203 holds an index into one of the value structures 202 in the constant table 201 that holds the name of the associated class.

[0083] The name and type information structure 204 holds an index to one of the value structures 202 in the constant table 201 that stores the name of the field / method, and an index to one of the value structures 202 in the constant table 201 that stores the descriptor.

[0084] In an embodiment, the class metadata 207 includes metadata of the class, such as (one or more) version numbers, the number of entries in the constant pool, the number of fields, the number of methods, access flags (whether the class is public, private, final, abstract, etc.), an index to one of the class information structures 203 that identifies the constant table 201 of the current class, an index to one of the class information structures 203 that identifies the constant table 201 of the superclass (if any), and the like.

[0085] In an embodiment, field structures 208 represent a set of structures that identify the various fields of a class. For each field of the class, field structures 208 store the field's accessor flags (whether the field is static, public, private, final, etc.), an index into one of the value structures 202 in constant table 201 that holds the field's name, and an index into one of the value structures 202 in constant table 201 that holds the field's descriptor.

[0086] In an embodiment, method structure 209 represents a set of structures that identify the various methods of a class. Method structure 209 stores, for each method of the class, the method's accessor flags (e.g., whether the method is static, public, private, synchronized, etc.), an index to one of the value structures 202 in constant table 201 that holds the method's name, an index to one of the value structures 202 in constant table 201 that holds the method's descriptor, and the virtual machine instructions corresponding to the method's body as defined in source code file 101.

[0087] In an embodiment, a descriptor represents the type of a field or method. For example, a descriptor can be implemented as a string that follows a specific syntax. Although the exact syntax is not important, several examples will be described below.

[0088] In the example where the descriptor represents the type of the field, the descriptor identifies the type of data held by the field. In an embodiment, a field can hold a basic type, an object, or an array. When a field holds a basic type, the descriptor is a string that identifies the basic type (e.g., "B" = byte, "C" = char, "D" = double, "F" = float, "I" = int, "J" = longint, etc.). When a field holds an object, the descriptor is a string that identifies the class name of the object (e.g., "L ClassName"). In this case, "L" indicates a reference, so "L ClassName" represents a reference to an object of class ClassName. When a field is an array, the descriptor identifies the type held by the array. For example, "[B" indicates an array of bytes, where "[" indicates an array and "B" indicates that the array holds a basic byte type. However, since arrays can be nested, the descriptor of an array can also indicate nesting. For example, "[[L ClassName" indicates an array, where each index holds an array of objects of class ClassName. In some embodiments, ClassName is fully qualified and includes the simple name of the class and the path name of the class. For example, ClassName may indicate where the file is stored in a package, library, or file system that hosts the class file 200 .

[0089] In the case of a method, the descriptor identifies the method's parameters and the method's return type. For example, a method descriptor may follow the general form "({ParameterDescriptor})ReturnDescriptor," where {ParameterDescriptor} is a list of field descriptors representing the parameters, and ReturnDescriptor is a field descriptor identifying the return type. For example, the string "V" may be used to represent a void return type. Thus, a method defined in source code file 101 as "Object m(int 1, doubled, Thread t) {...}" would match the descriptor "(ID L Thread) LObject."

[0090] In an embodiment, the virtual machine instructions held in method structure 209 include operations that reference entries of constant table 201. Using Java as an example, consider the following class:

[0091]

[0092] In the above example, the Java method add12and13 is defined in class A, takes no parameters, and returns an integer. The body of the method add12and13 calls the static method addTwo of class B, which takes the constant integer values ​​12 and 13 as parameters, and returns the result. Therefore, in the constant table 201, the compiler 102 includes, in addition to other entries, a method reference structure corresponding to the call to the method B.addTwo. In Java, calls to methods are compiled down to invoke commands in the bytecode of the JVM (invokestatic in this case, because addTwo is a static method of class B). The invoke command is provided with an index pointing to the method reference structure in the constant table 201 that identifies the class that defines addTwo "B", the name of addTwo "addTwo", and the descriptor of addTwo "(II)I". For example, assuming that the above method reference is stored at index 4, the bytecode instruction may appear as "invokestatic#4".

[0093] Because constant table 201 references classes, methods, and fields with structures that carry identifying information symbolically, rather than directly referencing memory locations, the entries of constant table 201 are called "symbolic references." One reason symbolic references are used for class files 103 is because in some embodiments, once a class is loaded into runtime environment 113, compiler 102 does not know how and where the class will be stored. As will be described in Section 2.3, after the referenced class (and associated structures) have been loaded into the runtime environment and assigned specific memory locations, the runtime representation of the symbolic reference is ultimately resolved by virtual machine 104 into an actual memory address.

[0094] 2.3 Example Virtual Machine Architecture

[0095] Figure 3 An example virtual machine memory layout 300 is illustrated in block diagram form according to an embodiment. To provide a clear example, the remaining discussion will assume that the virtual machine 104 follows Figure 3 . Furthermore, while components of virtual machine memory layout 300 may be referred to as memory "regions," there is no requirement that the memory regions be contiguous.

[0096] exist Figure 3 In the illustrated example, the virtual machine memory layout 300 is divided into a shared area 301 and a thread area 307. The shared area 301 represents an area in memory where structures shared between various threads executing on the virtual machine 104 are stored. The shared area 301 includes a heap 302 and a per-class area 303. In an embodiment, the heap 302 represents a runtime data area from which memory for class instances and arrays is allocated. In an embodiment, the per-class area 303 represents a storage area in which data related to each class is stored. In an embodiment, for each loaded class, the per-class area 303 includes a runtime constant pool 304 representing data from the constant table 201 of class, field, and method data 306 (for example, to hold static fields of the class), and method code 305 representing virtual machine instructions for the class's methods.

[0097] The thread area 307 represents a memory area in which structures specific to each thread are stored. Figure 3 , thread area 307 includes thread structure 308 and thread structure 311, which represent the per-thread structures utilized by different threads. To provide a clear example, Figure 3 The thread region 307 depicted in FIG. 3 assumes that two threads are executing on the virtual machine 104. However, in a practical environment, the virtual machine 104 may execute any arbitrary number of threads, with the number of thread structures scaling accordingly.

[0098] In an embodiment, thread structure 308 includes a program counter 309 and a virtual machine stack 310. Similarly, thread structure 311 includes a program counter 312 and a virtual machine stack 313. In an embodiment, program counter 309 and program counter 312 store the current address of the virtual machine instruction executed by their corresponding thread.

[0099] Thus, as a thread steps through instructions, the program counter is updated to maintain an index pointing to the current instruction. In an embodiment, virtual machine stack 310 and virtual machine stack 313 each store frames that hold local variables and partial results for their respective threads and also for method calls and returns.

[0100] In an embodiment, a frame is a data structure used to store data and partial results, return method values, and perform dynamic linking. Each time a method is called, a new frame is created. When the method that generated the frame is completed, the frame is destroyed. Therefore, when a thread executes a method call, the virtual machine 104 generates a new frame and pushes the frame onto the virtual machine stack associated with the thread.

[0101] When the method call is completed, the virtual machine 104 passes the result of the method call back to the previous frame and pops the current frame from the stack. In an embodiment, for a given thread, a frame is active at any point. This active frame is called the current frame, so that the method that generates the current frame is called the current method, and the class to which the current method belongs is called the current class.

[0102] Figure 4 An example frame 400 is illustrated in block diagram form according to an embodiment. To provide a clear example, the remaining discussion will assume that the frames of virtual machine stack 310 and virtual machine stack 313 adhere to the structure of frame 400.

[0103] In an embodiment, frame 400 includes local variables 401, operand stack 402 and runtime constant pool reference table 403. In an embodiment, local variables 401 are represented as a variable array, each variable holding a value, for example, a Boolean value, byte, char, short, int, float or reference value. In addition, some value types (such as long or double) can be represented by more than one entry in the array. Local variables 401 are used to pass parameters on method calls and store partial results. For example, when frame 400 is generated in response to calling a method, parameters can be stored in predetermined locations within local variables 401, such as indexes 1-N corresponding to the first to Nth parameters in the call.

[0104] In an embodiment, when the virtual machine 104 creates a frame 400, the operand stack 402 defaults to being empty. Then, the virtual machine 104 supplies instructions from the method code 305 of the current method to load constants or values ​​from local variables 401 onto the operand stack 402. Other instructions obtain operands from the operand stack 402, operate on them, and push the results back onto the operand stack 402. In addition, the operand stack 402 is used to prepare parameters to be passed to the method and to receive method results. For example, the parameters of the called method can be pushed onto the operand stack 402 before the method is called. Then, the virtual machine 104 generates a new frame for the method call, wherein the operands on the operand stack 402 of the previous frame are popped out and loaded into the local variables 401 of the new frame. When the called method terminates, the new frame is popped out from the virtual machine stack and the return value is pushed onto the operand stack 402 of the previous frame.

[0105] In an embodiment, the runtime constant pool reference table 403 contains a reference to the runtime constant pool 304 of the current class. The runtime constant pool reference table 403 is used to support resolution. Resolution is the process of converting symbol references in the constant pool 304 into specific memory addresses, thereby loading classes as needed to resolve undefined symbols and converting variable accesses to appropriate offsets in storage structures associated with the runtime locations of these variables.

[0106] 2.4 Loading, Linking, and Initialization

[0107] In an embodiment, virtual machine 104 dynamically loads, links, and initializes classes. Loading is the process of finding a class with a specific name and creating a representation of that class from its associated class file 200 within the memory of runtime environment 113. For example, a runtime constant pool 304, method code 305, and field and method data 306 are created for each class within region 303 of virtual machine memory layout 300. Linking is the process of taking the in-memory representation of a class and combining it with the runtime state of virtual machine 104 so that the class's methods can be executed. Initialization is the process of executing a class constructor to set the starting state of its fields and method data 306 and / or creating a class instance on heap 302 for the initialized class.

[0108] The following are examples of loading, linking, and initialization techniques that may be implemented by the virtual machine 104. However, in many embodiments, these steps may be interleaved such that an initial class is loaded, and then during linking, a second class is loaded to resolve symbolic references found in the first class, which in turn causes a third class to be loaded, and so on. Thus, the progress through the loading, linking, and initialization phases may vary from class to class. Furthermore, some embodiments may delay ("lazy" execution) one or more functions of the loading, linking, and initialization process until the class is actually needed. For example, resolution of a method reference may be delayed until the virtual machine instructions that call the method are executed. Thus, the exact timing of the execution steps for each class will vary significantly between various implementations.

[0109] To begin loading processing, virtual machine 104 starts by calling class loader 107 that loads the initial class. The technology of specifying the initial class will differ between embodiments. For example, one technology can make virtual machine 104 accept command line arguments specifying the initial class when starting.

[0110] To load a class, the class loader 107 parses the class file 200 corresponding to the class and determines whether the class file 200 is well-formed (meets the syntactic expectations of the virtual machine 104). If not, the class loader 107 generates an error. For example, in Java, an error may be generated in the form of an exception, which is thrown to an exception handler for processing. Otherwise, the class loader 107 generates an in-memory representation of the class by allocating a runtime constant pool 304, method code 305, and field and method data 306 for the class within a per-class area 303.

[0111] In some embodiments, when class loader 107 loads a class, class loader 107 also recursively loads the superclass of the loaded class. For example, virtual machine 104 can ensure that the superclass of a particular class is loaded, linked, and / or initialized before continuing the loading, linking, and initialization process of a particular class.

[0112] During linking, the virtual machine 104 validates the class, prepares the class, and performs resolution of symbolic references defined in the runtime constant pool 304 of the class.

[0113] To validate a class, the virtual machine 104 checks whether the class's memory representation is structurally correct. For example, the virtual machine 104 may check that each class, except for the general class Object, has a superclass; that final classes have no subclasses and that final methods are not overridden; that constant pool entries are consistent with one another; that the current class has the correct access permissions to classes / fields / structures referenced in the constant pool 304; and that the virtual machine 104 code for a method will not cause unexpected behavior (e.g., ensuring that a jump instruction does not send the virtual machine 104 beyond the end of the method). The exact checks performed during validation depend on the implementation of the virtual machine 104. In some cases, validation may result in additional classes being loaded, but it does not necessarily require that those classes be linked before proceeding. For example, suppose class A contains a reference to a static field of class B. During validation, the virtual machine 104 may check class B to ensure that the referenced static field actually exists. This may result in class B being loaded, but does not necessarily result in class B being linked or initialized. However, in some embodiments, certain validation checks may be deferred until a later stage, such as during the resolution of symbolic references. For example, some embodiments may delay checking access permissions on symbolic references until those references are resolved.

[0114] To prepare the class, the virtual machine 104 initializes the static fields located within the class and method data 306 to default values. In some cases, setting the static fields to default values ​​may be different from running the class's constructor. For example, the validation process may clear the static fields to zero or set them to the values ​​that the constructor expects those fields to have during initialization.

[0115] During parsing, the virtual machine 104 dynamically determines the specific memory address from the symbolic reference included in the class's runtime constant pool 304. To resolve the symbolic reference, the virtual machine 104 uses the class loader 107 to load the class identified in the symbolic reference (if not already loaded). Once loaded, the virtual machine 104 knows the memory location within the area 303 of each class of the referenced class and its fields / methods. The virtual machine 104 then replaces the symbolic reference with a reference to the specific memory location of the referenced class, field, or method. In an embodiment, the virtual machine 104 caches the parsed result so that it can be reused when the virtual machine 104 encounters the same class / name / descriptor when processing another class. For example, in some cases, class A and class B may call the same method of class C. Therefore, when parsing is performed on class A, the result can be cached and reused during the parsing of the same symbolic reference in class B to reduce overhead.

[0116] In some embodiments, the step of resolving symbol references during linking is optional. For example, an embodiment may perform symbol resolution in a "lazy" manner, thereby delaying the step of resolution until the virtual machine instructions requiring the referenced class / method / field are executed.

[0117] During initialization, the virtual machine 104 executes the constructor of a class to set the starting state of the class. For example, initialization can initialize the field and method data 306 of the class and generate / initialize any class instances on the heap 302 created by the constructor. For example, the class file 200 for a class can specify that a particular method is a constructor for setting the starting state. Therefore, during initialization, the virtual machine 104 executes the instructions of that constructor.

[0118] In some embodiments, the virtual machine 104 performs resolution of field and method references by initially checking whether the field / method is defined in the referenced class. Otherwise, the virtual machine 104 recursively searches the superclasses of the referenced class for the referenced field / method until the field / method is located or the top-level superclass is reached, in which case an error is generated.

[0119] 3. Symbolic Description Language

[0120] 3.1 General SDL Features

[0121] In an embodiment, a symbolic description language (SDL) uses a language of predefined symbols to describe the structure and functionality of a set of Java source code. SDL contains symbols, which are building blocks for describing Java language structures. For example, SDL can contain values, operations, bodies, and / or blocks as described below. Various arrangements of symbols can represent Java modules, packages, types (e.g., classes), methods, variables, instructions (e.g., method calls), assignments, etc. SDL can provide patterns that allow nested operations, so that SDL can represent loops and other non-linear language constructs. Therefore, the SDL representation of a particular set of Java source code preserves language constructs that are lost in bytecode. An example of an SDL pattern is described in detail below.

[0122] 3.2 Example SDL Mode

[0123] Figure 5 An example of a symbolic description language schema 500 according to one or more embodiments is illustrated. Schema 500 should be understood as a specific example and may not be applicable to certain embodiments. Accordingly, the components and / or operations described below should not be construed as limiting the scope of any claims. In this example, the building blocks of SDL schema 500 include value 526, operation 501, body 510, and block 512. Some embodiments may include more or fewer building blocks, building blocks with different names, and / or building blocks of different types.

[0124] In an embodiment, the SDL schema 500 does not specify any semantics for operation 501. Operation 501 may include:

[0125] A name 502 that uniquely identifies the definition of the operation 501 and describes the behavior of the operation.

[0126] Zero or more operands 504 , each operand 504 being a corresponding value 526 .

[0127] The result of the operation 506 is the value 526.

[0128] • Zero or more attributes 508, described in further detail below.

[0129] • Zero or more subjects 510, described in further detail below.

[0130] The body 510 includes one or more blocks 512. The first block 512 in the body 510 is referred to herein as the entry block. Each block 512 includes a unique name 514 for the block 512 and one or more operations 520. A block 512 may include zero or more arguments 516 that are values ​​526.

[0131] The last operation 520 in a block 512 is referred to herein as a terminating operation. A terminating operation contains zero or more block headers 522, which reference other blocks 512 in the same body 510 by name 514. A terminating operation can use a block header 522 to reference another block 512 as its successor. A block header 522 contains zero or more block arguments 524, each of which has a value 526 corresponding to the argument 516 of the referenced block 512. Depending on the definition of the operation 520, the block(s) 512 of the body 510 can form a control flow graph. In an embodiment, a terminating operation does not allow an entry block to be referenced as a successor. In this case, the entry block has no predecessor and is the root of the control flow graph.

[0132] In an embodiment, value 526 is assigned exactly once by operation result 506 or block parameter 516 and used by operation 501 as operand 504 and / or block argument 524. Thus, by definition, SDL supports the properties of static single assignment form (SSA), as that term applies to compiler design.

[0133] Based on the above, the symbolic description is conceptually a tree of operation → body * → block + → operation +, where * indicates zero or more nodes and + indicates one or more nodes. Depending on the definition of the operation (one or more), the blocks 512 of the body 510 can form a control flow graph. The values ​​526 form a data flow graph through their use.

[0134] A value 526 contains a type descriptor (often referred to simply as type 528) and zero or more attributes 530. A type 528 defines a set of values ​​526 such that a value 526 of that type 528 is a member of that set. Otherwise, SDL does not specify the semantics of type 528, i.e., how to determine the set of values ​​526 for a given type 528. Attributes 530 contain name / value pairs 532. Note that in this context, the values ​​contained in name / value pairs 532 are of a different kind than the values ​​526 used as operation results 506 and block parameters 516. SDL does not specify the semantics of attributes 530.

[0135] Operation 501 has a method type, whose parameter types include the types of operands 504 (in order) and a return type that is the type of operation result 506. Body 510 has a method type, whose parameter value types include the types of entry block parameters 516 (in order) and the return value type specified by the operation definition. Block 512 has a method type, whose parameter value types include the types of block parameters 516 (in order) and a return value type that is either void or "unit" (because block 512 itself has no explicit return value).

[0136] 3.2.1 Control Flow

[0137] In this example, an operation 501 containing one or more bodies 510 can enter the body 510 and pass control to an entry block 512, thereby assigning a value 526 to the entry block parameter(s) 516 (if any). Block 512 can pass control to its first operation 520. After operation 520 completes according to its definition, control is passed back to block 512. Block 512 then passes control to the next operation 520, and so on, until a terminal operation 520 is reached. A terminal operation 520 containing a block header 522 indicates that the operation 520, according to its definition, can pass control (or jump) to the referenced block 512 and pass the block arguments 524 assigned to the corresponding block parameters 516. A terminal operation 520 without a block header 522 passes control back to operation 501.

[0138] 3.2.2 Value Usage

[0139] The structural characteristics of a given value 526 determine whether the value 526 can be used as an operand 504 or a block argument 524.

[0140] As an example, a value V must be defined before an operation O can use it. O can use V if V is:

[0141] The result of an operation in O's block (e.g., B) that happens before O, or a block parameter of B, or

[0142] The result of an operation in block D that dominates B, or a block parameter of D. Here, "dominate" has the meaning used in graph theory; a node N1 dominates another node N2 if every path from the entry node to N2 passes through N1.

[0143] Otherwise, O becomes the parent of O, and the first two rules are applied recursively, traversing the tree upward. If O has no parent, then V is undefined and cannot be used. An operation definition may specify that one or more of its bodies is isolated, in which case the latter rule can be refined to terminate if the parent of O is isolated.

[0144] 3.2.3 Dialect

[0145] In an embodiment, the programming behavior of a symbolic description (i.e., a description written in SDL) is governed by the operations declared in the description (by name), the order of these operations in blocks, and the logical connections between these blocks. As used herein, a "dialect" is a set of operations and types that provides a unit of some combinatorial capability. A symbolic description may contain operations and value types from more than one dialect. The type correctness of a symbolic description is governed by the dialect types.

[0146] 3.2.4 Symbolic description form

[0147] One or more embodiments support at least two forms of symbolic descriptions: a runtime form in computer memory; and a textual form. The Java API can be configured to generate the runtime form (e.g., during compilation). Additionally or alternatively, a separate tool (e.g., a script or executable) can be configured to parse the textual description and generate the runtime form. If the Java API exposes commands for generating the runtime form, then the tool can be configured to use the Java API. The textual form can be specified using a syntax corresponding to the structure of SDL. The Java API and / or another tool can be configured to generate the textual form from the runtime form. Alternatively or additionally, a human user (e.g., a programmer) can manually generate the textual form. The textual form can be used for debugging, testing, storage, and / or transmission over a network, among other things. Furthermore, the textual form is a convenient, human-readable way to present symbolic descriptions for interpretation—including the examples of modeling Java language constructs described herein.

[0148] 4. Modeling Java source code using symbolic description language

[0149] Using a symbolic description language (SDL) such as the one described above, Java language constructs can be modeled as operations, either directly or through composition. Java language constructs can be modeled using a variety of SDL dialects (defined above). The examples described herein define two dialects: the core dialect and the advanced dialect.

[0150] Core dialects include:

[0151] Operations that model Java methods, lambda expressions, operations on primitive values, etc.

[0152] Exception regions covered by try / catch / finally blocks supported in the high-level dialect.

[0153] Definition of type descriptors, which model Java's built-in type system so that SDL type descriptions contain full type information (unlike bytecode, where reference types are erased and primitive types are reduced). In some cases, SDL cannot preserve full type fidelity because types that do not exist in the source code may appear in the compiler's abstract syntax tree (AST). Some of these types may be unrepresentable (i.e., cannot be expressed in the source code) and therefore may be too complex to support in SDL's type descriptors. Furthermore, for simplicity, SDL may use approximate representations for some representable types.

[0154] Definitions of method and field descriptors, which include type descriptors and are declared in the attributes of the operation. More details are described below in the section on reflective operations.

[0155] The high-level dialect includes operations that model Java language constructs such as loops, if / then / else code blocks, etc. These operations can use types defined by the core dialect.

[0156] One or more embodiments also include a line number attribute (e.g., "line.number") that can be applied to an operation. The line number attribute can be a non-negative integer corresponding to a given line number in the original source code. The line number attribute can be optional and / or user-configurable at, for example, a module, package, class, or method level.

[0157] The enhanced version of the Java compiler can be configured to generate symbolic descriptions representing Java programs, which correspond to the bodies of Java methods and / or lambda expressions. These symbolic descriptions are valid, well-typed Java programs and can contain operations from both dialects.

[0158] Operations in the high-level dialect have properties that can be transformed (or lowered) to one or more operations in the core dialect. Lowering preserves the semantics of the program and can allow for easier analysis of control flow and data flow. However, lowering erases structure that is difficult to accurately recover. Accordingly, lowering a high-level dialect to a core dialect can be optional and / or user-configurable.

[0159] In an embodiment, an SDL representation of a Java program can be compiled into bytecode. To compile SDL into bytecode, one or more embodiments transform the SDL so that all high-level operations are reduced to core operations. The resulting description contains only core operations and can therefore be more easily compiled into bytecode.

[0160] As the Java programming language evolves and new language features are added, the corresponding SDL representation can be modeled. The modeling of new language features can include new core operations, new high-level operations and / or existing operations (core operations and / or high-level operations).

[0161] In one example, a library using symbolic descriptions is compiled on version 1 of the Java platform. An application compiled on version 2 or later of the Java platform uses the library and provides the symbolic description to the library. Version 2 of the Java platform introduces new Java language features that are modeled as high-level operations, and the application uses this language feature in a body expressed in a symbolic description. The library does not understand the new high-level operations. However, it can still reduce the operations to the core operations that it does understand. This ability helps ensure a certain degree of forward compatibility. However, as with the addition of new bytecode instructions, libraries compiled on earlier versions of the Java platform may fail when encountering the new operations.

[0162] 4.1 Java Core Dialect

[0163] In an embodiment, the Java core dialect includes the operations listed in Table 1. In addition, one or more embodiments include a set of arithmetic operations (binary, unary, test) on primitive values. These operations are not listed in Table 1 because their names are self-explanatory; for example, the operation "cos" returns the trigonometric cosine of an angle. For the sake of brevity, some arithmetic operations may be modeled after java.lang.Math methods rather than as method calls. Some of the operations included in the core dialect are discussed in more detail below.

[0164] Table 1: Operations in the Java core dialect

[0165]

[0166]

[0167] 4.1.1 Modeling Java Methods

[0168] One or more embodiments use SDL to model static methods, instance methods, and method signatures. In this example, the func operation definition is used to model a Java method. The func operation that describes a Java method in symbolic form includes:

[0169] A symbolic name attribute whose value is m.

[0170] An optional method descriptor attribute, for example named "source", which describes the signature of method m.

[0171] As the result of a void operation.

[0172] The func operation has a standalone body. Java methods, like functions, cannot capture values. Therefore, any nested operations are not allowed to reference values ​​defined outside the body. If the Java programming language is modified in the future to support capturing values, the func operation can be adjusted accordingly, and / or new operations can be defined to support the expanded functionality.

[0173] The body of the func operation contains blocks and operations that describe the method's body code. The body's entry block contains N block parameters—one for each parameter of m, in order. For a given parameter p of type t, the block parameter's name is p and contains a type descriptor describing t. The body's method type contains a return type that describes the return type of m. The terminating return operation exits the function and passes control back to the callee.

[0174] 4.1.2 Static Methods

[0175] Here is an example to describe the modeling of static methods, where a static method m is declared in class Foo:

[0176]

[0177] The symbol m in text form is:

[0178] func@"m"@source="Foo::m(int,int)int"(%x:int,%y:int)int->{

[0179] …

[0180] }

[0181] Note that in this example, the textual form merges the body and entry blocks. Furthermore, no class modeling is required. To model the parameters of method m as local variables rather than pure single static assignment (SSA) form, each block parameter type can be Var <t>, where T is the corresponding method parameter type:

[0182] func@"m"@source="Foo::m(int,int)int"(%x:Var <int>,%y:Var <int>)int->{ ...

[0184] }

[0185] Alternatively or additionally, local variables can be modeled inline:

[0186]

[0187] 4.1.3 Instance Methods

[0188] One or more embodiments model instance methods by adding an additional block parameter before all other arguments. This argument corresponds to "this," whose type describes the type of the method being declared. Specifically, for example, class Foo is defined as:

[0189]

[0190] The corresponding SDL representation of instance method m is:

[0191] func@"m"@source="Foo::m(int,int)int"(%this:Foo,%x:int,%y:int)int->{ ...

[0193] }

[0194] Note that the source method descriptor contains one less argument than the entry block parameter. Since it cannot be assigned a value, it is not modeled as a local variable. Furthermore, since the source method descriptor can be parsed to determine the method modifiers, there is no need to model them directly.

[0195] 4.1.4 Generic Methods

[0196] The type parameters declared by a generic method do not need to be modeled directly. However, the model type variables declared in the method's parameter types do need to be modeled. So, for example, if class Foo is defined as:

[0197]

[0198] The corresponding SDL representation of instance method m is:

[0199]

[0200] The method type descriptor for the body has the parameter types and return type of the type variables T.

[0201] One or more embodiments parse the source method descriptor as an instance of java.lang.reflect.Method and query the type parameters to determine that the type parameters introduce a type variable T. If Foo's class declaration is also generic, then this type variable will shadow any type variable of the same name introduced by Foo's class declaration.

[0202] 4.1.5 Annotation Methods

[0203] In an embodiment, there is no need to model annotations declared on methods or parameters of methods. Instead, one or more embodiments obtain such declarations by parsing the source method descriptor.

[0204] 4.1.6 Exceptions

[0205] In an embodiment, there is no need to model a method "throws" clause. Instead, one or more embodiments parse the source method descriptor to obtain the exceptions declared to be thrown.

[0206] 4.1.7 Modeling Field Access

[0207] In an embodiment, the field.load and field.store operations model field access expressions used to read values ​​from and assign values ​​to fields.

[0208] The field.load operation, which symbolically describes field access to the value of a field, consists of:

[0209] Zero or one operand that is the receiver of the field (optional).

[0210] A field descriptor attribute describing the method to be called. Assuming the referenced class exists at parse time, the field descriptor can be resolved to an instance of java.lang.reflect.Field.

[0211] A result type that is compatible with the field type of the field descriptor.

[0212] The field.store operation, which symbolically describes field access to assign a value to a field, consists of:

[0213] One or two operands that are the receiver of the field (optional) and the value to be assigned to the field.

[0214] A field descriptor attribute describing the field to be called. Assuming the referenced class exists at parse time, the field descriptor can be resolved to an instance of java.lang.reflect.Field.

[0215] A result type of void.

[0216] If the operand of a field.load operation is 1, or the operand of a field.store operation is 2, then the field access is to an instance field. Otherwise, the field access is to a static field.

[0217] 4.1.8 Modeling Method Calls

[0218] In an embodiment, a call operation definition models a call expression. A call operation that describes a method call in symbolic form includes:

[0219] • Zero or more operands corresponding to (a) the optional receiver of the method and (b) zero or more arguments to the method.

[0220] A method descriptor attribute describing the method to be called. Assuming the referenced class exists at resolution time, the method descriptor can be resolved to an instance of java.lang.reflect.Method.

[0221] A result type that is compatible with the return type of the method descriptor.

[0222] If the number of operands of a call operation is one more than the number of parameters in the method descriptor, then: the call is to an instance method; the first operand is the receiver; and subsequent operands are arguments. Otherwise, the call is to a static method.

[0223] An example of a call operation is described in more detail below.

[0224] 4.1.9 Reflection Operations

[0225] In an embodiment, Java language constructs that interact with types, classes, and objects at runtime (e.g., instantiating new objects or invoking methods) are modeled as reflective operations whose behavior is specified by Java reflection.

[0226] A reflective operation declares a descriptor, type, method type, method, or field descriptor that describes reflective information. The descriptor can be unambiguously resolved to an instance of the reflection class in the java.lang, java.lang.reflect, and java.lang.invoke packages, with appropriate access rights as needed.

[0227] Descriptors can be translated into equivalent bytecode descriptors, which can be encoded in the constant pool of the class file. This scheme helps to interpret the reflection operation or translate it into equivalent bytecode instructions (for example, a method call can be translated into an invokevirtual instruction). For example, in an embodiment, the reflection operation modeling the method call includes a method descriptor, which can be parsed into an instance of java.lang.reflect.Method or java.lang.invoke.MethodHandle.

[0228] The set of reflection operations is:

[0229] new, used to instantiate objects and array objects, accepts a method type descriptor.

[0230] call, used to call static or instance methods, accepts a method descriptor.

[0231] field.load and field.store, for accessing static or instance fields, accept field descriptors.

[0232] array.load and array.store, for accessing arrays, accept type descriptors.

[0233] method-ref, used to target a method as a functional interface, accepts a method descriptor.

[0234] cast, used to convert an object into another type, accepts a type descriptor.

[0235] instanceof, used to determine whether an object is an instance of a type, accepts a type descriptor.

[0236] 4.1.10 Parsing and Access Control

[0237] In an embodiment, method, field, and method type descriptors are parsed as follows:

[0238] A method descriptor is resolved to an instance of java.lang.reflect.Method or java.lang.invoke.MethodHandle.

[0239] Field descriptors are resolved to instances of java.lang.reflect.Field or java.lang.invoke.MethodHandle.

[0240] Method type descriptors are resolved to instances of java.lang.invoke.MethodType. Type descriptors are resolved to instances of java.lang.Class.

[0241] A MethodHandles.Lookup instance can be granted the ability to resolve a method, field, method handle, or class from an operation's descriptor. For type descriptors, resolution can be done using MethodHandles.Lookup.findClass.

[0242] 4.1.11 Static and instance methods and fields

[0243] In an embodiment, method and field descriptors do not inherently distinguish between static and instance methods. This distinction can be determined by using each descriptor with reflection operations. For method-ref, this can be determined based on the function interface and its single abstract method.

[0244] In an embodiment, the method descriptor of an invoke operation includes additional information indicating whether it is translated into bytecode as an invokespecial operation.

[0245] 4.1.12 Coercion and Conversion

[0246] In an embodiment, the operand(s) and results of at least some reflective operations are specified to be cast or converted—specifically, by parsing the descriptor to find the MethodHandle, adapting it to the operand and result types using MethodHandle.asType, and then invoking it using MethodHandle.invokeWithArguments. This approach reduces the need for explicit casts or conversions in symbolic descriptions.

[0247] 4.1.13 Preventing Heap Pollution

[0248] In an embodiment, to prevent calling a method on an object whose result is an instance of a type variable in the source code (e.g., <string>::get) when performing reflection. The descriptor retains generic type information at the use site. Generic type information can be used to adapt the return type of the resolved method handle before adapting it to the operation result.

[0249] For method descriptor java.util.List <string>::get(inti)Object, the receiver type is generic and has List <string>Furthermore, parsing of the descriptor shows that the method has a declaring class whose type parameter is the type variable E, and that the method has a generic return type that is the type variable E. Therefore, the return value of the method is an instance of String, and the return type of the method handle needs to be adapted to String.

[0250] The same applies to generic methods such as:

[0251]

[0252] In this example, the method descriptor will be Foo. <integer>::get(Number v)Number. Similar to the previous example, the descriptor and its parsed result provide enough information to determine that the return type of the parsed method handle needs to be adjusted to Integer.

[0253] 4.1.14 Translation into bytecode

[0254] In an embodiment, the translation into bytecode parsing descriptors makes the necessary information available for bytecode generation. Similarly, in a Java source compiler, the class file must be present on the module path or class path. For example, the method descriptor of a method call can be parsed into an instance of java.lang.reflect.Method, from which the method's access modifiers can be queried. The access modifiers can then determine whether to generate an invokevirtual or invokespecial bytecode instruction.

[0255] 4.1.15 Modeling Lambda Expressions

[0256] In an embodiment, a lambda operation definition models a lambda expression. A lambda operation, which symbolically describes a lambda, includes an operation result whose type describes a functional interface that is a target type of the lambda expression.

[0257] A lambda operation contains a non-isolated body, and therefore any nested operations can capture values ​​defined outside the body. The body contains blocks and operations that describe the code of the lambda body. The entry block of the body contains N block parameters, each of which corresponds in order to an abstract method parameter of the functional interface. For a given parameter p of type t, the block parameter is named p and has a type (or supertype) that describes t. The method type descriptor of the body contains a return type that describes the abstract method return type (or subtype) of the functional interface. The terminating return operation exits the lambda expression and passes control back to the callee.

[0258] The following is an example of modeling a lambda expression whose target type is IntUnaryOperator. The lambda expression captures the arguments of a method f:

[0259]

[0260]

[0261] 4.1.16 Reference Operation

[0262] In an embodiment, a reference operation defines a reference operation. For example, a lambda operation can be referenced, thereby encapsulating the operation and its contents so that it can be presented symbolically in runtime form rather than being processed symbolically as code. The result of referencing a lambda operation can be passed as an argument to a method call, allowing symbolic analysis and transformation of lambda expressions in a broader context at runtime (e.g., creating a symbolic description that models a SQL query).

[0263]

[0264]

[0265] The quote operation encapsulates the lambda expression and produces Quoted <lambdaop>An instance of , from which the runtime form of the lambda symbol description can be obtained. In addition, any captured arguments can be obtained from this instance.

[0266] 4.1.17 Modeling Local Variables

[0267] In an embodiment, local variables may be modeled in SSA form by defining three operations:

[0268] 1. Local variable definition operation, which accepts an initial value, a variable type, and an optional name. The result of this operation is Var <x>A variable value of type X, where X is the variable type. A variable value represents a box that holds the variable's value. A variable value cannot be accessed from ordinary Java code or concurrently by multiple threads; it behaves as if it were confined to the stack (just like a Java local variable).

[0269] 2. Read variable operation, which accepts Var <x>Type variable value, and returns the value of the variable of type X.

[0270] 3. Write a variable operation that accepts a value v of type X and Var <x>Updates the value of a variable of type v.

[0271] Since the variable value is in SSA form, its usage can be inferred through the level of indirection. This can be called impure SSA.

[0272] In many cases, the definition and use of a local variable can be replaced by the value it holds; therefore, these are core operations (see, for example, the try operation). These operations serve as a useful modeling tool for capturing the location of local variable definitions in the source code (including the capture of names). In addition, these operations simplify the design of advanced operations, as discussed in further detail below.

[0273] 4.2 Advanced Java Dialects

[0274] In an embodiment, the Java high-level dialect includes the operations listed in Table 2. Although not shown in Table 2, one or more embodiments may also model switch statements and expressions using the modeling techniques described herein. The operation of the high-level dialect can be greatly simplified if the use of local variables within the body of the high-level dialect is explicitly modeled in a non-pure SSA form. For example, this approach can be used for local variables that are written to, because final (or effectively final) local variables can be modeled directly as values.

[0275] Java dialect operations do not need to return multiple values ​​to update all associated local variables. One or more embodiments model the return value of an expression (such as a switch expression). This simplification is obvious for nested code (e.g., nested loops where the inner loop updates a variable), where using pure SSA would require propagating the value up the nest.

[0276] In embodiments, operations lowered into the core dialect may result in local variables being omitted, assuming they do not need to be preserved (e.g., see the discussion herein regarding modeling try statements). Some operations are described in further detail below.

[0277]

[0278]

[0279] 4.2.1 Modeling the loop

[0280] One or more embodiments define operations for modeling loops. A graph of basic blocks can model loops and other forms of control flow. However, at this level, the structure of the code is erased. One or more embodiments include specific operations that preserve this structure.

[0281] 4.2.2 Modeling the Enhanced for Loop

[0282] In one embodiment, the enhancedFor operation definition models an enhanced for statement. The enhancedFor operation, which symbolically describes the enhanced for statement, includes an operand whose type is a subtype of java.lang.Iterable or an array type. The enhancedFor operation includes a body that models the statements contained in the loop. The body's entry block includes arguments, which are elements of the Iterable or an array for the current iteration of the loop.

[0283] The following is an example of modeling an enhanced for loop, where the body is an expression that returns an Iterable, the body of the element variable definition, and the loop body:

[0284]

[0285]

[0286] In this example, even though the elements of the iterable (list) are of type Integer, the body's input parameter is of type int. The same unboxing conversion is performed implicitly, following similar rules as for conversions of arguments and return values ​​of reflective operations.

[0287] One or more embodiments reduce the aforementioned symbolic description to a description containing only the core operations. In addition, one or more embodiments remove local variable operations:

[0288]

[0289] ^entryBlock(%5)

[0290] ^entryBlock(%i:int):

[0291] %nextSum:int=add%sum%i

[0292] br^header(%nextSum)

[0293] ^entryBlock_split:

[0294] return %sum

[0295] }

[0296] In an embodiment, lowering requires performing a method call on the iterable value to obtain an iterator, and then performing a method call on the iterator to check whether there are any elements, and if so, to obtain the next element. In this example, some method descriptors for the call operations describe methods with erased types, but the type of the value is not erased because one or more embodiments rely on implicit conversions as specified by the reflection operation. The call to obtain the next element ensures that the return value is an instance of Integer. No explicit casts need to be inserted, as would be the case if the original source were compiled into bytecode (or the description were converted into bytecode).

[0297] With this approach, it becomes more difficult to determine the original structure of the code. However, it is not necessary to understand the specific semantics of the enhancedFor operation.

[0298] 4.2.3 Modeling a Counting For Loop

[0299] In an embodiment, a countedFor operation definition models a for statement as a counted loop. The countedFor operation, which symbolically describes a for statement, includes three operands, all of the same integer type, corresponding to (1) a count start value (inclusive), (2) a count end value (inclusive), and (3) a count step size. The countedFor operation includes a body that models the statements contained in the loop. The entry block of the body includes an argument, which is the current count. Counted loops may be easier to identify than for statements, but may be more difficult to identify when represented differently (such as in a degraded form). Counted loops may be useful to identify for optimization and transformation purposes.

[0300] The following is an example of modeling a for statement as a counting loop, summing the count:

[0301] private static int f(int start,int end,int step){

[0302]

[0303] 4.2.4 Modeling the WHILE Loop

[0304] In an embodiment, the while operation definition models a while statement. The while operation, which symbolically describes a while statement, does not contain any operands and contains two bodies. The first body (the predicate body) models the expression of the while statement, and the second body (the loop action body) models the contained statements. The predicate body produces a Boolean value. The following is an example of modeling a while statement:

[0305]

[0306]

[0307]

[0308] 4.2.5 Modeling IF-THEN and IF-THEN-ELSE Statements

[0309] In an embodiment, the ifelseif operation definition models if-then and if-then-else statements. The ifelseif operation, which symbolically describes an if-then or if-then-else statement, contains zero operands and two or more bodies. The sequence of bodies contains a predicate-body pair and an action-body that models the if-then. Optionally, at the end of the sequence of bodies, an action-body models the else. The predicate body models the if expression, and the action body models the contained if or else statement. The following is an example of modeling an if-then-else statement:

[0310]

[0311]

[0312]

[0313] One or more embodiments reduce the symbolic description to a description containing only the core operations. In addition, one or more embodiments remove local variable operations:

[0314]

[0315]

[0316] Note that in this example, the ifelseif operation is more expressive than the corresponding Java construct it models, because the predicate body is not restricted to modeling only Java expressions.

[0317] 4.2.6 Modeling TRY / CATCH / FINALLY

[0318] In an embodiment, the try operation definition models the try / catch / finally statement. The try operation, which symbolically describes a try statement, contains up to three bodies: the body of the try code statement; the catch clause and optional body of the statement; and the optional body of the finally statement. Multiple catch regions constructed in the Java language are merged into a single body, and an instanceof check is performed for each exception class. The entry block of the catch body contains an argument of type Throwable and thus distinguishes itself from the finally body, which does not contain any arguments.

[0319] The try operation specifies how control passes from the try body to the catch body, and then to the finally body. If an exception occurs in the try body, control passes to the catch body, which passes the exception as a value to the catch body's entry block. If a finally body is present, then (a) control passes to the finally body before processing the termination action exiting the try or catch body; and (b) if the finally body is exited via break, control passes back to the try or catch block to process the termination action.

[0320] As an example:

[0321]

[0322]

[0323] The above method can be modeled using the following notation:

[0324]

[0325]

[0326] The second body of the try operation, the catch body, models the catch clauses and statements that catch and handle ArrayIndexOutOfBoundsException and NullPointerException exceptions. The body tests whether the Throwable value is an instance of ArrayIndexOutOfBoundsException or NullPointerException; otherwise, it rethrows the value. In this example, the body contains multiple (basic) blocks, some of which are reached by conditional branching on the result of the tryof operation. Alternatively, the model can declare multiple catch bodies in sequence, one for each exception type, thereby modeling more closely the source structure. Modeling multiple catch clauses will still follow the same scheme as described above.

[0327] One or more embodiments use the var operation to model local variables that are explicitly updated by a try statement. The try body performs a var.store operation, which should never throw any exceptions (but may still throw an Error for exceptional circumstances).

[0328] The var operation and related operations can be omitted, and the try operation updated to yield the value currently stored in the local variable, as follows:

[0329]

[0330] If, in a try body, a local variable is stored into and loaded from a catch or finally body, one or more embodiments may not omit the var.load operation because the value that should replace the result of the operation is unknown.

[0331] 5. Example operation of modeling Java source code with SDL

[0332] Figure 6 Illustrated is a set of example operations for modeling Java source code in a symbolic description language in accordance with one or more embodiments. Figure 6 One or more operations shown in may be modified, rearranged, or omitted altogether. Accordingly, Figure 6 The particular order of operations shown should not be construed as limiting the scope of one or more embodiments.

[0333] In the following discussion, the term "system" refers to any system or its component(s) that is configured to generate an SDL model of Java source code. For example, a system may refer to a compiler and / or a standalone modeling tool. A system may also include a runtime environment (e.g., a JRE) that is configured to execute bytecode corresponding to the Java source code.

[0334] In an embodiment, the system obtains Java source code (operation 602).For example, the system can obtain code as input to a compiler, command line arguments (e.g., referencing one or more names of files containing Java source code), etc.

[0335] The system determines the type(s) represented in the Java source code (operation 604). For example, the system can parse the code to identify type declarations and / or calls to native Java types. The system further determines the semantic structure of the Java source code (operation 606), including whether there are any conditional branches, loops, lambda expressions, etc. The system generates an SDL model of the Java source code (operation 608), such that the SDL model describes the identified type(s) and semantic structure. Some examples of SDL models for various Java language constructs are described in detail above.

[0336] One or more embodiments are configured to use an SDL model at runtime. The system may execute bytecode corresponding to Java source code (operation 610). During runtime, the system may encounter a request to reflect on types defined in the Java source code (operation 612). In response to encountering the request, the system may reflect on the SDL model corresponding to the Java source code (operation 614). Because the SDL model describes language constructs that would otherwise be lost during compilation to bytecode, reflecting on the SDL model allows for a wider range of reflection operations.

[0337] As described above, the system can first use standard reflection to obtain an SDL representation. For example, to access the SDL representation of a method body (if one exists), the system can first obtain a java.lang.reflect.Method instance. The system can then query the reflection object to obtain its SDL representation. The following is a code example for obtaining an SDL representation of a method body according to one or more embodiments:

[0338]

[0339]

[0340] In the above example, according to one or more embodiments, m.getTree() is a new method added to Method that returns an Optional<CoreOps.FuncOp> Furthermore, in this example, the system uses annotations to identify methods that have corresponding SDL representations.

[0341] Alternatively or additionally, one or more embodiments "target type" a lambda expression to an instance of a particular type (e.g., Quoted). This instance contains the runtime SDL representation of the lambda expression as a Closure operation (which is similar to a lambda operation, but lacks a function interface). For example:

[0342]

[0343]

[0344] 6. Example Embodiments

[0345] For the sake of clarity, a detailed example is described below. The components and / or operations described below should be understood as a specific example and may not be applicable to certain embodiments. Accordingly, the components and / or operations described below should not be understood as limiting the scope of any claims.

[0346] Specifically, Figures 7A-7B The figure shows an example of modeling Java source code using a symbolic description language according to one or more embodiments. Figure 7A As shown in FIG, a Java compiler 704 receives a Java source code 702 and generates (a) a Java bytecode 706 corresponding to the Java source code 702 and (b) an SDL model 708 corresponding to the Java source code 702. The Java bytecode 706 can be a component of the SDL model 708, such that the SDL model 708 contains information required to execute the Java bytecode 706 and to perform reflective operations that are not possible with the Java bytecode 706 alone. For example, the SDL model 708 can model class file code attributes, which contain Java bytecode 706 instruction sequences.

[0347] Continuing with the example, Figure 7B As shown in FIG, execution platform 710 includes a virtual machine 712 configured to execute Java bytecodes. Virtual machine 712 receives Java bytecodes 706 and SDL model 708 (or only SDL model 708 if it contains Java bytecodes 706). Virtual machine 712 executes Java bytecodes 706 and uses SDL model 708 to reflect on one or more types defined in Java source code 702 at runtime.

[0348] 7. Sample Application

[0349] The SDL model of Java source code generated as described herein can be used in a variety of applications, including but not limited to the following examples. In general, one or more embodiments support transforming Java source code into some other form, such as source code written in another language, or Java source code that differs from the original source code in some respect (e.g., via differentiation and / or optimization).

[0350] SDL models can be used in Java-based machine learning applications where understanding types and semantic structure is important. For example, gradient descent techniques start from the original Java source code and generate differentiated versions of the Java source code. The differentiation process can use SDL models so that the differentiated Java source code is based on a complete description of the types and semantic structure defined therein.

[0351] The SDL model can be used to generate code written in another programming language other than Java. Because the SDL model accurately represents the types and semantic structures defined in the Java source code, the corresponding code in the other language can be functionally equivalent to the Java version. The system can then (if necessary) compile the new code into an executable form (e.g., an executable file, non-Java bytecode, etc.). The other programming language can be a domain-specific language, that is, a language designed for a specific operating domain. ParallelGraph AnalytiX (PGX) is an example of a domain-specific language designed for graph analysis.

[0352] The SDL model can be used to optimize Java programs. The process of generating an SDL model based on Java source code can eliminate redundant and / or inefficient semantic structures. Alternatively or additionally, the SDL model can clarify multi-threading opportunities. The SDL model can be compiled into Java bytecode (optionally, first by generating transformed Java source code based on the optimized SDL model), which operates more efficiently than bytecode compiled from the original source code. For example, the technology described herein can be integrated into an accelerated VM such as TornadoVM.

[0353] 8. Additional Examples

[0354] Appendix A of this specification, which is incorporated by reference in its entirety, describes additional examples according to one or more embodiments.

[0355] The class definitions provided in Appendix A contain the tests and SDL representation output for the examples described in this article. Each example is presented in text form as the SDL representation of the code in the following Java methods, in order:

[0356] (1) Contains high-level representations of operations in both high-level and core dialects.

[0357] (2) Transform (1) into a representation that contains only operations in the core dialect.

[0358] (3) Transform (2) into pure SSA.

[0359] (4) Transform (3) into operations in a bytecode dialect from which the system can generate bytecode.

[0360] In these examples, the programming meaning is preserved.

[0361] The examples in Appendix A model each catch as a separate body. Alternatively, one or more embodiments may combine multiple catches into a single catch body. One or more embodiments may swap existing clauses in an application and focus first on the separate catch bodies. For multiple catches in an embodiment that does not model the union type "IndexOutOfBounds|IOException" (e.g., catch(IndexOutOfBounds|IOException e)), one or more embodiments may perform intnanceof checks and casts. In these examples, the bodies have names for ease of identification.

[0362] In one embodiment, the downgraded try operation requires inlining code in the finally body, just before the exit point of the try and catch bodies. One or more embodiments further identify regions in the code where exceptions may be thrown and caught. The exception.region.enter and exception.region.exit operations support this approach.

[0363] One or more embodiments model the enhanced for statement as a java.enhancedFor operation, which contains three bodies:

[0364] The first body corresponds to an expression whose result is an instance of Iterable. This body produces the value of the iterable.

[0365] The second body accepts an element from the iterable and produces a variable for that element.

[0366] The third body accepts a variable for the element. The body terminates with a java.continue call, which passes control back to the operation. The operation then continues to obtain the next element from the iterable. Otherwise, if a java.break call is present, control passes back to the operation, which then passes control back to the parent block.

[0367] One or more embodiments model the for statement as a java.for operation, which contains four bodies:

[0368] • The first body corresponds to the init statement that initializes the loop variables. The loop variables are generated from the body (if there is more than one loop variable, one or more embodiments return a tuple holding two or more variables).

[0369] The second body corresponds to the conditional expression, which accepts the loop variable and produces a Boolean value.

[0370] The third body corresponds to an update expression statement, which takes the loop variables and modifies one or more of them.

[0371] The fourth body corresponds to the body statement, which accepts the loop variable and can choose to continue loop iterations or exit the loop until the conditional expression returns false.

[0372] One or more embodiments model a while statement as a java.while operation, which contains two bodies:

[0373] The first body corresponds to the conditional expression, which produces a Boolean value.

[0374] The second body corresponds to the body statement, which can choose to continue loop iterations or exit the loop until the conditional expression returns false.

[0375] In the example in Appendix A, the conditional body of the loop operation contains the java.cand operation that models the binary conditional && expression, and the loop body of the operation contains the java.if operation that models the if statement.

[0376] One or more embodiments model an if statement as a java.if operation, which contains 2N+1 bodies, where N corresponds to the number of conditional (or corresponding then) expressions. The bodies are arranged in order, corresponding to the conditional expression that produces a Boolean value, the then statement that produces a void value, and so on, and finally ends with the body corresponding to the last else statement.

[0377] 9. Machine Learning

[0378] In one or more embodiments, a machine learning algorithm is an iterative algorithm that uses a set of training data to learn a target model that optimally maps a set of input variables to one or more output variables. The training data includes a dataset and associated labels. The dataset is associated with the input variables of the target model. The associated labels are associated with the output variable(s) of the target model. The training data can be updated based on, for example, feedback on the accuracy of the current target model. The updated training data can be fed back into the machine learning algorithm, which can in turn update the target model.

[0379] The machine learning algorithm can generate a target model such that the target model best fits the dataset of training data to the labels of the training data. Specifically, the machine learning algorithm can generate the target model such that, when the target model is applied to the dataset of training data, the maximum number of results determined by the target model matches the labels of the training data. Different target models can be generated based on different machine learning algorithms and / or different sets of training data.

[0380] Machine learning algorithms can include supervised and / or unsupervised components. Various types of algorithms can be used, such as linear regression, logistic regression, linear discriminant analysis, classification and regression trees, naive Bayes, k-nearest neighbors, learning vector quantization, support vector machines, bagging and random forests, boosting, back propagation and / or clustering.

[0381] 10. Computer Networks and Cloud Networks

[0382] In one or more embodiments, a computer network provides connections between a collection of nodes. The nodes can be local and / or remote from each other. The nodes are connected via a set of links. Examples of links include coaxial cables, unshielded twisted pair cables, copper cables, optical fibers, and virtual links.

[0383] A subset of nodes implements a computer network. Examples of such nodes include switches, routers, firewalls, and network address translators (NATs). Another subset of nodes utilizes a computer network. Such nodes (also referred to as "hosts") can execute client processes and / or server processes. A client process makes a request for a computing service (such as the execution of a specific application and / or the storage of a specific amount of data). The server process responds by, for example, executing the requested service and / or returning the corresponding data.

[0384] A computer network can be a physical network, comprising physical nodes connected by physical links. A physical node is any digital device. A physical node can be a function-specific hardware device, such as a hardware switch, a hardware router, a hardware firewall, or a hardware NAT. Additionally or alternatively, a physical node can be a general-purpose machine configured to execute various virtual machines and / or applications that perform corresponding functions. A physical link is the physical medium that connects two or more physical nodes. Examples of links include coaxial cable, unshielded twisted cable, copper cable, and optical fiber.

[0385] A computer network can be an overlay network. An overlay network is a logical network implemented on top of another network (such as a physical network). Each node in the overlay network corresponds to a corresponding node in the underlying network. Therefore, each node in the overlay network is associated with both an overlay address (addressing the overlay node) and an underlying address (addressing the underlying node that implements the overlay node). Overlay nodes can be digital devices and / or software processes (such as virtual machines, application instances, or threads). The links connecting overlay nodes are implemented as tunnels through the underlying network. The overlay nodes at either end of the tunnel treat the underlying multi-hop path between them as a single logical link. Tunneling is performed by encapsulation and decapsulation.

[0386] The client can be located locally on the computer network and / or remotely from the computer network. The client can access the computer network through other computer networks, such as a private network or the Internet. The client can transmit requests to the computer network using a communication protocol, such as the Hypertext Transfer Protocol (HTTP). The request is transmitted through an interface, such as a client interface (e.g., a web browser), a program interface, or an application programming interface (API).

[0387] In one or more embodiments, a computer network provides connections between clients and network resources. Network resources include hardware and / or software configured to execute server processes. Examples of network resources include processors, data storage devices, virtual machines, containers, and / or software applications. Network resources are shared among multiple clients. Clients independently request computing services from the computer network. Network resources are dynamically allocated to requests and / or clients on demand. The network resources allocated to each request and / or client can be expanded or reduced based on, for example, (a) the computing services requested by a specific client, (b) the aggregated computing services requested by a specific tenant, and / or (c) the requested aggregated computing services of the computer network. Such a computer network may be referred to as a "cloud network."

[0388] In one or more embodiments, a service provider provides a cloud network to one or more end users. The cloud network can implement various service models, including but not limited to Software as a Service (SaaS), Platform as a Service (PaaS), and Infrastructure as a Service (IaaS). In SaaS, the service provider provides the end user with the ability to use the service provider's applications that are being executed on network resources. In PaaS, the service provider provides the end user with the ability to deploy customized applications on network resources. Custom applications can be created using programming languages, libraries, services, and tools supported by the service provider. In IaaS, the service provider provides the end user with the ability to supply processing, storage, network, and other basic computing resources provided by the network resources. Any arbitrary application, including an operating system, can be deployed on the network resources.

[0389] Computer networks can be deployed in various ways, including but not limited to private clouds, public clouds, and / or hybrid clouds. In a private cloud, network resources are provisioned for exclusive use by a specific group of one or more entities (as used herein, the term "entity" refers to an enterprise, organization, individual, or other entity). Network resources can be local to and / or remote from the premises of the specific entity group. In a public cloud, cloud resources are provisioned to multiple independent entities (also referred to as "tenants" or "clients"). A computer network and its network resources can be accessed by clients corresponding to different tenants. Such a computer network may be referred to as a "multi-tenant computer network." Several tenants can use the same specific network resources at different times and / or at the same time. Network resources can be local to and / or remote from the tenant's premises. In a hybrid cloud, the computer network includes both private and public clouds. The interface between the private and public clouds allows for data and application portability. Data stored in the private cloud and data stored in the public cloud can be exchanged via the interface. Applications implemented in the private cloud and applications implemented in the public cloud may have dependencies on each other. Calls from an application at the private cloud to an application at the public cloud (and vice versa) may be performed through the interface.

[0390] In one or more embodiments, the tenants of a multi-tenant computer network are independent of each other. For example, the business or operations of one tenant can be separated from the business or operations of another tenant. Different tenants may have different network requirements for the computer network. Examples of network requirements include processing speed, data storage capacity, security requirements, performance requirements, throughput requirements, latency requirements, resilience requirements, quality of service (QoS) requirements, tenant isolation and / or consistency. The same computer network may need to implement different network requirements required by different tenants.

[0391] In a multi-tenant computer network, tenant isolation can be implemented to ensure that applications and / or data of different tenants are not shared with each other. Various tenant isolation schemes can be used. Each tenant can be associated with a tenant identifier (ID). Each network resource of the multi-tenant computer network can be tagged with a tenant ID. A tenant can only be allowed to access a specific network resource if the tenant and the specific network resource are associated with the same tenant ID.

[0392] For example, each application implemented by a computer network can be tagged with a tenant ID, and a tenant can only be allowed to access a specific application if the tenant and the specific application are associated with the same tenant ID. Each data structure and / or dataset stored by a computer network can be tagged with a tenant ID, and a tenant can only be allowed to access a specific data structure and / or dataset if the tenant and the specific data structure and / or dataset are associated with the same tenant ID. Each database implemented by a computer network can be tagged with a tenant ID, and a tenant can only be allowed to access data in a specific database if the tenant and the specific database are associated with the same tenant ID. Each entry in a database implemented by a multi-tenant computer network can be tagged with a tenant ID, and a tenant can only be allowed to access a specific entry if the tenant and the specific entry are associated with the same tenant ID. However, a database can be shared by multiple tenants.

[0393] In one or more embodiments, a subscription list indicates which tenants are authorized to access which network resources. For each network resource, a list of tenant IDs that are authorized to access the network resource may be stored. A tenant may be allowed to access a particular network resource only if the tenant ID of the tenant is included in the subscription list corresponding to the particular network resource.

[0394] In one or more embodiments, network resources corresponding to different tenants (such as digital devices, virtual machines, application instances, and threads) are isolated to tenant-specific overlay networks maintained by a multi-tenant computer network. As an example, packets from any source device in a tenant overlay network can only be sent to other devices within the same tenant overlay network. Encapsulation tunnels can be used to prohibit any transmission from a source device on a tenant overlay network to devices in other tenant overlay networks. Specifically, a packet received from a source device can be encapsulated within an outer packet. The outer packet is sent from a first encapsulation tunnel endpoint (communicating with a source device in the tenant overlay network) to a second encapsulation tunnel endpoint (communicating with a destination device in the tenant overlay network). The second encapsulation tunnel endpoint decapsulates the outer packet to obtain the original packet sent by the source device. The original packet is sent from the second encapsulation tunnel endpoint to a destination device in the same specific overlay network.

[0395] 11. Hardware Overview

[0396] In one or more embodiments, the technology described herein is implemented by one or more special-purpose computing devices. The special-purpose computing device(s) may be hard-wired to perform the technology, and / or may include digital electronic devices that are permanently programmed to perform the technology, such as one or more application-specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or network processing units (NPUs), or may include one or more general-purpose hardware processors that are programmed to perform the technology according to program instructions in firmware, memory, other storage devices, or a combination thereof. Such special-purpose computing devices may also combine customized hard-wired logic, ASICs, FPGAs, or NPUs with customized programming to implement the technology. The special-purpose computing device may be a desktop computer system, a portable computer system, a handheld device, a networked device, or any other device that combines hard-wiring and / or program logic to implement the technology.

[0397] For example, Figure 8 8 is a block diagram illustrating a computer system 800 upon which one or more embodiments of the present invention may be implemented. Computer system 800 includes a bus 802 or other communication mechanism for communicating information, and a hardware processor 804 coupled with bus 802 for processing information. Hardware processor 804 may be, for example, a general-purpose microprocessor.

[0398] Computer system 800 also includes a main memory 806, such as a random access memory (RAM) or other dynamic storage device, coupled to bus 802 for storing information and instructions to be executed by processor 804. Main memory 806 may also be used to store temporary variables or other intermediate information during execution of instructions to be executed by processor 804. When such instructions are stored in a non-transitory storage medium accessible to processor 804, such instructions cause computer system 800 to become a special-purpose machine customized to perform the operations specified in the instructions.

[0399] Computer system 800 also includes a read only memory (ROM) 808 or other static storage device coupled to bus 802 for storing static information and instructions for processor 804. A storage device 810, such as a magnetic or optical disk, is provided and coupled to bus 802 for storing information and instructions.

[0400] The computer system 800 may be coupled via bus 802 to a display 812, such as a cathode ray tube (CRT), for displaying information to a computer user. An input device 814, including alphanumeric and other keys, is coupled to bus 802 for communicating information and command selections to processor 804. Another type of user input device is a cursor control 816, such as a mouse, trackball, or cursor direction keys, for communicating direction information and command selections to processor 804 and for controlling cursor movement on display 812. Such input devices typically have two degrees of freedom along two axes, a first axis (e.g., x) and a second axis (e.g., y), to allow the device to specify a position in a plane.

[0401] Computer system 800 can implement the techniques described herein using custom hardwired logic, one or more ASICs or FPGAs, firmware, and / or program logic that, in combination with computer system 800, makes computer system 800 a special-purpose machine or programs computer system 800 to be a special-purpose machine. In one or more embodiments, the techniques herein are performed by computer system 800 in response to processor 804 executing one or more sequences of one or more instructions contained in main memory 806. These instructions can be read into main memory 806 from another storage medium, such as storage device 810. Execution of the sequences of instructions contained in main memory 806 causes processor 804 to perform the process steps described herein. Alternatively, hardwired circuitry can be used in place of or in combination with software instructions.

[0402] As used herein, the term "storage medium" refers to any non-transitory medium that stores data and / or instructions that cause a machine to operate in a specific manner. Such storage media may include non-volatile media and / or volatile media. Non-volatile media include, for example, optical or magnetic disks, such as storage device 810. Volatile media include dynamic memory, such as main memory 806. Common forms of storage media include, for example, floppy disks, flexible disks, hard disks, solid-state drives, magnetic tape or any other magnetic data storage medium, compact disk read-only memory (CD-ROM), any other optical data storage medium, any physical medium with a hole pattern, RAM, PROM and EPROM, FLASH-EPROM, NVRAM, any other memory chip or cassette tape, content addressable memory (CAM), and ternary content addressable memory (TCAM).

[0403] Storage media are distinct from, but may be used in conjunction with, transmission media. Transmission media participate in the transfer of information between storage media. For example, transmission media include coaxial cables, copper wire, and optical fiber, including the wires of bus 802. Transmission media may also take the form of acoustic or light waves, such as those generated during radio frequency (RF) and infrared data communications.

[0404] Various forms of media may be involved in carrying one or more sequences of one or more instructions to processor 804 for execution. For example, the instructions may initially be carried on a magnetic disk or solid-state drive of a remote computer. The remote computer can load the instructions into its dynamic memory and send the instructions over a telephone line or other communications medium using a modem. A modem local to computer system 800 can receive the data on the telephone line or other communications medium and use an infrared transmitter to convert the data into an infrared signal. An infrared detector can receive the data carried in the infrared signal, and appropriate circuitry can place the data on bus 802. Bus 802 carries the data to main memory 806, from which processor 804 retrieves and executes the instructions. The instructions received by main memory 806 may optionally be stored on storage device 810 before or after execution by processor 804.

[0405] Computer system 800 also includes a communication interface 818 that is coupled to bus 802. Communication interface 818 provides two-way data communication that is coupled to network link 820, and wherein network link 820 is connected to local network 822. For example, communication interface 818 can be an integrated services digital network (ISDN) card, a cable modem, a satellite modem, or a modem that provides data communication connections to the telephone line of the corresponding type. As another example, communication interface 818 can be a LAN card that is configured to provide data communication connections to a compatible local area network (LAN). Wireless links can also be implemented. In any such implementation, communication interface 818 sends and receives electrical signals, electromagnetic signals, or optical signals that carry the digital data stream representing various types of information.

[0406] The network link 820 typically provides data communication to other data devices through one or more networks. For example, the network link 820 can provide a connection to a host computer 824 or to data equipment operated by an Internet Service Provider (ISP) 826 through a local network 822. The ISP 826, in turn, provides data communication services through the global packet data communication network now commonly referred to as the "Internet" 828. Both the local network 822 and the Internet 828 use electrical, electromagnetic, or optical signals that carry digital data streams. The signals through the various networks and the signals on the network link 820 and through the communication interface 818 are example forms of transmission media that carry digital data to and from the computer system 800.

[0407] Computer system 800 can send messages and receive data, including program code, through the network(s), network link 820, and communication interface 818. In the Internet example, server 830 can transmit the requested code for an application program through Internet 828, ISP 826, local network 822, and communication interface 818.

[0408] The received code may be executed by processor 804 as it is received, and / or may be stored in storage device 810 or other non-volatile storage for later execution.

[0409] 12. Other Matters; Extension

[0410] Embodiments are directed to a system having one or more devices comprising a hardware processor and configured to perform any of the operations described herein and / or in any of the following claims.

[0411] In one or more embodiments, a non-transitory computer-readable storage medium stores instructions that, when executed by one or more hardware processors, cause performance of any of the operations described herein and / or claimed.

[0412] Any combination of the features and functions described herein may be used in accordance with one or more embodiments. In the foregoing specification, embodiments have been described with reference to numerous specific details that may vary from implementation to implementation. Accordingly, the specification and drawings are to be regarded as illustrative rather than restrictive. The sole and exclusive indicator of the scope of the invention, and the scope intended by the applicants to be the scope of the invention, is the literal and equivalent scope of the set of claims set forth in this application, in the specific form of that set of claims, including any subsequent amendments.

[0413] Appendix A

[0414]

[0415]

[0416]

[0417]

[0418]

[0419]

[0420]

[0421]

[0422]

[0423]

[0424]

[0425]

[0426]

[0427]

[0428]

[0429]

[0430]

[0431]

[0432]

[0433]

[0434]

[0435]

[0436]

[0437] < / x> < / x> < / x> < / lambdaop> < / integer> < / string> < / string> < / string> < / int> < / int> < / t>

Claims

1. One or more non-transitory computer-readable media storing instructions that, when executed by one or more processors, cause operations comprising: Get a set of Java source code; determining that the set of Java source codes includes a user-defined type; determining that the set of Java source codes includes a loop; A symbolic description language (SDL) model is generated based on the set of Java source codes, the symbolic description language model including a first SDL representation of the user-defined type and a second SDL representation of the loop.

2. The one or more non-transitory computer-readable media of claim 1 , wherein the SDL model represents the set of Java source code using a pattern, the pattern comprising: An operation, which consists of a name, zero or more operands, an operation result, zero or more attributes, and a body; The body comprises one or more blocks; Each of the one or more blocks contains one or more corresponding operations.

3. The one or more non-transitory computer-readable media of claim 1, wherein the loop is a for loop.

4. The one or more non-transitory computer-readable media of claim 1, wherein the loop is a while loop.

5. One or more non-transitory computer-readable media as recited in claim 1: wherein the set of Java source codes further comprises an if-then statement; The SDL model further includes a third SDL representation of the if-then statement.

6. One or more non-transitory computer-readable media as recited in claim 1: wherein the set of Java source codes further includes a try / catch / finally block; The SDL model further includes a third SDL representation of the try / catch / finally block.

7. One or more non-transitory computer-readable media as recited in claim 1: wherein the set of Java source codes further includes a lambda expression; The SDL model further includes a third SDL representation of the lambda expression.

8. A system comprising: at least one device comprising one or more hardware processors, The system is configured to perform the operations as claimed in any one of claims 1-7.

9. A method comprising the operations according to any one of claims 1 to 7.

10. A system comprising means for performing the operations of any one of claims 1-7.

Citation Information

Patent Citations

  • Transforming a JAVA program using a symbolic description language model

    US20240272884A1