A code optimization method, apparatus, computing device, and computer program product
Patent Information
- Application Number
- CN202510399896.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2026-09-29
AI Technical Summary
然而,在该SQL代码中存在java语言的用户自定义函数(user-defined function,UDF)的情况下,需要native执行引擎与JVM通过java本地接口(java native interface,JNI)频繁交互;从而导致代码运行效率的提升非常有限
[0045]第八方面,本申请提供一种计算设备集群,包括至少一个计算设备,每个计算设备包括处理器和存储器;该至少一个计算设备的处理器用于执行至少一个计算设备的存储器中存储的指令,以使得计算设备集群执行第一方面和第二方面及其可能的实现方式中任意之一所述的方法。
Smart Images

Figure CN122837837A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to a code optimization method, apparatus, computing device, and computer program product. Background Technology
[0002] As is well known, big data systems mainly run in the Java Virtual Machine (JVM); however, due to the low efficiency of the JVM in running code, the operating efficiency of big data systems is relatively low.
[0003] A common code optimization approach is to retrieve Structured Query Language (SQL) code from the JVM using a native execution engine and execute that SQL code. However, when the SQL code contains user-defined functions (UDFs) in Java, the native execution engine needs to frequently interact with the JVM through the Java Native Interface (JNI), resulting in very limited improvements in code execution efficiency. Summary of the Invention
[0004] This application provides a code optimization method, apparatus, computing device, and computer program product that can improve code execution efficiency.
[0005] To achieve the above objectives, the embodiments of this application adopt the following technical solutions:
[0006] In a first aspect, embodiments of this application provide a code optimization method, the method comprising: obtaining first code; the first code being binary code obtained by compiling original code, the original code being code based on a first programming language; converting the first code into second code, the second code being code based on a second programming language having the same function as the first code; wherein the code in the second programming language has higher running efficiency on the processor than the code in the first programming language; the second programming language is a programming language supported by a native execution engine, the native execution engine running on the aforementioned processor.
[0007] In this embodiment, the compiled code in a first programming language (i.e., the first code) is converted into second code in a second programming language with the same functionality. Since the code in the second programming language is more efficient on the processor than the code in the first programming language, and the second programming language is supported by the native execution engine, the native execution engine can directly run this more efficient second code without the intervention of the runtime platform (e.g., JVM) of the first programming language (e.g., Java). Therefore, there is no problem of frequent interaction between the runtime platform and the native execution engine, thus improving the code's execution efficiency.
[0008] Furthermore, since the embodiments of this application convert the first code into second code supported by the native execution engine, so that the native execution engine can implement the function of the first code by executing the second code; and so that the native execution engine can implement the function of the first code by executing the second code corresponding to the first code, thereby increasing the applicability of the native execution engine.
[0009] In one possible implementation, the above-mentioned conversion of the first code into the second code includes: converting the dependency object on which the first code depends into a dependency object with the same function written in a second programming language to obtain the third code; wherein the type of the dependency object includes: interface and / or class; and converting the third code into the second code based on the syntax rules of the second programming language.
[0010] In one possible implementation, the process of converting the dependency object on which the first code depends into a functionally identical dependency object written in a second programming language to obtain the third code includes: obtaining the identifier of the first dependency object on which the first code depends from the first code; determining the second dependency object corresponding to the first dependency object from a first correspondence based on the identifier of the first dependency object; the second dependency object being written in the second programming language; wherein the first dependency object and the second dependency object have the same function; the first correspondence is used to indicate the correspondence between dependency objects written in the first programming language and functionally identical dependency objects written in the second programming language; and replacing the first dependency object on which the first code depends with the second dependency object to obtain the third code.
[0011] Based on the principle of substitution, the embodiments of this application directly replace the first dependent object on which the first code depends with a second dependent object with the same function written in a second programming language, thereby realizing the replacement of the dependent object in the first code without needing to convert the syntax rules of the code in the first dependent object, thus improving the code conversion efficiency.
[0012] In one possible implementation, the first dependency object corresponds to multiple candidate dependency objects, and the different candidate dependency objects have different running efficiencies on the processor; wherein, the candidate dependency object is a dependency object with the same function as the first dependency object and written in a second programming language; the above-mentioned determination of the second dependency object corresponding to the first dependency object from the first correspondence based on the identifier of the first dependency object includes: determining multiple candidate dependency objects corresponding to the first dependency object from the first correspondence based on the identifier of the first dependency object; and determining the candidate dependency object whose running efficiency on the processor meets the condition as the second dependency object.
[0013] Based on the identifier of the first dependent object, this application determines multiple candidate dependent objects corresponding to the first dependent object from the first correspondence. Then, among these multiple candidate dependent objects, the candidate dependent object whose running efficiency on the processor meets the condition (e.g., the highest running efficiency) is determined as the second dependent object; thereby, when the first dependent object on which the first code depends is replaced with the second dependent object that meets the condition, the running efficiency of the code is improved.
[0014] In one possible implementation, the above-mentioned conversion of third code into second code based on the syntax rules of the second programming language includes: converting code blocks in the third code into code blocks of the second programming language with the same function based on the syntax rules of the third code and the syntax rules of the second programming language, thereby obtaining the second code.
[0015] The embodiments of this application are based on the syntax rules of the third code and the syntax rules of the second programming language. The code blocks in the third code are converted into code blocks that conform to the syntax rules of the second programming language and have the same function, thus obtaining the second code. It is not necessary to decompile the third code into code of the first programming language, thereby improving the code conversion efficiency.
[0016] In one possible implementation, before converting the third code into the second code based on the syntax rules of the second programming language, the method further includes: converting the memory allocation code in the third code into a target allocation code; wherein the target allocation code is used to determine a free storage unit that meets the allocation conditions in a storage unit pool; the storage unit pool includes: at least one storage unit; the conversion of the third code into the second code includes: converting the converted third code into the second code.
[0017] This application embodiment converts the memory allocation code in the third code into a target allocation code, so that the data can directly obtain storage units from the storage unit pool based on the target allocation code to store the data; it does not need to search for free storage units in the linked list; therefore, the running efficiency of the code is improved.
[0018] In one possible implementation, before converting the third code into second code based on the syntax rules of the second programming language, the method further includes: determining the reference relationship of the heap object in the third code; if the reference relationship indicates that the heap object is only referenced by a local variable, then inserting release code for releasing the heap object at the end of the local code block where the local variable is located; the above conversion of the third code into second code includes: converting the third code into second code after inserting the above release code.
[0019] This application embodiment determines the reference relationship of the heap object in the third code, and when the heap object is only referenced by local variables, it inserts release code for releasing the heap object at the end of the local code block where the local variable is located, so as to realize the timely release of the heap object, thereby improving the code running efficiency while saving memory space.
[0020] In one possible implementation, the first code is a user-defined function (UDF); or, the first code is an intermediate representation (IR) converted from the user-defined function (UDF).
[0021] In one possible implementation, the method further includes triggering the native execution engine to execute the second code.
[0022] Secondly, embodiments of this application provide a code optimization method.
[0023] In one implementation, the code optimization method includes: obtaining first code; the first code is binary code obtained by compiling the original code, and the original code is code based on a first programming language; obtaining the identifier of a first dependent object that the first code depends on from the first code; determining multiple candidate dependent objects corresponding to the first dependent object from a first correspondence based on the identifier of the first dependent object; wherein, different candidate dependent objects have different running efficiencies on the processor; determining the candidate dependent object whose running efficiency on the processor meets the condition as a second dependent object; and replacing the first dependent object that the first code depends on with the second dependent object.
[0024] In this embodiment, based on the identifier of the first dependent object upon which the first code depends, multiple candidate dependent objects corresponding to the first dependent object are determined from a first correspondence. Then, among these multiple candidate dependent objects, the candidate dependent object whose running efficiency on the processor meets the condition (e.g., the highest running efficiency) is determined as the second dependent object; thereby, by replacing the first dependent object upon which the first code depends with the second dependent object that meets the condition, the running efficiency of the code is improved.
[0025] In one implementation, the code optimization method includes: obtaining first code; the first code is binary code obtained by compiling the original code, and the original code is code based on a first programming language; converting the memory allocation code in the first code into target allocation code; wherein the target allocation code is used to determine a free storage unit that meets the allocation conditions in a storage unit pool; the storage unit pool includes at least one storage unit.
[0026] This application embodiment converts the memory allocation code in the first code into a target allocation code, so that data can directly obtain storage units from the storage unit pool based on the target allocation code to store the data; it does not need to search for free storage units in the linked list; therefore, the running efficiency of the code is improved.
[0027] In one implementation, the code optimization method includes: obtaining first code; the first code is binary code obtained by compiling the original code, and the original code is code based on a first programming language; determining the reference relationship of heap objects in the first code; if the reference relationship indicates that the heap object is only referenced by local variables, then inserting release code for releasing the heap object at the end of the local code block where the local variable is located.
[0028] This application embodiment determines the reference relationship of the heap object in the first code, and when the heap object is only referenced by local variables, it inserts release code for releasing the heap object at the end of the local code block where the local variable is located, so as to realize the timely release of the heap object, thereby improving the code running efficiency while saving memory space.
[0029] Thirdly, embodiments of this application provide a code optimization apparatus, which includes: an acquisition module and a processing module; the acquisition module is used to acquire first code; the first code is binary code obtained by compiling original code, and the original code is code based on a first programming language; the processing module is used to convert the first code into second code, the second code is code based on a second programming language with the same function as the first code; wherein, the code in the second programming language has higher running efficiency on the processor than the code in the first programming language; the second programming language is a programming language supported by a native execution engine, and the native execution engine runs on the processor.
[0030] In one implementation, a processing module is used to convert the dependency objects on which the first code depends into dependency objects with the same functionality written in a second programming language, thereby obtaining third code; wherein the types of dependency objects include: interfaces and / or classes; the processing module is used to convert the third code into second code based on the syntax rules of the second programming language.
[0031] In one implementation, an acquisition module is used to obtain the identifier of a first dependency object that the first code depends on from the first code; a processing module is used to determine a second dependency object corresponding to the first dependency object from a first correspondence based on the identifier of the first dependency object; the second dependency object is written in a second programming language; wherein the first dependency object and the second dependency object have the same function; the first correspondence is used to indicate the correspondence between a dependency object written in the first programming language and a dependency object written in the second programming language with the same function; the processing module is used to replace the first dependency object that the first code depends on with the second dependency object to obtain the third code.
[0032] In one implementation, a first dependent object corresponds to multiple candidate dependent objects, and different candidate dependent objects have different running efficiencies on the processor; wherein, a candidate dependent object is a dependent object with the same function as the first dependent object and written in a second programming language; the processing module is used to determine multiple candidate dependent objects corresponding to the first dependent object from a first correspondence based on the identifier of the first dependent object; wherein, different candidate dependent objects have different running efficiencies on the processor; a candidate dependent object is a dependent object with the same function as the first dependent object and written in a second programming language; the processing module is used to determine the candidate dependent object whose running efficiency on the processor meets the condition as the second dependent object.
[0033] In one implementation, the processing module is used to convert code blocks in the third code into code blocks in the second programming language with the same function, based on the syntax rules of the third code and the syntax rules of the second programming language, to obtain the second code.
[0034] In one implementation, a processing module is used to convert the memory allocation code in the third code into a target allocation code; wherein the target allocation code is used to determine a free storage unit in the storage unit pool that meets the allocation conditions; the storage unit pool includes at least one storage unit; the processing module is specifically used to convert the converted third code into a second code.
[0035] In one implementation, the processing module is used to determine the reference relationship of the heap object in the third code; the processing module is used to insert release code for releasing the heap object at the end of the local code block where the local variable is located if the reference relationship indicates that the heap object is only referenced by local variables; the processing module is specifically used to convert the third code with inserted release code into second code.
[0036] In one implementation, the first code is a user-defined function (UDF); or, the first code is an intermediate representation (IR) converted from the user-defined function (UDF).
[0037] In one implementation, the code optimization device further includes: a triggering module; and a processing module for triggering the native execution engine to execute the second code.
[0038] Fourthly, embodiments of this application provide a code optimization apparatus.
[0039] In one implementation, the code optimization device includes: an acquisition module and a processing module; the acquisition module is used to acquire first code; the first code is binary code obtained by compiling original code, and the original code is code based on a first programming language; the acquisition module is also used to acquire the identifier of a first dependent object that the first code depends on from the first code; the processing module is used to determine multiple candidate dependent objects corresponding to the first dependent object from a first correspondence based on the identifier of the first dependent object; wherein, different candidate dependent objects have different running efficiencies on the processor; the processing module is used to determine the candidate dependent object whose running efficiency on the processor meets the condition as a second dependent object; the processing module is also used to replace the first dependent object that the first code depends on with the second dependent object.
[0040] In one implementation, the code optimization device includes: an acquisition module and a processing module; the acquisition module is used to acquire first code; the first code is binary code obtained by compiling original code, and the original code is code based on a first programming language; the processing module is used to convert memory allocation code in the first code into target allocation code; wherein, the target allocation code is used to determine free storage units that meet the allocation conditions in a storage unit pool; the storage unit pool includes: at least one storage unit.
[0041] In one implementation, the code optimization device includes: an acquisition module and a processing module; the acquisition module is used to acquire first code; the first code is binary code obtained by compiling the original code, and the original code is code based on a first programming language; the processing module is used to determine the reference relationship of heap objects in the first code; the processing module is further used to insert release code for releasing the heap object at the end of the local code block where the local variable is located if the reference relationship indicates that the heap object is only referenced by a local variable.
[0042] Fifthly, embodiments of this application provide a computing device, which includes a memory and a processor, the memory being coupled to the processor; the memory is used to store computer program code, the computer program code including computer instructions; when the computer instructions are executed by the processor, the computing device performs the method described in the first aspect and the second aspect and any one of their possible implementations.
[0043] In a sixth aspect, embodiments of this application provide a computer storage medium including computer instructions that, when executed on a computing device, cause the computing device to perform the method described in the first aspect and the second aspect and any of their possible implementations.
[0044] In a seventh aspect, embodiments of this application provide a computer program product comprising computer instructions that, when executed on a computer, perform any one of the methods of the first aspect and the second aspect and their possible implementations.
[0045] Eighthly, this application provides a computing device cluster including at least one computing device, each computing device including a processor and a memory; the processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device, such that the computing device cluster performs the method described in the first aspect and the second aspect and any of their possible implementations.
[0046] It should be understood that the beneficial effects of the technical solutions of the third to eighth aspects of this application and their corresponding possible implementations can be found in the above description of the technical effects of the first and second aspects and their corresponding possible implementations, and will not be repeated here. Attached Figure Description
[0047] Figure 1 A schematic diagram illustrating the interaction between a JVM and a native execution engine, provided as an embodiment of this application;
[0048] Figure 2 A schematic diagram of a code processing system provided in an embodiment of this application;
[0049] Figure 3 This application provides a schematic diagram of the hardware structure of a computing device.
[0050] Figure 4 This is a schematic diagram of a code optimization method provided in an embodiment of this application;
[0051] Figure 5 This is a schematic flowchart of a first code conversion method provided in an embodiment of this application;
[0052] Figure 6 This application provides a schematic flowchart of a method for converting first code into third code in an embodiment of the present application.
[0053] Figure 7 This is a schematic diagram of another code optimization method provided in an embodiment of this application;
[0054] Figure 8 This is a schematic diagram of another code optimization method provided in an embodiment of this application;
[0055] Figure 9 This is a schematic diagram of another code optimization method provided in an embodiment of this application;
[0056] Figure 10 This is a schematic diagram of a code optimization device provided in an embodiment of this application. Detailed Implementation
[0057] In this article, the term "and / or" is merely a description of the relationship between related objects, indicating that there can be three relationships. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone.
[0058] The terms "first code" and "second code," etc., used in the specification and claims of this application are used to distinguish different codes, not to describe a specific order of codes. For example, "first programming language" and "second programming language," etc., are used to distinguish different programming languages, not to describe a specific order of programming languages.
[0059] In the embodiments of this application, the terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or design that is described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of the terms "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.
[0060] In the description of the embodiments in this application, unless otherwise stated, "multiple" means two or more. For example, multiple candidate dependency objects means two or more candidate dependency objects.
[0061] First, some concepts involved in the embodiments of this application will be explained as follows:
[0062] Intermediate representation (IR): A transitional form generated during the process of converting source code into machine code.
[0063] UDF: These are functions written by developers according to their needs, used to encapsulate specific logic.
[0064] Dependency: If a module (such as a class or function) needs to rely on the functionality of other modules (such as code blocks in third-party code) to complete its operation, then the module is considered to depend on the other module, and the other module is called the module's dependent object. For example, a class uses the functionality of another class in a specific scenario.
[0065] A dependency library refers to a third-party library or module that needs to be included during project runtime or compilation. It typically contains predefined classes, functions, or tools. A dependency library includes multiple modules (or dependency objects) that are used by the code in a UDF (User-Defined Function).
[0066] Heap: This is the area of memory used for dynamic allocation during program runtime; the heap is used to cache dynamically allocated objects.
[0067] A stack is a linear data structure used to cache local variables, function parameters, and other data.
[0068] JVM: An abstract computer that runs Java code.
[0069] Local variables are variables declared within a method, code block, or constructor. Their scope and lifetime are strictly limited to the code in which they are declared. Local variables are only valid within their respective code blocks, cannot be accessed from outside, and are destroyed (i.e., released) when the code block finishes execution.
[0070] Global variables: These are variables that can be accessed from anywhere in the program, and their scope covers the entire application.
[0071] Existing big data systems mainly run on the JVM; however, due to the low affinity between the JVM and the processor running the big data system (such as the processor's machine language), the operating efficiency of the big data system is low.
[0072] Common code optimization methods include Figure 1 As shown, the native execution engine retrieves the SQL code from the Java code package (i.e., the JAR file) from the JVM and executes it. However, if the SQL statement contains a Java UDF (User-Defined Function), the native execution engine cannot execute Java code. Therefore, the native execution engine needs to send the UDF to the JVM via JNI so that the JVM can execute the UDF and send the execution result back to the native execution engine via JNI so that the native execution engine can execute the SQL code based on the result.
[0073] Therefore, when Java UDFs exist in SQL code, the native execution engine needs to frequently interact with the JVM through JNI, resulting in very limited improvement in code execution efficiency.
[0074] Based on this, embodiments of this application provide a code optimization method. This method converts compiled code in a first programming language (e.g., Java) (i.e., the first code) into second code in a second programming language (e.g., C++) that has higher execution efficiency and the same functionality. Since the second programming language is a programming language supported by the native execution engine, the native execution engine can directly execute the second code without the intervention of the runtime platform (e.g., JVM) corresponding to the first programming language. This eliminates the problem of frequent interaction between the runtime platform and the native execution engine, thus improving the code's execution efficiency.
[0075] Furthermore, since the embodiments of this application convert the first code into second code supported by the native execution engine, so that the native execution engine can implement the function of the first code by executing the second code; and so that the native execution engine can implement the function of the first code by executing the second code corresponding to the first code, thereby increasing the applicability of the native execution engine.
[0076] The code optimization method provided in this application embodiment is applied to... Figure 2 The code processing system shown includes a code conversion device 101 and a native execution engine 102.
[0077] The code conversion device 101 is used to convert compiled code in a first programming language (e.g., Java) (i.e., the first code) into code in a second programming language (e.g., C++) with higher execution efficiency.
[0078] The following describes the conversion process of converting the first code into the second code, using the code conversion device 101, which includes a code parser 1011, a code optimizer 1012 (optional), a code converter 1013, and a code compiler 1014, as an example:
[0079] It should be noted that the embodiments of this application use Java as the first programming language and C++ as the second programming language as an example for illustration, and will not be repeated hereafter.
[0080] The code parser 1011 is used to parse the compiled Java code (i.e., the first code) from the code package (e.g., jar package) of the software system and send the first code to the code optimizer 1012.
[0081] The code optimizer 1012 is used to optimize the first code; for the specific optimization process, please refer to code optimization schemes 1 to 3 in the following method embodiments, which will not be repeated here.
[0082] The code converter 1013 is used to convert optimized first code into second code. For example, by parsing the optimized first code, the dependent interfaces in the first code are identified. Then, the dependent interfaces in the first code are replaced with functionally identical dependent interfaces written in a second programming language; and the code blocks in the first code are converted into functionally identical code blocks written in the second programming language, thereby obtaining the second code.
[0083] The code compiler 1014 is used to obtain the second code generated by the code converter 1013 and compile the second code into a native release.
[0084] For example, a native distribution includes: a binary code set (e.g., a .so file containing C++ code), a dependency set, and a native management table. The binary code set includes: the binary code of the second codebase and resource files (e.g., images and configuration files).
[0085] The dependency set includes the dependency objects that the second code depends on; the types of these dependency objects include interfaces and / or classes.
[0086] The native management table includes the correspondence between the first code and the second code, which is used to indicate when the second code is executed to implement the function of the first code.
[0087] The native execution engine 102 includes an executor 1021. The executor 1021 is used to execute native releases to implement the functionality of the first code.
[0088] It should be noted that the code conversion device 101 and the native execution engine 102 can be functional modules deployed on different computing devices, or they can be different functional modules integrated on the same computing device. Specifically, this application does not limit the computing devices on which the code conversion device 101 and the native execution engine 102 are deployed.
[0089] In one implementation, the code conversion device 101 may be a functional module integrated into the native execution engine 102; that is, the native execution engine 102 is an engine with functions such as code optimization and code execution.
[0090] For example, Figure 3 It can be deployed Figure 2 The diagram shows the hardware structure of a computing device including a code conversion device 101 and / or a native execution engine 102. This computing device may include a processor 201, a memory 202, and a communication interface 203. The processor 201, memory 202, and communication interface 203 can be connected via a bus 204 or other means.
[0091] Processor 201 includes one or more central processing units (CPUs). The CPU can be a single-core CPU or a multi-core CPU. Optionally, processor 201 may also include a graphics processing unit (GPU), a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), etc.
[0092] The memory 202 includes, but is not limited to, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, or optical memory, disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer.
[0093] In the embodiments of this application, if Figure 3 If the computing device shown is equipped with a code conversion device 101, then the memory 202 can store the first code. If Figure 3 If the computing device shown is a device with a native execution engine 102 deployed, then the memory 202 can store the second code.
[0094] In one possible implementation, the memory 202 can exist independently of the processor 201. The memory 202 can be connected to the processor 201 via a bus 204 and is used to store data, operating system, instructions, or program code. When the processor 201 calls and executes the instructions or program code stored in the memory 202, it can implement the relevant steps in the code optimization method provided in the embodiments of this application.
[0095] In another possible implementation, the memory 202 can also be integrated with the processor 201.
[0096] The communication interface 203 can be a transceiver unit used to communicate with other devices or communication networks, such as Ethernet, RAN, wireless local area networks (WLAN), etc. The communication interface 203 can receive commands, messages, or data. The transceiver unit can be a transceiver or similar device.
[0097] Optionally, the communication interface 203 can also be a transceiver circuit located within the processor 201, used to implement signal input and signal output of the processor 201. The communication interface 203 can be a wired interface (port), such as a fiber distributed data interface (FDDI) or a gigabit Ethernet (GE) interface, or the communication interface 203 can also be a wireless interface.
[0098] Bus 204 can be an industry standard architecture (ISA) bus, a peripheral component interconnect (PCI) bus, or an extended industry standard architecture (EISA) bus, etc. This bus can be divided into address bus, data bus, control bus, etc. The bus can also be divided into serial bus and parallel bus. For ease of representation, Figure 3 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.
[0099] It should be understood that, Figure 3 The computing device mentioned is merely one example of a computing device; it can have more than Figure 3 The more or fewer components shown can be combined into two or more components, or they can have different component configurations. For example, a computing device can also include a smart network card, such as a data processing unit (DPU).
[0100] The embodiments of this application can be based on a computing device (such as...) Figure 3 The code optimization method provided in this application can be executed on multiple computing devices (i.e., a computing device cluster). Specifically, the number of computing devices executing the code optimization method is not limited in the specific embodiments of this application.
[0101] When the code optimization method provided in this application is executed on multiple computing devices, the memory 202 of the multiple computing devices may store the same instructions for executing the code optimization method, so that the multiple computing devices execute the code optimization method provided in this application independently. Alternatively, the multiple computing devices may each store partial instructions for executing the code optimization method, so that the combination of the multiple computing devices jointly executes the instructions for executing the code optimization method.
[0102] It should be understood that the system architecture and application scenarios described in the embodiments of this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided in the embodiments of this application. As those skilled in the art will know, with the evolution of system architecture and the emergence of new business scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.
[0103] It should be noted that the code optimization method provided in this application optimizes the code from two aspects: code syntax conversion and code semantic optimization. Specifically, the code syntax conversion aspect improves the execution efficiency by converting the compiled code in a first programming language (hereinafter referred to as "first code") into code in a second programming language, which has higher execution efficiency; see the section on code syntax conversion for details. The code semantic optimization aspect improves the execution efficiency by optimizing the less efficient semantics (such as code logic) in the code; see the section on code semantic optimization for details.
[0104] Code syntax conversion
[0105] This application provides a code optimization method, which is applied to... Figure 2 The code conversion device 101 (hereinafter referred to as the code conversion device); such as Figure 4 As shown, the method includes: S110-S130.
[0106] S110, Obtain the first code.
[0107] The first code is the binary code obtained by compiling the original code, which is written in the first programming language; that is, the first code is the compiled code in the first programming language, i.e., the first code is the compiled original code.
[0108] For example, if the first programming language is Java, then the first code is the compiled Java code; specifically, the first code is the binary code (i.e., bytecode) in the .class file (i.e., bytecode file) after compiling the Java code. If the first programming language is Python, then the first code is the compiled Python code; specifically, the first code is the binary code in the .pyc file after compiling the Python code.
[0109] The programming language in this application embodiment can be a procedural programming language (such as C language) or an object-oriented programming language (such as Java language). Specifically, this application embodiment does not limit the type of programming language.
[0110] The first code in this application can be the entire compiled source code, a UDF in the compiled source code, or an IR converted from a UDF in the compiled source code. Specific embodiments of this application do not impose specific limitations on the first code.
[0111] It should be understood that, taking Java code as an example, Java code includes: code written by the user based on requirements (i.e., UDF) and the dependent objects that the UDF depends on (such as the third-party interfaces called by the UDF); that is, Java code includes UDF and non-UDF (i.e., the dependent objects of the third parties called by the UDF); in other words, the first code includes both UDF and non-UDF.
[0112] It should be noted that when the first code is the compiled original code, the implementation of S110 is as follows: Method 1; when the first code is a UDF, the implementation of S110 is as follows: Method 2; and when the first code is an IR, the implementation of S110 is as follows: Method 3.
[0113] Method 1: Obtain the first code, which is the binary code after compiling the original code.
[0114] It should be noted that Method 1 can be implemented by the code conversion device receiving the first code sent by the native execution engine; or by the code conversion device obtaining the first code from a preset location; the specific implementation of Method 1 is not limited in this application embodiment.
[0115] Method 2: Obtain the compiled source code (i.e., binary code) and extract the UDF from the compiled source code.
[0116] In this application, the specific implementation of obtaining a UDF from the compiled source code can be based on the keywords or tags corresponding to the UDF, identifying and obtaining the UDF from the compiled source code; alternatively, it can be based on regular expressions used to represent the UDF. This application does not specifically limit the implementation method for obtaining the UDF.
[0117] Method 3: Based on obtaining the UDF using Method 2, convert the UDF to IR.
[0118] For the specific implementation of converting UDF to IR in this application, please refer to the relevant technologies, which will not be repeated here.
[0119] It should be noted that the embodiments of this application take the IR converted from UDF as an example for explanation, and will not be repeated hereafter.
[0120] S120, Convert the first code to the second code.
[0121] The second code in this application is code based on a second programming language that has the same function as the first code; that is, the second code is code written in a second programming language that has the same function as the first code.
[0122] In this application, the code in the second programming language runs more efficiently on the processor than the code in the first programming language; that is, the second code in this application runs more efficiently on the processor than the first code. The second programming language is a programming language (e.g., C++) supported by the native execution engine, which runs on the functional modules of the aforementioned processor.
[0123] The second code in this application can be code in a second programming language, or it can be binary code corresponding to the code in the second programming language (i.e., binary code after compiling the code in the second programming language). The specific embodiments of this application do not limit the form of the second code.
[0124] It should be noted that the embodiments of this application use code in the second programming language as an example for illustration, and will not be repeated hereafter.
[0125] The first programming language in this application can be any one of the following programming languages: C, C++, C#, Java, and Python. The second programming language can be any one of the following programming languages: C, C++, C#, Java, and Python, and its execution efficiency is higher than that of the first programming language. Specific embodiments of this application do not impose specific limitations on the first and second programming languages.
[0126] In the embodiments of the present application, the first programming language is Java programming language, the second programming language is C++ programming language, and the first code is jimple IR in the IR corresponding to the bytecode of Java code, which is taken as an example for description, and will not be repeated subsequently.
[0127] In the present application, S120 can be implemented by the following method 1 or method 2, which is specifically as follows:
[0128] Method 1: decompile the first code into original code (e.g., Java code); then, convert the original code into the second code based on the grammatical rules of the first programming language and the grammatical rules of the second programming language (e.g., C++).
[0129] In the present application, decompiling the first code into original code can be implemented based on a decompiler (e.g., javap -c -v Demo); for specific implementation, please refer to the related art, which will not be repeated here.
[0130] In the present application, based on the grammatical rules of the first programming language (e.g., Java) and the grammatical rules of the second programming language (e.g., C++), the specific implementation of converting the original code into the second code includes at least one of the following: 1) converting the definition mode of classes in the original code into a definition mode of classes that conforms to the grammatical rules of the second programming language. 2) converting the lambda expression in the original code into a code logic that conforms to the grammatical rules of the second programming language. 3) converting the nested method body into a method body that conforms to the grammatical rules of the second programming language. 4) converting the expression (e.g., arithmetic expression) in the original code into an expression with the same function that conforms to the grammatical rules of the second programming language. 5) converting the data type in the original code into a data type that conforms to the grammatical rules of the second programming language. 6) converting the value access mode in the original code into a value access mode that conforms to the grammatical rules of the second programming language.
[0131] For example, assuming that the first programming language is Java, the second programming language is C++, and the original code is a function for printing the "Hello World" string as shown in the following code block 1. Then, the process of converting the original code into the second code includes: replacing the function signature "public void hellWord(){}" in the original code with the C++ function signature "void hellWord(){}"; replacing the print statement "System.out.println("Hello World");" in the original code with the C++ print statement "std::cout<<"Hello World"<<std::endl;"; thereby obtaining the second code as shown in the following code block 2.
[0132]
[0133] Method 2: such as Figure 5 As shown, the implementation of S120 includes: S121-S122.
[0134] S121. Convert the dependency objects that the first code depends on into dependency objects with the same function written in the second programming language to obtain the third code.
[0135] The types of dependent objects in this application include interfaces and / or classes. These dependent objects can be interfaces and / or classes from third-party libraries, or interfaces and / or classes from standard libraries; specific embodiments of this application do not limit the source of the dependent objects.
[0136] The specific implementation of S121 is described below and will not be repeated here.
[0137] S122. Based on the syntax rules of the second programming language, convert the third code into the second code.
[0138] In one implementation, S122 includes: based on the syntax rules of a third code (e.g., the syntax rules of Jimple IR) and the syntax rules of a second programming language (e.g., C++), converting code blocks in the third code into code blocks that conform to the syntax rules of the second programming language and have the same function, thus obtaining the second code.
[0139] The third code in this application is the first code after converting the dependent object; that is, the syntax rules of the code blocks in the third code are consistent with the syntax rules of the first code. Based on this, the specific implementation of converting the code blocks in the third code into code blocks that conform to the syntax rules of the second programming language and have the same function is described in the relevant description of method 1 in the implementation of S121 below, and will not be repeated here.
[0140] The embodiments of this application are based on the syntax rules of the third code and the syntax rules of the second programming language. The code blocks in the third code are converted into code blocks that conform to the syntax rules of the second programming language and have the same function, thus obtaining the second code. It is not necessary to decompile the third code into code of the first programming language, thereby improving the code conversion efficiency.
[0141] In another implementation, S122 includes: decompiling the code block of the third code into the original code block; and then converting the original code block into a code block of the second programming language with the same functionality.
[0142] S130, triggers the native execution engine to execute the second code.
[0143] The implementation of S130 in this application includes: compiling the second code to obtain the compiled second code; and then, notifying the native execution engine to execute the compiled second code.
[0144] It should be noted that the specific implementation of notifying the native execution engine to execute the compiled second code in this application may involve generating a mapping relationship between the first code and the compiled second code (e.g., a native management table) and sending a notification message to the native execution engine. This allows the native execution engine to retrieve the compiled second code corresponding to the first code to be executed from the mapping relationship and execute the compiled second code. Alternatively, it may involve sending a notification message including an identifier of the compiled second code to the native execution engine, enabling the native execution engine to retrieve and execute the compiled second code. Specifically, this application does not limit the specific implementation of notifying the native execution engine to execute the compiled second code.
[0145] It should be understood that, in the case where the code conversion device in this application is a functional module integrated into the native execution engine, the native execution engine compiles and executes the second code after S120.
[0146] In this embodiment, the compiled code in a first programming language (i.e., the first code) is converted into second code in a second programming language with the same functionality. Since the code in the second programming language is more efficient on the processor than the code in the first programming language, and the second programming language is supported by the native execution engine, the native execution engine can directly run this more efficient second code without the intervention of the runtime platform (e.g., JVM) of the first programming language (e.g., Java). Therefore, there is no problem of frequent interaction between the runtime platform and the native execution engine, thus improving the code's execution efficiency.
[0147] Furthermore, since the embodiments of this application convert the first code into second code supported by the native execution engine, so that the native execution engine can implement the function of the first code by executing the second code; and so that the native execution engine can implement the function of the first code by executing the second code corresponding to the first code, thereby increasing the applicability of the native execution engine.
[0148] The following explains the specific implementation of "converting the first code into the third code", i.e., S121:
[0149] Method 1: Convert the syntax rules of the dependency objects on which the first code depends (e.g., the syntax rules of Jimple IR) to the syntax rules of the second programming language (e.g., C++) to obtain the third code.
[0150] Specifically, based on the syntax rules of the first code (e.g., the syntax rules of Jimple IR) and the syntax rules of the second programming language (e.g., C++), the code in the dependent objects on which the first code depends is converted line by line into code that conforms to the syntax rules of the second programming language.
[0151] The specific implementation includes: converting the variable declaration method in the above-mentioned dependent object into a variable declaration method that conforms to the syntax rules of the second programming language; converting the function signature method in the dependent object into a function signature method that conforms to the syntax rules of the second programming language; converting the class definition method in the dependent object into a class definition method that conforms to the syntax rules of the second programming language; and converting the keywords in the dependent object into keywords with the same semantics in the second programming language; the specific details will not be listed one by one in this application.
[0152] Method 2: Directly replace the dependency objects of the first code with dependency objects of the same functionality in the second programming language to obtain the third code. Specifically, as follows... Figure 6 The diagram shows S121A-S121C.
[0153] S121A. Obtain the identifier of the first dependent object that the first code depends on from the first code.
[0154] The first dependent object in this application can be any dependent object that the first code depends on, or it can be multiple dependent objects that the first code depends on (e.g., all dependent objects); specifically, the embodiments of this application do not limit the first dependent object.
[0155] It should be understood that when you need to call a dependent object in your code, you need to first import the dependency package containing the dependent object, and then call the dependent object by specifying the name of the dependent object in that dependency package.
[0156] It should be understood that, for example, to avoid the existence of dependency objects with the same name in different dependency packages, the identifier of the first dependency object in this application includes: the absolute address of the first dependency object (including: the name of the dependency package) and the name of the first dependency object.
[0157] The implementation of S121A in this application can be based on a regular expression used to characterize the identifier of the dependent object, to obtain the identifier of the first dependent object from the first code, or it can be based on a keyword used to indicate the import or invocation of the dependent object, to obtain the identifier of the first dependent object from the first code; the specific implementation of S121A in this application does not limit the implementation of S121A.
[0158] S121B. Based on the identifier of the first dependent object, determine from the first correspondence whether there is a second dependent object corresponding to the first dependent object.
[0159] If a second dependent object exists in the first correspondence, execute S121C as follows.
[0160] If no second dependent object exists in the first correspondence, execute the termination action to end the current method flow.
[0161] In this application, the second dependent object is an interface and / or class written in a second programming language; wherein, the first dependent object and the second dependent object have the same function; that is, the second dependent object is a dependent object written in a second programming language with the same function as the first dependent object.
[0162] For example, if the first dependency object is used to implement the addition operation of x and y, then the second dependency object is an interface written in C++ to implement the addition operation of x and y.
[0163] The first correspondence is used to indicate the correspondence between dependency objects written in the first programming language and dependency objects with the same functionality written in the second programming language.
[0164] S121C: Replace the first dependency object that the first code depends on with the second dependency object to obtain the third code.
[0165] For example, suppose the first code implements the addition operation of x and y by calling the first dependent object's add interface; and suppose the dependent object used in the second programming language to implement the addition operation of x and y is the addxy interface; then, replace the add interface called in the first code with the addxy interface, and replace the code in the first code that imports the dependent package A containing the addxy interface with the code that imports the dependent package B containing the addxy interface.
[0166] Based on the principle of substitution, the embodiments of this application directly replace the first dependent object on which the first code depends with a second dependent object with the same function written in a second programming language, thereby realizing the replacement of the dependent object in the first code without needing to convert the syntax rules of the code in the first dependent object, thus improving the code conversion efficiency.
[0167] Code semantic optimization
[0168] The code semantic optimization aspect of this application's embodiments is approached from three perspectives; specifically as follows:
[0169] Firstly, from the perspective of the affinity between the second dependent object and the processor running the code optimization method of this application, the dependent object whose running efficiency on the processor meets the condition is determined as the second dependent object, thereby improving the running efficiency of the code. See the code optimization scheme 1 below for details.
[0170] Secondly, starting from the timing of releasing heap objects (i.e., objects cached in the heap), the heap objects are released in a timely and reasonable manner, thereby saving heap storage resources while improving code running efficiency. See code optimization scheme 2 below for details.
[0171] Thirdly, from the perspective of allocating memory space for the variable, the storage space used to store the variable can be quickly determined, thereby improving the running efficiency of the code. See code optimization scheme 3 below for details.
[0172] It should be noted that any number of code optimization schemes from code optimization scheme 1 to code optimization scheme 3 can be used in combination or independently. In addition, code optimization schemes from code optimization scheme 1 to code optimization scheme 3 (i.e., code semantic optimization schemes) and code syntax transformation schemes can be used in combination or independently (i.e., only at least one code optimization scheme from code optimization scheme 1 to code optimization scheme 3 is executed on the first code).
[0173] Code optimization solution 1
[0174] This application embodiment provides multiple dependent objects with the same function written in a second programming language for the dependent objects on which the first code depends. These multiple dependent objects have different running efficiencies on different types of processors.
[0175] Based on this, embodiments of this application provide another code optimization method, such as... Figure 7 As shown, the method includes: S210-S270.
[0176] S210, Obtain the first code.
[0177] It should be noted that the implementation of S210 is the same as that of S110. For a detailed description of S210, please refer to the relevant description of S110 above. It will not be repeated here.
[0178] S220. Obtain the identifier of the first dependent object from the first code.
[0179] It should be noted that the implementation of S220 is consistent with the implementation of S121A. For a detailed description of S220, please refer to the relevant description of S121A above, which will not be repeated here.
[0180] S230. Based on the identifier of the first dependent object, determine multiple candidate dependent objects corresponding to the first dependent object from the first correspondence.
[0181] In this application, any one of the multiple candidate dependency objects is a dependency object written in a second programming language with the same function as the first dependency object. These multiple candidate dependency objects have different running efficiencies on the processor; specifically, they have different running efficiencies on different types of processors.
[0182] For example, suppose the above candidate dependencies include candidate dependency A and candidate dependency B. Candidate dependency A has a runtime efficiency of 7 on an i3 CPU (meaning it completes in 7 seconds); candidate dependency A has a runtime efficiency of 4 on an i5 CPU. Candidate dependency B has a runtime efficiency of 3 on an i3 CPU; candidate dependency B has a runtime efficiency of 6 on an i5 CPU.
[0183] S240: The candidate dependency objects that meet the running efficiency requirements on the processor are identified as the second dependency objects.
[0184] The processor in S240 of this application is a processor that runs second code; that is, the processor is a processor that runs the native execution engine.
[0185] In this application, S240 is implemented by: obtaining the target type of the processor running the native execution engine; and then, based on the target type, determining the candidate dependent objects in the second correspondence that meet the conditions (e.g., the highest running efficiency on the processor of the target type) as the second dependent objects.
[0186] In this application, the specific implementation of obtaining the target type of the processor used to run the native execution engine includes the following:
[0187] When the data converter and the native execution engine are two independent functional modules, the data converter sends a type retrieval request to the native execution engine; in response to the type retrieval request, the native execution engine sends the target type of the processor in which the native execution engine resides to the data converter.
[0188] When the data converter is a functional module integrated into the native execution engine, the data converter obtains the type of the processor running the data converter as the target type.
[0189] The second correspondence in this application includes: the correspondence between the candidate dependent object and the running efficiency of the candidate dependent object on different types of processors.
[0190] For example, based on the example in S230, when the processor is an i3-CPU, candidate dependency object B is determined as the second dependency object. When the processor is an i5-CPU, candidate dependency object A is determined as the second dependency object.
[0191] S250. Replace the first dependent object that the first code depends on with the second dependent object to obtain the third code.
[0192] It should be noted that the implementation of S250 is consistent with that of S121C. For a detailed description of S250, please refer to the relevant description of S121C above. It will not be repeated here.
[0193] It should be understood that S220-S250 is another implementation of S121.
[0194] Based on the identifier of the first dependent object, this application determines multiple candidate dependent objects corresponding to the first dependent object from the first correspondence. Then, among these multiple candidate dependent objects, the candidate dependent object whose running efficiency on the processor meets the condition (e.g., the highest running efficiency) is determined as the second dependent object; thereby, when the first dependent object on which the first code depends is replaced with the second dependent object that meets the condition, the running efficiency of the code is improved.
[0195] S260, based on the syntax rules of the second programming language, converts the third code into the second code.
[0196] It should be noted that the implementation of S260 is the same as that of S122. For a detailed description of S260, please refer to the relevant description of S122 above. It will not be repeated here.
[0197] It should be understood that S220-S250 can be executed before or after S260. Specifically, this application does not limit the execution order of S220-S250 and S260 in any particular embodiment.
[0198] S270, triggers the native execution engine to execute the second code.
[0199] It should be noted that the implementation of S270 is the same as that of S130. For a detailed description of S270, please refer to the relevant description of S130 above. It will not be repeated here.
[0200] This embodiment of the application obtains third code by converting the dependent objects in the first code into functionally identical dependent objects written in a second programming language; then, based on the syntax rules of the second programming language, the third code is converted into second code in the second programming language. Since the code in the second programming language runs more efficiently on the processor than the code in the first programming language, and the second programming language is supported by the native execution engine, the native execution engine can directly run this more efficient second code without the intervention of the runtime platform (e.g., JVM) of the first programming language (e.g., Java); thus, there is no problem of frequent interaction between the runtime platform and the native execution engine, thereby improving the code's running efficiency.
[0201] Code optimization solution 2
[0202] It's important to note that because local variables are stored on the stack, they are released after the local code block (e.g., a function) containing them has finished executing (e.g., the stack space occupied by the local variable is marked as overwriteable). However, because heap objects are stored in the heap, their lifecycle is not affected by local variables referencing them on the stack. Consequently, even when no variables reference them, heap objects are not immediately released but are instead freed through periodic garbage collection (GC). This results in unreferenced heap objects (i.e., invalid data) occupying heap space. When heap space is insufficient, heap objects to be created must wait for invalid data in the heap to be released before creation, thus reducing code execution efficiency.
[0203] Furthermore, programming languages are divided into those with automatic garbage collection (e.g., Java) and those with manual garbage collection (e.g., C++). Therefore, if the first programming language is one with automatic garbage collection, the third code will not contain any code for manual garbage collection. Consequently, when the third code is finally converted to C++ (i.e., a language with manual garbage collection) as the second code, the lack of manual garbage collection code in this second code will cause it to malfunction.
[0204] Based on this, embodiments of this application provide a code optimization method, such as... Figure 8 The method shown includes: S310-S370.
[0205] S310, Obtain the first code.
[0206] It should be noted that the implementation of S310 is the same as that of S110. For a detailed description of S310, please refer to the relevant description of S110 above. It will not be repeated here.
[0207] S320. Convert the dependency objects that the first code depends on into dependency objects with the same function written in the second programming language to obtain the third code.
[0208] It should be noted that the implementation of S320 is the same as that of S121. For a detailed description of S320, please refer to the relevant description of S121 above. It will not be repeated here.
[0209] S330. Determine the reference relationships of heap objects in the third code.
[0210] The implementation of S330 in this application includes: identifying heap objects and variables (such as local variables and / or global variables) from third-party code, and then determining the reference relationship between the heap objects and variables based on keywords or symbols used to indicate the creation of reference relationships (such as the "=" assignment operator). For specific implementation details, please refer to relevant technologies, which will not be elaborated here.
[0211] It should be understood that heap objects can be identified from third-party code based on the keywords used to create them (e.g., "new"). Similarly, variables can be identified from third-party code based on their declaration method (e.g., data type variable name); specifically, since local and global variables are declared in different locations, their declaration location can be used to determine whether they are local or global.
[0212] For example, suppose the third code includes "temp$0 = new Student"; then, since "new" is the keyword used to create heap objects, "new Student" represents the created heap object, temp$0 is a variable declaration, and "=" is the symbol used to create a reference relationship; therefore, the code includes the heap object "new Student" and the variable "temp$0"; the reference relationship is that the variable "temp$0" references (i.e. points to) the heap object "new Student".
[0213] S340. Determine the reference relationships of heap objects, indicating whether heap objects are referenced only by local variables.
[0214] If the heap object is only referenced by local variables, execute S350 as follows.
[0215] If a heap object is not only referenced by local variables, perform the termination action.
[0216] S350. At the end of the local code block containing the local variable, insert release code to release the heap object.
[0217] In this application, the release code is used to release the storage space of the heap object; subsequently, when the subsequent code runs to the end of the local code block, it releases the unreferenced heap object by running the release code.
[0218] It should be noted that S330-S350 can be executed after S320 or before S320; specifically, the execution order of S330-S350 and S320 is not limited in this application embodiment.
[0219] This application embodiment determines the reference relationship of the heap object in the third code, and when the heap object is only referenced by local variables, it inserts release code for releasing the heap object at the end of the local code block where the local variable is located, so as to realize the timely release of the heap object, thereby improving the code running efficiency while saving memory space.
[0220] Furthermore, since this application executes after converting the first code into the third code, it is not necessary to execute the heap object release step if the conversion of the first code into the third code fails, thus saving the processing resources of the code conversion device.
[0221] S360, based on the syntax rules of the second programming language, converts the third code that inserts release code into the second code.
[0222] It should be noted that the implementation of S360 is similar to that of S122. For a detailed description of S360, please refer to the relevant description of S122 above. It will not be repeated here.
[0223] It should be understood that S330-S350 can be executed before or after S360; specifically, the execution order of S330-S350 and S360 is not limited in the embodiments of this application.
[0224] S370 triggers the native execution engine to execute the second code.
[0225] It should be noted that the implementation of S370 is the same as that of S130. For a detailed description of S370, please refer to the relevant description of S130 above. It will not be repeated here.
[0226] This embodiment of the application obtains third code by converting the dependent objects in the first code into functionally identical dependent objects written in a second programming language; and when the heap object in the third code is only referenced by local variables, release code for releasing the heap object is inserted at the end of the local code block where the local variable is located. Then, based on the syntax rules of the second programming language, the third code with the inserted release code is converted into second code. Since the code in the second programming language runs more efficiently on the processor than the code in the first programming language, and the second programming language is a programming language supported by the native execution engine, the native execution engine can directly run the more efficient second code without the intervention of the runtime platform (e.g., JVM) of the first programming language (e.g., Java); thus, there is no problem of frequent interaction between the runtime platform and the native execution engine, thereby improving the running efficiency of the code.
[0227] Code optimization solution 3
[0228] It's important to note that each time a variable is declared, memory allocation code (such as the `malloc` interface) needs to be called to allocate storage space for that variable. However, since memory storage units are represented as linked lists, allocating storage space for variable A involves searching through the linked list one by one using a pointer. This search efficiency is low, thus reducing the overall efficiency of the code.
[0229] Based on this, the embodiments of this application provide yet another code optimization method, such as... Figure 9 The method shown includes: S410-S460.
[0230] S410, Obtain the first code.
[0231] It should be noted that the implementation of S410 is the same as that of S110. For a detailed description of S410, please refer to the relevant description of S110 above. It will not be repeated here.
[0232] S420. Convert the dependency objects that the first code depends on into dependency objects with the same function written in the second programming language to obtain the third code.
[0233] It should be noted that the implementation of S420 is the same as that of S121. For a detailed description of S420, please refer to the relevant description of S121 above. It will not be repeated here.
[0234] S430, Identify memory allocation codes from third-party code.
[0235] In this application, memory allocation code is code that allocates storage units in memory for data such as variables, objects, or arrays (e.g., an interface). The storage units allocated based on the memory allocation code are used to store the variable, object, or array data.
[0236] In this application, S430 is implemented by identifying memory allocation codes from the third code based on the syntax rules of the third code. Specifically, this can be done by identifying memory allocation codes in the third code based on keywords used for memory allocation; or by identifying memory allocation codes from the third code using regular expressions used to characterize memory allocations. The specific implementation of S430 in this application does not limit the specific implementation method of S430.
[0237] S440. Convert the memory allocation code in the third code to the target allocation code.
[0238] In this application, the target application code is used to determine a free storage unit that meets the application conditions in a storage unit pool; the storage unit pool includes at least one storage unit (i.e., a free storage unit); wherein the application conditions are used to indicate the size and / or whether the storage unit to be applied for is contiguous or other attributes.
[0239] It should be noted that the storage unit pool in this application is used to provide free storage units for data. When storage unit A is requested from this storage unit pool for data A through the object code, storage unit A is deleted from the storage unit pool. After releasing storage unit A that stores data A, storage unit A is added back to the storage unit pool.
[0240] It should be understood that when executing S430-S440, S130 converts the third code (i.e., the memory allocation code into the target allocation code) into the second code.
[0241] Optionally, if the number of storage units in the storage unit pool is less than the first preset number, a second preset number of free storage units in memory are added to the storage unit pool.
[0242] This application embodiment converts the memory allocation code in the third code into a target allocation code, so that the data can directly obtain storage units from the storage unit pool based on the target allocation code to store the data; it does not need to search for free storage units in the linked list; therefore, the running efficiency of the code is improved.
[0243] It should be noted that S430-S440 can be executed after S420 or before S420; specifically, the execution order of S430-S440 and S420 is not limited in this application embodiment.
[0244] S450 converts the transformed third code into second code based on the syntax rules of the second programming language.
[0245] The converted third code in S450 of this application is the third code that converts the memory allocation code into the target allocation code.
[0246] It should be noted that the implementation of S450 is similar to that of S122. For a detailed description of S450, please refer to the relevant description of S122 above. It will not be repeated here.
[0247] It should be understood that S430-S440 can be executed before or after S450; specifically, the execution order of S430-S440 and S450 is not limited in the embodiments of this application.
[0248] S460 triggers the native execution engine to execute the second code.
[0249] It should be noted that the implementation method of S460 is the same as that of S130. For a detailed description of S460, please refer to the relevant description of S130 above. It will not be repeated here.
[0250] This embodiment of the application obtains third code by converting the dependent objects in the first code into functionally identical dependent objects written in a second programming language; and converting the memory allocation code in the third code into target allocation code. Then, based on the syntax rules of the second programming language, the converted third code is converted into second code. Since the code in the second programming language runs more efficiently on the processor than the code in the first programming language, and the second programming language is a programming language supported by the native execution engine, the native execution engine can directly run the more efficient second code without the intervention of the runtime platform (e.g., JVM) of the first programming language (e.g., Java); thus, there is no problem of frequent interaction between the runtime platform and the native execution engine, thereby improving the running efficiency of the code.
[0251] The foregoing mainly describes the solutions provided by the embodiments of this application from a methodological perspective. To achieve the above functions, the code optimization apparatus includes hardware structures and / or software modules corresponding to the execution of each function. Those skilled in the art should readily recognize that, based on the modules and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed in hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0252] This application embodiment can, based on the above method, exemplarily divide the code optimization device into functional modules. For example, the code optimization device may include functional modules corresponding to each functional division, or two or more functions may be integrated into one processing module. The integrated module can be implemented in hardware or as a software functional module. It should be noted that the module division in this application embodiment is illustrative and only represents one logical functional division; in actual implementation, there may be other division methods.
[0253] When dividing each function into modules according to its corresponding function. Figure 10 A schematic diagram of a possible structure of the code optimization apparatus involved in the above embodiments is shown. For example... Figure 10 As shown, the code optimization device includes an acquisition module 1110 and a processing module 1120.
[0254] The acquisition module 1110 is used to acquire the first code; for example, by executing step S110 in the above method embodiment.
[0255] The processing module 1120 is used to convert the first code into the second code; for example, to execute step S120 in the above method embodiment.
[0256] Optionally, the processing module 1120 is used to convert the dependency objects on which the first code depends into dependency objects with the same function written in the second programming language to obtain the third code; for example, by executing step S121 in the above method embodiment.
[0257] The processing module 1120 is used to convert the third code into the second code based on the syntax rules of the second programming language; for example, it executes step S122 in the above method embodiment.
[0258] Optionally, the acquisition module 1110 is used to obtain the identifier of the first dependent object that the first code depends on from the first code; for example, by executing step S121A in the above method embodiment.
[0259] The processing module 1120 is used to determine the second dependent object corresponding to the first dependent object from the first correspondence based on the identifier of the first dependent object; for example, by executing step S121B in the above method embodiment.
[0260] The processing module 1120 is used to replace the first dependent object on which the first code depends with the second dependent object to obtain the third code; for example, by executing step S121C in the above method embodiment.
[0261] Optionally, the processing module 1120 is used to determine multiple candidate dependent objects corresponding to the first dependent object from the first correspondence based on the identifier of the first dependent object; for example, by executing step S230 in the above method embodiment.
[0262] The processing module 1120 is used to determine the candidate dependency objects that meet the running efficiency conditions on the processor as the second dependency objects; for example, to execute step S240 in the above method embodiment.
[0263] Optionally, the processing module 1120 is used to convert code blocks in the third code into code blocks in the second programming language with the same function, based on the syntax rules of the third code and the syntax rules of the second programming language, to obtain the second code.
[0264] Optionally, the processing module 1120 is used to convert the memory allocation code in the third code into the target allocation code; for example, by performing step S440 in the above method embodiment.
[0265] The processing module 1120 is used to convert the converted third code into second code; for example, by performing step S450 in the above method embodiment.
[0266] Optionally, the processing module 1120 is used to determine the reference relationships of heap objects in the third code; for example, by performing step S330 in the above method embodiment.
[0267] The processing module 1120 is used to insert release code for releasing the heap object at the end of the local code block where the local variable is located if the reference relationship indicates that the heap object is only referenced by the local variable; for example, it executes steps S340-S350 in the above method embodiment.
[0268] The processing module 1120 is used to convert the third code of the inserted release code into the second code; for example, to perform step S360 in the above method embodiment.
[0269] Optionally, the first code is a UDF; or, the first code is an IR converted from a UDF.
[0270] Optionally, the code optimization device also includes: trigger module 1130.
[0271] The trigger module 1130 is used to trigger the native execution engine to execute the second code; for example, to execute step S130 in the above method embodiment.
[0272] This application provides a computing device including a memory and at least one processor connected to the memory. The memory is used to store computer program code, which includes computer instructions. When the computer instructions are executed by the at least one processor, the computing device performs the method described in the above embodiments.
[0273] This application also provides a computing device cluster. The computing device cluster includes at least one computing device, which can be a desktop computer, laptop computer, or smartphone, among other terminal devices.
[0274] This application also provides a computing device cluster, which includes at least one computing device, each computing device including a processor and a memory; the processor of the at least one computing device is used to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster executes a code optimization method.
[0275] This application also provides a computer program product containing instructions. The computer program product may be a software or program product containing instructions, capable of running on a computing device or stored on any usable medium. When the computer program product is run on at least one computing device, it causes the at least one computing device to execute a code optimization method.
[0276] This application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that a computing device can store, or a data storage device such as a data center that includes one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive). The computer-readable storage medium includes instructions that instruct a computing device to execute a code optimization method.
[0277] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the protection scope of the technical solutions of the embodiments of the present invention.
Claims
1. A code optimization method, characterized in that, The method includes: Obtain the first code; the first code is the binary code obtained by compiling the original code, and the original code is code based on the first programming language; The first code is converted into second code, which is code based on a second programming language with the same function as the first code; wherein the code in the second programming language has higher running efficiency on the processor than the code in the first programming language; the second programming language is a programming language supported by the native execution engine, which runs on the processor.
2. The method according to claim 1, characterized in that, The step of converting the first code into the second code includes: The dependency objects on which the first code depends are converted into dependency objects with the same functionality written in the second programming language to obtain the third code; wherein, the types of the dependency objects include: interfaces and / or classes; Based on the syntax rules of the second programming language, the third code is converted into the second code.
3. The method according to claim 2, characterized in that, The step of converting the dependency objects on which the first code depends into dependency objects with the same functionality written in the second programming language to obtain the third code includes: Obtain the identifier of the first dependent object that the first code depends on; Based on the identifier of the first dependent object, a second dependent object corresponding to the first dependent object is determined from the first correspondence; the second dependent object is written in the second programming language; wherein, the first dependent object and the second dependent object have the same function; the first correspondence is used to indicate: the correspondence between a dependent object written in the first programming language and a dependent object written in the second programming language with the same function; The first dependent object that the first code depends on is replaced with the second dependent object to obtain the third code.
4. The method according to claim 3, characterized in that, The first dependency object has multiple candidate dependency objects, and the different candidate dependency objects have different running efficiencies on the processor; wherein, the candidate dependency objects are dependency objects with the same function as the first dependency object and written in the second programming language; The step of determining the second dependent object corresponding to the first dependent object from the first correspondence based on the identifier of the first dependent object includes: Based on the identifier of the first dependent object, determine the plurality of candidate dependent objects corresponding to the first dependent object from the first correspondence; Candidate dependency objects that meet the operating efficiency requirements on the processor are identified as the second dependency object.
5. The method according to any one of claims 2-4, characterized in that, The conversion of the third code into the second code based on the syntax rules of the second programming language includes: Based on the syntax rules of the third code and the syntax rules of the second programming language, code blocks in the third code are converted into code blocks of the second programming language with the same function to obtain the second code.
6. The method according to any one of claims 2-5, characterized in that, Before converting the third code into the second code based on the syntax rules of the second programming language, the method further includes: The memory allocation code in the third code is converted into a target allocation code; wherein the target allocation code is used to determine a free storage unit in the storage unit pool that meets the allocation conditions; the storage unit pool includes at least one storage unit; The step of converting the third code into the second code includes: Convert the transformed third code into the second code.
7. The method according to any one of claims 2-6, characterized in that, Before converting the third code into the second code based on the syntax rules of the second programming language, the method further includes: Determine the reference relationships of heap objects in the third code; If the reference relationship indicates that the heap object is only referenced by local variables, then at the end of the local code block where the local variable is located, release code for releasing the heap object is inserted; The step of converting the third code into the second code includes: The third code inserted into the release code is converted into the second code.
8. The method according to any one of claims 1-7, characterized in that, The first code is a user-defined function (UDF); Alternatively, the first code is an intermediate representation IR converted from a user-defined function (UDF).
9. The method according to any one of claims 1-8, characterized in that, The method further includes: This triggers the native execution engine to execute the second code.
10. A code optimization device, characterized in that, The device includes: an acquisition module and a processing module; The acquisition module is used to acquire first code; the first code is binary code obtained by compiling the original code, and the original code is code based on a first programming language; The processing module is used to convert the first code into second code, the second code being code with the same function as the first code based on a second programming language; wherein the code in the second programming language has higher running efficiency on the processor than the code in the first programming language; the second programming language is a programming language supported by a native execution engine, and the native execution engine runs on the processor.
11. The apparatus according to claim 10, characterized in that, The processing module is used to convert the dependency objects on which the first code depends into dependency objects with the same function written in the second programming language, to obtain the third code; wherein, the types of the dependency objects include: interfaces and / or classes; The processing module is used to convert the third code into the second code based on the syntax rules of the second programming language.
12. The apparatus according to claim 11, characterized in that, The acquisition module is used to obtain the identifier of the first dependent object that the first code depends on from the first code; The processing module is configured to determine, based on the identifier of the first dependent object, a second dependent object corresponding to the first dependent object from a first correspondence; the second dependent object is written in the second programming language; wherein the first dependent object and the second dependent object have the same function; the first correspondence is used to indicate the correspondence between a dependent object written in the first programming language and a dependent object written in the second programming language with the same function; The processing module is used to replace the first dependency object that the first code depends on with the second dependency object to obtain the third code.
13. The apparatus according to claim 12, characterized in that, The first dependency object has multiple candidate dependency objects, and the different candidate dependency objects have different running efficiencies on the processor; wherein, the candidate dependency objects are dependency objects with the same function as the first dependency object and written in the second programming language; The processing module is configured to determine, based on the identifier of the first dependent object, the plurality of candidate dependent objects corresponding to the first dependent object from the first correspondence; wherein, different candidate dependent objects have different running efficiencies on the processor; the candidate dependent objects are dependent objects with the same function as the first dependent object and written in the second programming language. The processing module is used to determine the candidate dependency objects whose running efficiency on the processor meets the conditions as the second dependency objects.
14. The apparatus according to any one of claims 11-13, characterized in that, The processing module is used to convert the memory allocation code in the third code into a target allocation code; wherein the target allocation code is used to determine a free storage unit in the storage unit pool that meets the allocation conditions; the storage unit pool includes at least one storage unit; The processing module is specifically used to convert the converted third code into the second code.
15. The apparatus according to any one of claims 11-14, characterized in that, The processing module is used to determine the reference relationships of heap objects in the third code; The processing module is configured to, if the reference relationship indicates that the heap object is only referenced by local variables, insert release code for releasing the heap object at the end of the local code block where the local variables are located; The processing module is specifically used to convert the third code inserted into the release code into the second code.
16. The apparatus according to any one of claims 10-15, characterized in that, The code optimization device further includes: a triggering module; The processing module is used to trigger the native execution engine to execute the second code.
17. A computing device, characterized in that, The device includes a memory and a processor, the memory being coupled to the processor; the memory is used to store computer program code, the computer program code including computer instructions; when the computer instructions are executed by the processor, the computing device causes the device to perform the method as described in any one of claims 1 to 9.
18. A computer program product, characterized in that, When the computer program product is run on a computer, it causes the computer to perform the method as described in any one of claims 1 to 9.