A code processing method, device and storage medium

By decompiling the source platform's software code into intermediate representations and compiling it into target platform code, the problem of high difficulty in building the ARMv8 platform software ecosystem is solved, and efficient code porting and development efficiency are achieved.

CN114253554BActive Publication Date: 2025-05-30HUAWEI TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202011066288.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-09-21
Filing Date
2020-09-30
Publication Date
2025-05-30
Estimated Expiration
2040-09-30

AI Technical Summary

Technical Problem

It is difficult to build a software ecosystem for ARMv8 platform or other platforms, and it is difficult for existing technologies to effectively reduce this difficulty.

Method used

Provides a code processing method, which realizes the porting of code by decompiling the software code of the source platform into an intermediate representation (IR) and compiling it according to the target platform. This method requires no developer involvement, reducing the possibility of developers getting in touch with software code.

Benefits of technology

The ability to port software code on the source platform to the target platform is realized, reducing the difficulty of building a platform software ecosystem, and improving the efficiency of software development.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114253554B_ABST
    Figure CN114253554B_ABST
Patent Text Reader

Abstract

The present application discloses a code processing method, apparatus and storage medium, including: obtaining first code based on a low-level language applied to a source platform; decompiling the obtained first code to obtain an intermediate representation IR; and then compiling the IR into second code based on a low-level language applied to a first target platform, wherein the source platform and the target platform have different instruction sets. For example, code applicable to the x86 platform can be converted into code applicable to the ARM platform, without the need for technicians to implement cross-platform migration of software code by manually writing program code. In this way, the software code of the source platform can be transplanted to the target platform for running, thereby reducing the difficulty of building the software ecosystem of the first target platform.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims the priority of a Chinese patent application with the application number 202010996387.6 and the application title "A Code Processing Method, Device and Storage Medium", which was filed with the China National Intellectual Property Administration on September 21, 2020, and the entire content of which is incorporated herein by reference. Technical Field

[0002] Embodiments of this application relate to the technical field of code processing, and in particular, to a code processing method, device and storage medium. Background Art

[0003] The x86 platform is a general computing platform mainly developed by Intel Corporation, specifically a general reference to a series of central processing unit instruction set architectures based on the Intel 8086 and backward compatible. Since its introduction in 1978, the x86 platform has evolved into a large and complex instruction set after years of accumulation.

[0004] In practical applications, developers usually develop various software that can be applied to the x86 platform based on the x86 platform, thus building a large software ecosystem. Therefore, many current application software may only be applicable to the x86 platform and cannot be applicable to other platforms, such as the ARMv8 platform (a processor architecture supporting 64-bit instruction sets released by ARM Corporation), etc.

[0005] When building the software ecosystem of the ARMv8 platform or other platforms, it is usually the technical personnel who write software code according to the code rules of the platform (such as instruction sets, etc.) so that the developed software can be applicable to the platform. However, the difficulty of developing new software for this platform is usually relatively high, and the software development efficiency is relatively slow, which makes it difficult to build the software ecosystem of the ARMv8 platform or other platforms. For this reason, there is an urgent need for a method that can reduce the difficulty of building the platform software ecosystem. Summary of the Invention

[0006] Embodiments of this application provide a code processing method, device and storage medium to reduce the difficulty of building the software ecosystem of the platform.

[0007] In a first aspect, an embodiment of the present application provides a code processing method. By transplanting the software code of the source platform to the first target platform, the difficulty of building the software ecosystem of the platform is reduced. The source platform and the first target platform belong to different platforms, specifically, they may have different instruction sets. When specifically implemented, the first code based on a low-level language applied to the source platform can be obtained first. The first code can be, for example, code based on assembly language or machine language and can be recognized by the source platform. Then, the obtained first code can be decompiled to obtain a first intermediate representation (IR). The first IR can be an IR related to the first target platform or an IR unrelated to the first target platform. Next, the first IR can be compiled to obtain code based on a low-level language applied to the first target platform, and the obtained code can be recognized and run by the first target platform, thereby realizing the transplantation of the software code on the source platform to the first target platform.

[0008] At the same time, the processes of decompiling and compiling the software code do not require the participation of developers, so the isolation between developers and the software code can be achieved, and the possibility of developers contacting the software code is reduced. For software operators, they can optimize and re-develop the software code transplanted to the first target platform, which is convenient for software operators to maintain the software code transplanted to the first target platform.

[0009] Among them, the source platform can be, for example, an x86 platform, and the first target platform can be an ARM platform, specifically an ARMv8 platform, etc. Of course, in practical applications, the source platform can be any platform, and the first target platform can be any platform different from the source platform.

[0010] The above method can be applied locally or in the cloud. When applied locally, it can be specifically applied to local terminals or servers, etc. When applied in the cloud, it can be specifically presented to users in the form of cloud services.

[0011] In a possible implementation, the first code of the source platform can be ported to any target platform. Specifically, taking the porting to the first target platform and the second target platform as an example, in addition to obtaining the low-level language-based code applied to the first target platform through the above decompilation and compilation processes, it can also be that during the decompilation process, the IR corresponding to the second platform is obtained according to the first code. The IR corresponding to the second target platform is different from the first IR, and the applicable target platforms are different. The first target platform and the second target platform have different instruction sets, and moreover, the second target platform and the source platform also have different instruction sets. That is, when porting the software code on the source platform to any platform, the above decompilation and compilation processes can be used for implementation.

[0012] In a possible implementation, a target platform selection interface can be presented to the user. The target platform selection interface can present candidates of multiple target platforms for the user to select the target platform. Then, in response to the user's selection operation for the target platform, the target platform selected by the user can be determined from the multiple target platforms. For example, the above first target platform can be determined according to the selection operation performed by the user on the target platform selection interface.

[0013] In a possible implementation, during the decompilation of the first code, the annotation information of the first code can be obtained first. The annotation information can include, for example, any one or more of the types, quantities, and jump address types of the parameters in the first code. In this way, during decompilation, the first code can be decompiled according to the annotation information to obtain the first IR. Taking the parameter type in the annotation information as an example, assuming that the length of the parameter in the first code is 64 bits, then in the first IR obtained by decompiling the first code, the length of the parameter can still be 64 bits (such as floating-point type) and will not change to 32 bits (such as integer type). In this way, the type of the parameter can be kept consistent before and after decompilation, reducing the possibility of logical errors in the first IR.

[0014] In a possible implementation, during the decompilation of the first code, the initially obtained IR can also be optimized. Specifically, after decompiling the first code, a second IR can be obtained. Generally, the second IR may have a certain optimization space. For example, the data reading method in the generated second IR is to read 64-bit data each time, while the first target platform has the ability to read 128-bit data at one time. The data reading method in the second IR can be optimized to read 128-bit data each time. When specifically implementing, the second IR can be optimized according to the first target platform, such as according to the hardware / logical capabilities of the first target platform, so as to obtain the above first IR. In this way, more efficient code can be obtained after compiling the first IR subsequently.

[0015] In a possible implementation, prompt information can be further generated and presented. The prompt information can be used to prompt for items to be checked, where the items to be checked can be generated based on the differences between the first target platform and the source platform. For example, the prompt information can specifically be the instructions of the source platform displayed in a highlighted manner and the instructions of the first target platform with the same semantics as the instructions, so that the user can determine whether the instructions of the first target platform translated from the instructions of the source platform are accurate based on the prompt information.

[0016] In a possible implementation, when the above method is applied to the cloud, the user can provide the first code to the cloud. For example, the user can send a code processing request containing the first code to the cloud through a terminal or a client, etc., so that the cloud can obtain the first code. Correspondingly, after performing corresponding decompilation and compilation processing on the received first code, the cloud can send the generated low-level language-based code applied to the first target platform to the user, so that the user can obtain the required code.

[0017] In a possible implementation, when decompiling the first code, specifically, the first code can be decompiled according to the instruction semantic library corresponding to the source platform. For example, the instruction semantic library can contain the semantics of multiple instructions of the source platform. In this way, during the decompilation process, each instruction in the first code can be traversed, and the semantics of the instruction can be identified according to the instruction semantic library, so that the first code can be decompiled into the corresponding compiler IR according to the identified instruction semantics.

[0018] In a possible implementation, the instruction semantic library can also be modified by the user. For example, the user can add single instruction multiple data (SIMD) instructions, etc. to the instruction semantic library to be used to identify the SIMD instructions in the first code. Thus, during the decompilation process, corresponding decompilation processing can be performed on the SIMD instructions in the first code, so that the first IR contains instructions with vectorized semantics. Correspondingly, the terminal or the server can respond to the user's modification operation on the instruction semantic library and perform corresponding modification on the instruction semantic library.

[0019] In a possible implementation, in an embedded assembly scenario, the software code to be ported may include both first code based on a low-level language and variables based on a high-level language. Therefore, when the terminal or server obtains the first code, it also obtains the variables based on the high-level language. Thus, when decompiling the first code, the first code can be translated into a first IR including functions, and the functions include formal parameters, and the actual parameters corresponding to the formal parameters can be variables. That is, during the decompilation process, variables based on the high-level language can be passed as actual parameters to the formal parameters in the functions.

[0020] In a possible implementation, when translating the first code into a first IR including functions, specifically, it can be to determine the semantics of each instruction string in the first code, so that according to the correspondence between the semantics and the functions, the functions corresponding to the semantics of each instruction string in the first code can be determined, and then a first IR including the functions can be generated.

[0021] In a possible implementation, before decompiling the first code, the variables in the first code can also be relocated. In this way, after decompiling the first code, each variable in the obtained first IR can have different logical addresses. Taking the first variable and the second variable included in the first IR as an example, the first variable can have a first logical address, and the second variable can have a second logical address, and the first logical address and the second logical address are different logical addresses. In specific implementation, before decompiling the first code, a preset first logical address can be configured for the first variable in the first code, and a preset second logical address can be configured for the second variable in the first code, where both the first logical address and the second logical address can be abstract logical addresses.

[0022] In a possible implementation, since there may be differences in function call conventions or SIMD instructions between the source platform and the first target platform, when decompiling the first code, specifically, the first code can be decompiled according to the function call convention of the first target platform or the SIMD instructions of the first target platform, so that the finally obtained code can meet the function call convention or SIMD instruction requirements of the first target platform. In practical applications, specifically, decompilation processing can be performed according to the differences in function call conventions or SIMD instructions between the first target platform and the source platform.

[0023] In a possible implementation, when the first code includes SIMD instructions, the decompilation of the first code may be performed by means of direct vectorization. Specifically, a third code based on a low-level language for the target platform may be generated first, and the third code can be used to describe the vectorization semantics of the SIMD instructions in the first code; Exemplarily, the third code may be, for example, a code including intrinsic functions. Then, the third code may be decompiled to obtain a first IR corresponding to the SIMD instructions of the first target platform with vectorization marks. At this time, the first IR is associated with the first target platform.

[0024] In a possible implementation, when the first code includes SIMD instructions, the decompilation of the first code may be performed by means of indirect vectorization. In specific implementation, a fourth code based on a high-level language may be generated first, and the fourth code can be used to describe the vectorization semantics of the SIMD instructions in the first code; Then, the fourth code may be compiled to obtain a first IR with vectorization marks. Then, at the compilation stage, the IR with vectorization marks may be automatically vectorized and compiled to generate a code based on a low-level language applied to the first target platform, and the code may include SIMD instructions of the first target platform. In this way, the indirect vectorization processing of the SIMD instructions in the first code is realized.

[0025] In a second aspect, the embodiments of the present application further provide a code processing method. During the process of implementing software code transplantation, a first code based on a low-level language applied to a source platform may be obtained, and then, a second code may be output according to the first code, where the second code is a code based on a low-level language that can be applied to a first target platform, and the second code is obtained by processing the acquired first code, for example, it may be obtained by first decompiling the first code and then compiling it, and the source platform and the first target platform have different instruction sets. In this way, the second code applied to the first target platform can be obtained according to the first code of the source platform, and the obtained second code can be recognized and run by the first target platform, thereby realizing the transplantation of the software code on the source platform to the first target platform.

[0026] At the same time, the above process of processing the first code may not require the participation of developers, thereby realizing the isolation between developers and software code and reducing the possibility of developers contacting software code. For software operators, they can optimize and re-develop according to the software code transplanted to the first target platform, which is convenient for software operators to maintain the software code transplanted to the first target platform.

[0027] Among them, the source platform can be, for example, an x86 platform, and the first target platform can be an ARM platform, specifically an ARMv8 platform or the like. Of course, in actual applications, the source platform can be any platform, and the first target platform can be any platform different from the source platform.

[0028] In a possible implementation manner, when outputting the second code, specifically, the second code can be presented through a code display interface. In this way, the user can view the translated code on the code display interface, and thus can perform operations such as secondary development and optimization based on the translated code.

[0029] In a possible implementation manner, the process of processing the first code can be applied to the cloud. When obtaining the first code, specifically, the first code received from the user can be received. For example, the user can send a code processing request to the cloud through a terminal or a client, etc. The code processing request can carry the first code, etc. Of course, the user can also provide the first code to the cloud through other means. In this way, after the cloud processes the first code and obtains the second code, when outputting the second code, specifically, the second code can be output to the user. For example, it can be presented on the code display interface of the terminal used by the user, etc.

[0030] In a possible implementation manner, before processing the first code, a target platform selection interface can also be presented. The target platform selection interface can present candidates of multiple target platforms for the user to select the target platform. Then, in response to the user's selection operation for the target platform, the target platform selected by the user can be determined from the multiple target platforms. For example, the above-mentioned first target platform can be determined according to the selection operation performed by the user on the target platform selection interface.

[0031] In a possible implementation manner, when processing the first code, specifically, the instruction semantic library corresponding to the source platform can be obtained first, and the first code can be processed using the instruction semantic library. For example, the instruction semantic library can contain the semantics of multiple instructions of the source platform. In this way, when processing the first code, each instruction in the first code can be traversed, and the semantics of the instruction can be identified according to the instruction semantic library, so that the first code can be decompiled into the corresponding compiler IR according to the identified instruction semantics.

[0032] In a possible implementation, the user can also modify the instruction semantic library corresponding to the source platform. For example, the user can add SIMD instructions to the instruction semantic library to identify SIMD instructions in the first code, etc. Then, in response to the modification operation performed by the user on the instruction semantic library, the instruction semantic library corresponding to the source platform can be modified. In this way, during the decompilation of the first code, decompilation can be performed according to the modified instruction semantic library.

[0033] In a possible implementation, after obtaining the second code from the first code, a prompt message can also be generated and presented. The prompt message is used to prompt for items to be checked, where the items to be checked are generated based on the differences between the first target platform and the source platform. For example, the prompt message can specifically be the instructions of the source platform and the instructions of the first target platform with the same semantics displayed in a highlighted manner, so that the user can determine whether the instructions of the first target platform translated from the instructions of the source platform are accurate based on the prompt message.

[0034] In a possible implementation, not only can the second code be output, but also the first IR obtained by decompiling the first code can be presented. Correspondingly, the output second code is obtained by compiling the first IR. In this way, the user can perform operations such as debugging and observing on the presented first IR and conduct corresponding analysis.

[0035] In a possible implementation, a second IR can also be presented. The second IR is obtained by decompiling the first code. Specifically, during the decompilation process, the second IR can be first obtained by decompiling the first code, and then the second IR can be optimized to obtain the first IR. In this way, the first IR obtained after the optimization process can be more efficient during the code execution stage. For example, the data reading method in the generated second IR is to read 64-bit data each time, while the first target platform has the ability to read 128-bit data at once. The data reading method in the second IR can be optimized to read 128-bit data each time, so that when reading the same amount of data, the data reading operation does not need to be performed twice.

[0036] In a possible implementation, the output first IR can also be modified by the user. For example, when the user determines that there are logical errors or optimizable code in the output first IR, the first IR can be modified. Then, the terminal or the server, in response to the modification operation on the first IR, can obtain the modified first IR, and thus can compile according to the modified first IR to obtain a relatively better third code. The third code is a low-level language-based code applied to the first target platform, and thus the third code can be presented to the user.

[0037] In a possible implementation, when generating low-level language-based code applied to a second target platform according to the first code, when decompiling the first code, the first code can be decompiled into a third IR, and subsequently, the low-level language code applied to the second target platform can be obtained by compiling the third IR. Among them, the IRs corresponding to different target platforms can be different, so that different IRs and codes can be generated for different target platforms.

[0038] In a possible implementation, the user can also modify the output second code. For example, the user can perform secondary development and optimization on the basis of the output second code. Correspondingly, by responding to the user's modification operation on the second code, the modified second code can be obtained, and at the same time, the modified second code can be presented to the user in real time for the user to view.

[0039] In a possible implementation, when obtaining the first code, not only can the first code based on the low-level language applied to the source platform be obtained, but also variables based on the high-level language can be obtained at the same time. For example, in the embedded assembly scenario, it not only includes the assembly language code as the low-level language, but also can include variables of the high-level language, such as variables of the C / C++ language.

[0040] In a third aspect, based on the same inventive concept as the method embodiment of the first aspect, an embodiment of the present application provides a computing device. The device has the functions corresponding to the respective implementations of the above first aspect. The functions can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions.

[0041] In a fourth aspect, based on the same inventive concept as the method embodiment of the second aspect, an embodiment of the present application provides two computing devices. The device has the functions corresponding to the respective implementations of the above second aspect. The functions can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions.

[0042] In a fifth aspect, an embodiment of the present application provides a computing device, including: a processor and a memory; the memory is used to store instructions, and when the computing device runs, the processor executes the instructions stored in the memory so that the device executes the code processing method in the above first aspect or any implementation manner of the first aspect. It should be noted that the memory can be integrated into the processor or can be independent of the processor. The device may further include a bus. Among them, the processor is connected to the memory through the bus. Among them, the memory may include a readable memory and a random access memory.

[0043] Sixthly, an embodiment of the present application provides a computing device, including: a processor and a memory; the memory is used to store instructions, and when the computing device runs, the processor executes the instructions stored in the memory, so that the device executes the code processing method in the second aspect or any implementation manner of the second aspect. It should be noted that the memory may be integrated into the processor or independent of the processor. The device may further include a bus. Among them, the processor is connected to the memory through the bus. Among them, the memory may include a readable memory and a random access memory.

[0044] Seventhly, an embodiment of the present application further provides a readable storage medium, in which a program or instructions are stored, and when it runs on a computer, the code processing method in the first aspect or any implementation manner of the first aspect is executed.

[0045] Eighthly, an embodiment of the present application further provides a readable storage medium, in which a program or instructions are stored, and when it runs on a computer, the code processing method in the second aspect or any implementation manner of the second aspect is executed.

[0046] Ninthly, an embodiment of the present application further provides a computer program product containing instructions, and when it runs on a computer, the computer executes any code processing method in the first aspect or any implementation manner of the first aspect.

[0047] Tenthly, an embodiment of the present application further provides a computer program product containing instructions, and when it runs on a computer, the computer executes any code processing method in the second aspect or any implementation manner of the second aspect.

[0048] In addition, for the technical effects brought by any implementation manner in the third aspect to the tenth aspect, reference may be made to the technical effects brought by different implementation manners in the first aspect, or reference may be made to the technical effects brought by different implementation manners in the second aspect, which will not be elaborated here. Description of the Drawings

[0049] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments recorded in the present application. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings.

[0050] Figure 1 It is a schematic diagram of an exemplary system architecture in an embodiment of the present application;

[0051] Figure 2It is a schematic flowchart of a code processing method in an embodiment of the present application;

[0052] Figure 3 It is a schematic diagram of code in an embedded assembly scenario in an embodiment of the present application;

[0053] Figure 4 It is a schematic flowchart of a parameterized translation in an embodiment of the present application;

[0054] Figure 5 It is a schematic diagram for adjusting parameters of a function in an embodiment of the present application;

[0055] Figure 6 It is a schematic diagram of IR before and after optimization in an embodiment of the present application;

[0056] Figure 7 It is a schematic diagram of a target platform selection interface in an embodiment of the present application;

[0057] Figure 8 It is a schematic diagram of a user 801 interacting with a computing device 802 in an embodiment of the present application;

[0058] Figure 9 It is a schematic diagram of the structure of a computing device in an embodiment of the present application;

[0059] Figure 10 It is a schematic diagram of the structure of another computing device in an embodiment of the present application;

[0060] Figure 11 It is a schematic diagram of the hardware structure of a computing device in an embodiment of the present application;

[0061] Figure 12 It is a schematic diagram of the hardware structure of yet another computing device in an embodiment of the present application. Detailed implementation manners

[0062] When building the software ecosystem of a platform, usually technicians write software that can run on this platform. However, this way of manually writing software program code not only has low software development efficiency, but also various program errors are likely to occur when writing the program code, which makes the development of software more difficult, and thus makes the construction of the platform software ecosystem more difficult.

[0063] To this end, the embodiments of the present application provide a code processing method, which may be to transplant software on other platforms to this platform, so as to enrich the software that can successfully run on this platform, thereby reducing the difficulty of building the software ecosystem of this platform. Specifically, during implementation, the software code of the source platform (i.e., the above-mentioned other platform) can be decompiled to obtain the internal representation (intermediate representation, IR) of the compiler, and then the IR can be compiled into code based on a low-level language applied to the target platform (the above-mentioned this platform), so that the code can successfully run on the target platform. In this way, it is possible to transplant the software of the source platform to the target platform for operation. Of course, the source platform and the target platform are different platforms, and the difference between these two platforms is at least that they have different instruction sets.

[0064] Moreover, the process of decompiling and compiling the software code does not require the participation of developers, so that the isolation between developers and the software code can be achieved, reducing the possibility of developers contacting the software code. For software operators, they can optimize and re-develop the software code transplanted to the target platform, which is convenient for software operators to maintain the software code transplanted to the target platform.

[0065] As an example, the above code processing method can be applied to Figure 1 the system architecture shown. As Figure 1 shown, the system architecture 100 includes a decompilation module 101 and a compilation module 102. For the code 1 applied to the source platform 103, the decompilation module 101 decompiles it to obtain the IR, and then passes the decompiled IR to the compilation module 102, and the compilation module 102 compiles the IR into code 2 based on a low-level language applied to the target platform 104. In this way, the obtained code 2 can run on the target platform 104.

[0066] In practical applications, the decompilation module 101 can be a software-based functional module; or it can also be implemented by a device with decompilation capabilities, such as a decompiler. Similarly, the compilation module 102 can be implemented by a device with compilation capabilities, such as a compiler. The decompilation module 101 and the compilation module 102 can be deployed on the target platform 104, or can also be deployed on the terminal 105 or the server 106, etc. Exemplarily, when the decompilation module 101 and the compilation module 102 are deployed on the terminal 105 and the server 106, the decompilation module 101 can be deployed on the server 106, while the compilation module is deployed on the terminal 105, etc.; or, both the decompilation module 101 and the compilation module 102 can be deployed in the server 106 located in the cloud. At this time, during the software transplantation process, the terminal 105 can send software code to the cloud (specifically, it can be the server 106 in the cloud), for example, it can send a code processing request to the cloud, and the code processing request includes the software code to be transplanted; then, after the above-mentioned decompilation and compilation processing by the cloud (server 106), code based on a low-level language applicable to the target platform can be obtained, and this code is sent to the terminal 105. In this way, the terminal 105 can obtain software code that can be applied to the target platform, realizing software transplantation. Of course, in practical applications, the server 106 can also be a local server.

[0067] To make the above objects, features, and advantages of the present application more obvious and understandable, the following will exemplarily illustrate various non-limiting embodiments in the embodiments of the present application with reference to the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.

[0068] As Figure 2 shown, it is a schematic flowchart of a code processing method in an embodiment of the present application. This method can be applied to the above-mentioned target platform 104, or can be applied to a computing device, such as the above-mentioned terminal 105 or server 106. And this computing device can include a decompilation module 101 and a compilation module 102. Specifically, this method can include:

[0069] S201: Obtain the first code based on a low-level language applied to the source platform.

[0070] In practical applications, when software runs on a source platform, the program code of the software is usually compiled into code in a low-level language that can be directly recognized by the source platform. Among them, a low-level language refers to a programming language or instruction code that a machine can directly recognize, and can specifically be an assembly language or a machine language, etc. A machine language is a language represented by binary codes. A machine language is the only language that a computer can recognize and execute. An assembly language is a language that uses names and symbols that are easy to understand and remember to represent the operation codes in machine instructions in order to solve the disadvantages of machine language being difficult to understand and remember. An assembly language uses symbols to replace the binary codes of machine language, so an assembly language is essentially a symbolic language.

[0071] Relative to low-level languages, there is also a high-level language that is independent of the machine and is process- or object-oriented. High-level programming languages are usually close to natural languages and can use mathematical expressions, so they have stronger expressive power, can conveniently represent data operations and program control structures, and can better describe various algorithms, such as languages like C, C++, and Java. High-level languages can be applicable to different platforms, such as platforms with the x86 instruction set architecture (a general computing platform mainly developed by Intel Corporation, hereinafter referred to as the x86 platform), and can also be applicable to platforms with the advanced RISC machines (ARM) architecture (hereinafter referred to as the ARM platform), a platform with the Performance Optimization With Enhanced RISC–Performance Computing architecture, etc. Usually, high-level programming languages cannot be directly recognized and executed by machines. Developers can compile code files based on high-level languages through a compiler program such as a compiler so that they can be recognized and executed by machines.

[0072] In this embodiment, the obtained first code can be code based on a low-level language, and this code can be applicable to the source platform. For example, the obtained first code can be a code in the ".obj" format file or other format files obtained by assembling based on the assembly language corresponding to the source platform.

[0073] Or, in other possible implementation manners, when obtaining the first code, variables based on a high-level language can be obtained simultaneously. For example, in Figure 3In the embedded assembly scenario shown, the first code obtained is assembly language code in the format of "movdqa % % xmm4, 16(%0)", etc., and the variables based on the high-level language are specifically "&ff_inverse[0]" and "&ff_inverse[8]" in ""r"(&ff_inverse[0]), "r"(&ff_inverse[8])". Among them, embedded assembly is usually an encoding method adopted to improve the code execution efficiency and execute the proprietary instructions of the processor. Then, after assembling the embedded assembly code, the first code in the format of a file such as ".obj" can be obtained.

[0074] Exemplarily, the first code in this embodiment can be the code in the "obj" format obtained after assembling the entire assembly language program code of the software on the source platform; or the code in the "obj" format obtained after assembling part of the program code, such as one of the multiple code files of the software, etc.; or, it can also be a section of code in a code file, such as the code obtained by compiling the Figure 3 code block shown, etc., that is, in this embodiment, local decompilation processing can be performed on the software code.

[0075] In this embodiment, the source platform to which the first code is applied and the target platform to which the first code is transplanted are not the same, and specifically, they can have different instruction sets. The source platform and the target platform can belong to different types of platforms. For example, the source platform can be the x86 platform, and the target platform can be the ARMv8 platform (a processor architecture supporting the 64-bit instruction set released by ARM); or, the source platform and the target platform can also be two different platforms of the same type. The source platform can be Pentium II under the x86 platform, and the target platform can be Pentium III under the x86 platform (introducing the new SSE instruction set), or the source platform can be the x86 platform supporting 32 bits, and the target platform can be the x86-64 platform supporting 64 bits, etc.

[0076] S202: Decompile the obtained first code to obtain the first IR.

[0077] In this embodiment, during the process of transplanting the first code applied to the source platform to run on the target platform, the decompilation module 101 can first decompile the first code. Specifically, it can perform lexical analysis, syntax analysis, and semantic analysis on the input first code and convert it into a compiler IR, and this compiler IR is the decompilation result corresponding to the first code. Among them, the compiler IR can represent the semantics and syntax structure of the first code, and it can be regarded as another high-level language.

[0078] As an example, only the first code based on a low-level language can be obtained. In this case, the decompilation module 101 can decompile the first code into compiler IR. In other examples, it can also be that while obtaining the first code based on a low-level language, variables based on a high-level language are also obtained. In this case, when the decompilation module 101 decompiles the first code, parameterized translation processing also needs to be performed on the variables of the high-level language to avoid loss of variable information in the high-level language in the first code.

[0079] Specifically, refer to the Figure 4 parameterized translation process shown as follows:

[0080] S401: For the instruction string based on a low-level language in the mixed code block, the decompilation module 101 can translate the instruction string into a first compiler IR including functions according to the semantics of the instruction string. The semantics expressed by the functions included in the first IR are consistent with the semantics of the instruction string. For example, when the instruction string is "ADD %x %y", the function included in the compiler IR obtained by translating the instruction string can be a function for summation.

[0081] In an exemplary implementation manner of translating an instruction string, the decompilation module 101 can pre-obtain the correspondence between instruction semantics and functions. For example, it can be pre-saved in the decompilation module 101. Then, the decompilation module 101 can determine the semantics of the instruction string based on a low-level language in the mixed code block. The instruction string can include one or more instructions, so the decompilation module can determine the semantics of each instruction in the instruction string. Then, the decompilation module 201 determines the functions corresponding to the semantics of each instruction in the instruction string by looking up the correspondence between semantics and functions, and thus can further generate a first IR including the function according to the determined functions.

[0082] As an example, the decompilation module 101 can determine the semantics of each instruction in the instruction string by looking up the instruction semantics library corresponding to the source platform (such as pre-semantically marking the instructions of the source platform by technicians and importing the instructions and instruction semantics into the decompilation module 101, etc.). The instruction semantics library includes multiple instructions of the source platform, and the semantics of each instruction have been pre-completed semantic annotation. When instruction A in the first code has the same syntax structure as instruction a in the instruction semantics library, it can be determined that the semantics of instruction a in the instruction semantics library are the semantics of instruction A in the first code. In practical applications, other methods can also be used to determine the semantics of each instruction, and this embodiment does not limit this.

[0083] S402: When the decompilation module 101 determines the function corresponding to the instruction string, it can create virtual registers as parameters for the function.

[0084] S403: The decompilation module 101 references the virtual register created during the translation process for the function determined in the translation, and uses this virtual register as the formal parameter of the function.

[0085] S404: The decompilation module 101 establishes the correspondence between the variables in the high-level language and the formal parameters in the function, so as to pass the variables in the high-level language in the hybrid code as actual parameters to the formal parameters in the semantic function.

[0086] Exemplarily, the decompilation module 101 can count the variables in the high-level language in the hybrid code block, so as to obtain a list composed of multiple variables in the high-level language, and then correspond the variables in the list with the formal parameters in the function according to the position of the formal parameter in the IR and the calling convention of the compiler. For example, the first variable in the list can be corresponded to the first formal parameter of the first function in the IR, the second variable in the list can be corresponded to the second formal parameter of the first function in the IR, etc.

[0087] In this way, during the decompilation process, the variable information in the high-level language can be retained as actual parameters in the function without information loss, so as to ensure that the code information before and after the decompilation process is consistent in the embedded assembly scenario.

[0088] In practical applications, the decompilation module 101 can first distinguish whether the code to be decompiled only includes the first code based on the low-level language, or includes both the first code and the high-level language variables. In a possible implementation manner, the decompilation module 101 can first detect the compilation command of the code to be decompiled and the file type of the code to determine whether the first code is a code based on the high-level language or a code based on the low-level language. Further, when it is determined that the code is a code based on the low-level language according to the compilation command and the file type, the decompilation module 101 can further determine whether the code is the first code entirely based on the low-level language or a hybrid code block that includes both the first code in the low-level language and the high-level language variables, such as the embedded assembly code composed of C / C++ language and assembly language (low-level language).

[0089] Further, if the first code includes a hybrid code block of the second code based on the low-level language and the variables based on the high-level language, during the decompilation process of the first code by the decompilation module 101, for the high-level language variables in the first code, the above parameterized translation process can be performed.

[0090] In this embodiment, the first code usually includes at least one function call. The function being called is essentially a variable. For example, in the function y = a + b * c, the function y is essentially a variable. In some possible implementation manners, other variables are also associated in the function body of the function. For example, the variables a, b, and c are also associated with the function y. Based on this, the variables associated with the function include the function itself and the variables associated in the function body. For example, the variables associated with the function y include y and a, b, and c. Since in the first code, the addresses of the variables associated with the function are relative addresses, that is, the addresses of each variable are uncertain. Therefore, before translating the instructions in the first code into the first IR of the compiler based on the instruction semantic library corresponding to the source platform, the decompilation module 101 can also relocate the variables (different from the variables in the high-level language) in the IR to determine the absolute addresses of each variable in the first code.

[0091] Exemplarily, taking the first code including a first variable and a second variable as an example, before decompiling, the decompilation module 101 can configure a preset first logical address for the first variable in the first code, and configure a preset second logical address for the second variable in the first code. Moreover, the first logical address and the second logical address can be different. In specific implementation, the decompilation module can access a relocation table, which stores the logical address information of multiple variables (including the first variable and the second variable). The decompilation module 101 determines the logical addresses (that is, the absolute addresses corresponding to the variables) corresponding to the first variable and the second variable respectively according to the logical address information in the redirection table, and associates the logical address with the symbol of the variable. Among them, the address information configured for the first variable and the second variable can be false logical address information. Then, the decompilation module 101 can decompile the first code with the variable address configured by using the instruction semantic library corresponding to the source platform to obtain the first IR of the compiler. In this way, in the subsequent compilation stage, the logical address of the variable will be recompiled into the relocation information of the target platform, that is, in the compilation stage, the variables in the first IR of the compiler can point to specific logical addresses.

[0092] In practical applications, in addition to different instruction sets between the source platform and the target platform, there may be other differences. For example, there may be differences in function call conventions or Single Instruction Multiple Data (SIMD) instructions between the source platform and the target platform. Therefore, in some embodiments, when decompiling the first code, the decompilation module 101 may first determine the differences between the source platform and the target platform, and decompile the first code according to the differences. Among them, the differences between the source platform and the target platform may be determined by technicians in advance by comparing the function call conventions or SIMD instructions of the source platform and the target platform, and then imported into the decompilation module 101. This embodiment does not limit this.

[0093] In other possible embodiments, the decompilation module 101 may also directly decompile the first code according to the function call convention or SIMD instruction of the target platform, etc. This embodiment does not limit this. In this way, through the above differential processing or directly decompiling based on the information of the target platform, the function calls or SIMD instructions in the obtained IR can conform to the function call convention and SIMD instruction of the target platform.

[0094] For ease of understanding, the following takes the adjustment of the parameters in the first code according to the difference in function call conventions between the target platform and the source platform as an example for illustrative description.

[0095] Specifically, the parameters of a function are stored in registers or in the stack of memory. The storage methods of parameters are different in different platforms. The decompilation module 101 can adjust the register information or stack information of the function in the source code block according to the differences in the function call rules of the source platform and the target platform. For example, the decompilation module 101 can adjust the parameters stored in registers and store them in the stack; or adjust the parameters stored in the stack and store them in registers.

[0096] Before adjusting the register information or stack information, the decompilation module 101 can first use a decoding tool, such as intel xed, to decode the source code block to obtain the instruction control flow. Then the decompilation module 101 can execute a data flow analysis algorithm for the instruction control flow to analyze the active registers and stack, so as to obtain the parameter types and quantities of the function in the source code block. Among them, the parameter type is mainly used to indicate whether the parameter is stored in a register or in the stack.

[0097] For ease of understanding, the processes of register analysis and stack analysis will be described in detail below.

[0098] First, the following several data sets are defined in this embodiment:

[0099] Use[n]: The set of variables used by n;

[0100] Def[n]: The set of variables defined by n;

[0101] In[n]: The variables that live on entry to n;

[0102] Out[n]: The variables that live on exit to n;

[0103] Among them, variables represents the registers corresponding to variables. In[n] and Out[n] respectively represent the sets of registers corresponding to input and output, and Def[n] and Use[n] respectively represent the sets of registers corresponding to definition and use.

[0104] The decompilation module 101 can traverse the blocks in the source code block and construct the use set and def set for each block. The specific construction process is as follows:

[0105] a) Traverse the instructions in the block in the execution order of the instructions in the block;

[0106] b) If the type of the operand of the instruction is Register and the action is kActionRead, then it will be added to the use set;

[0107] c) If the type of the operand of the instruction is Register and the action is kActionWrite, then it will be added to the def set;

[0108] d) If the operand type of the instruction is Address, then its base_reg and index_reg will be added to the use set;

[0109] The decompilation module 101 can establish a data flow analysis equation based on the above sets, as shown below:

[0110]

[0111]

[0112] Among them, n represents a block, and the symbol means that the set on the right side of the symbol is a subset of the set on the left side of the symbol. succ[n] represents the registers that are still valid in the block.

[0113] The decompilation module 101 can solve the above equation through a fixed-point algorithm as follows:

[0114]

[0115] After the fixed-point algorithm, the intersection of the in set at the function entry and the input parameter Reg specified by the Calling Convention is the input parameter register; the intersection of the out set at the function exit and the output parameter Reg specified by the Calling Convention is the possible return value register.

[0116] When performing stack analysis, the decompilation module 101 can analyze the instruction control flow by using an algorithm based on the extended stack pointer register (RSP) or an algorithm based on the extended base pointer register (RBP).

[0117] The process of the decompilation module 101 using the RSP-based algorithm for analysis can specifically include the following steps:

[0118] a. Based on the function prelogue part (entry basic block), check whether there is an offset in the RSP and record the offset value off;

[0119] Among them, the decompilation module 101 can judge the offset through the sub instruction or the push instruction. The decompilation module 101 also records the register associated with the RSP.

[0120] b. Traverse all instructions in all blocks, and find the usage scenario where the operand type of the operand is kTypeAddress, the action is kActionRead, the base_reg = RSP (associated register), and the memory offset (displacement, dis) is a positive number. This parameter is the (dis - off) / 8th stack parameter, and then count the total number of parameters S.

[0121] c. In the case that the rule b is met, further distinguish the parameter types:

[0122] If the other register operands of the same instruction are integer registers (RXX), and the instruction is not related to non-floating-point -> integer type conversion instructions, judge that the stack parameter is an integer. If the other register operands of the same instruction are floating-point registers (XMM), and the instruction is not related to non-integer -> floating-point type conversion instructions, judge that the stack parameter is a floating point.

[0123] The process of decompilation module 101 using the RBP-based algorithm for analysis can specifically include the following steps:

[0124] a. Traverse all instructions of all blocks to find the usage scenarios where the operand type of the operand is kTypeAddress, the action is kActionRead, the base_reg = RBP, and the dis is a positive number. This parameter is the (dis - 8) / 8th stack parameter, and count the total number of parameters X.

[0125] b. In the case that the rule in a is met, further distinguish the parameter types:

[0126] If the other register operands of the same instruction are integer registers (RXX), and the instruction is not related to floating-point -> integer type conversion instructions, determine that this stack parameter is an integer. If the other register operands of the same instruction are floating-point registers (XMM), and the instruction is not related to integer -> floating-point type conversion instructions, determine that this stack parameter is a floating point.

[0127] In some possible implementation manners, the decompilation module 101 may execute the above two algorithms simultaneously, and then take the maximum value of the total number of parameters S determined by the two algorithms.

[0128] After obtaining the total number of parameters and the parameter types, the decompilation module 101 can adjust the storage positions of the parameters according to the differences in function call rules. Specifically, the decompilation module 101 performs cross-platform processing on the input parameter registers and the stack according to the differences in function call rules, such as pushing several parameters in the register onto the stack, switching the stack pointer, etc., so that the perspectives of the input parameter registers and the stack space during runtime on different platforms are consistent.

[0129] For the sake of easy understanding, a specific example is described below.

[0130] Refer to Figure 5 the schematic diagram of parameter adjustment for a function shown. In this example, the function test includes a total of 10 parameters from i0 to i9. During runtime on the x86 platform, the parameters i0 to i5 are stored in registers, and i6 to i9 are stored in the stack. During runtime on the ARM platform, the parameters i0 to i7 are stored in registers, and i8 to i9 are stored in the stack. The decompilation module 101 can push the parameters i6 and i7 onto the stack and switch the stack pointer, so that the perspectives of the input parameter registers and the stack space during runtime on different platforms are consistent.

[0131] This method obtains accurate input parameter registers and stack input parameters by using compiler active register analysis and stack analysis, and reduces unnecessary register conversions of function call conventions.

[0132] Further, when the source platform and the target platform belong to different types of platforms, such as the source platform and the target platform being an x86 platform and an ARM platform respectively, etc., before decompiling the first code, the decompilation module 101 may first obtain the annotation information corresponding to the first code. The annotation information of the first code may include, for example, any one or more of the type, quantity, and type of jump address (internal or external jump in the assembly code, etc.) of the parameters in the first code. Thus, when decompiling the first code, the decompilation module 101 may determine the type, quantity, and jump address type of the parameters in the compiler IR according to the annotation information. Among them, the annotation information in the first code may be generated during the compilation of the assembly language code and is used to carry relevant information of the assembly language. Taking the parameter type in the annotation information as an example, assuming that the length of the parameter in the first code is 64 bits, then in the first IR obtained by decompiling the first code, the length of the parameter may still be 64 bits (such as floating-point type) and will not change to 32 bits (such as integer type). In this way, decompiling the first code according to the annotation information can keep the parameter type consistent before and after decompilation and reduce the possibility of logical errors in the first IR.

[0133] Of course, in other possible implementation manners, when decompiling the first code, the decompilation module 101 may not need to consider the annotation information. For example, when the source platform and the target platform are of the same type of platform, the similarity between these two platforms is relatively high and the difference is relatively small. For example, for the assembly language with the same instruction semantics between the source platform and the target platform, only the instruction format is different, etc. The decompilation module 101 can directly decompile the first code without relying on the annotation information of the first code.

[0134] Further, the instruction semantic library for translating the instruction string may further include vectorized instruction semantics, such as SIMD instruction semantics (the SIMD instruction can perform the same operation on each data in a group of data respectively to achieve parallel processing in space), etc. The vectorized instruction semantics can be used to perform vectorized translation on some instructions in the first code, so as to obtain the vectorized IR corresponding to the instruction.

[0135] Under normal circumstances, vectorized code (instructions) can be used to replace loop execution structures, which makes the program code more concise and has higher code execution efficiency. For example, when the first code includes SIMD instructions (used to sum multiple data), if the decompilation module 101 does not perform vectorization processing on the SIMD instructions, in the first IR obtained after decompilation, the corresponding code execution process is to perform a serial operation of reading data one by one from a set of data and summing them. After vectorizing the SIMD instructions, in the first IR obtained after decompilation, the corresponding code execution process is to read all the data from a set of data and perform a parallel summation calculation on all the data in the set.

[0136] Taking the case where the first code includes SIMD instructions as an example, when the decompilation module 101 translates the first code, for other instructions in the first code, it can be translated into an IR including corresponding functions according to the semantics of the instruction; for the SIMD instructions in the first code, it can perform vectorized translation on the SIMD instructions to obtain a first IR with a vectorization mark. In practical applications, this vectorization mark can be, for example, a special symbol in the first IR, such as symbols like "^", "!", "<", etc. Exemplarily, technicians can pre-add the semantics of SIMD instructions of the source platform to the imported instruction semantics library, so that at the compilation stage, the decompilation module 101 can identify the semantics of SIMD instructions in the first code according to this instruction semantics library.

[0137] Among them, the decompilation module 101 can directly or indirectly translate the SIMD instructions in the first code into an IR with a vectorization mark.

[0138] In a direct vectorization implementation, the decompilation module 101 can generate a third code based on a low-level language applicable to the target platform. This third code can be, for example, a low-level language code including intrinsic functions corresponding to the target platform. This intrinsic function can wrap language extensions or platform-related capabilities and be defined in high-level language header files such as C / C++. In this way, the generated third code related to the target platform can be used to describe the vectorization semantics of SIMD instructions of the source platform. Then, the decompilation module 101 can decompile this third code to obtain the first IR corresponding to the SIMD instructions of the target platform with a vectorization mark. In this way, after the subsequent compiler compiles this first IR, the SIMD instructions of the target platform can be obtained.

[0139] In an implementation of indirect vectorization, the decompilation module 101 may generate fourth code based on a high-level language that is independent of the target platform, and this fourth code may be used to describe the vectorization semantics of SIMD instructions. Then, the decompilation module 101 may decompile the fourth code based on the high-level language to obtain a first IR that is platform-independent and has vectorization tags. In this way, the subsequent compiler may perform automatic vectorization compilation on this first IR to generate SIMD instructions for the target platform.

[0140] S203: Compile the first IR obtained by decompilation into second code based on a low-level language applicable to the target platform, where the source platform and the target platform have different instruction sets.

[0141] In specific implementation, the compilation module 102 may compile the obtained first IR to obtain second code that can run on the target platform. Of course, this second code is low-level language code supported by the target platform, such as the assembly code corresponding to the target platform, etc.

[0142] Generally, the IR obtained by decompiling the first code may have certain optimization space. For example, when the decompilation module 101 does not decompile the first code based on the capabilities of the target platform, since the capabilities of different platforms usually vary, in order to make the decompiled IR applicable to multiple platforms, the decompilation module 101 may decompile the first code based on the lowest capabilities of multiple platforms, which enables the decompilation module 101 to also optimize the obtained IR according to the higher capabilities of the target platform, so that when the target platform executes the code corresponding to this IR, the code execution efficiency is relatively higher. Among them, the capabilities of the platform may include the data reading speed supported by the platform, the data access mode, etc.

[0143] For example, assume that when reading data on existing platforms, some platforms can read 64-bit (bit) data at one time, while some other platforms can read 128-bit data at one time. Therefore, after the decompilation module 101 decompiles the first code, the obtained IR may be the code shown above. If the target platform reads data only 64 bits at a time based on this code, however, the target platform can actually have the ability to read 128-bit data at one time, which causes the target platform to need to read 128-bit data in two times when reading 128-bit data, thus reducing the code execution efficiency. For this reason, the decompilation module 101 may, according to the ability of the target platform to read 128 bits at one time, optimize the code shown above to Figure 6 the code shown above. If the target platform reads data only 64 bits at a time based on this code, however, the target platform can actually have the ability to read 128-bit data at one time, which causes the target platform to need to read 128-bit data in two times when reading 128-bit data, thus reducing the code execution efficiency. For this reason, the decompilation module 101 may, according to the ability of the target platform to read 128 bits at one time, optimize the code shown above to Figure 6 the code shown above. If the target platform reads data only 64 bits at a time based on this code, however, the target platform can actually have the ability to read 128-bit data at one time, which causes the target platform to need to read 128-bit data in two times when reading 128-bit data, thus reducing the code execution efficiency. For this reason, the decompilation module 101 may, according to the ability of the target platform to read 128 bits at one time, optimize the code shown above to Figure 6The code shown below enables the target platform to read 128-bit data each time when reading data based on this code.

[0144] Based on this, in a further possible implementation, the decompilation module 101 can decompile the first code to generate a second IR, and then optimize the second IR according to the first target platform to obtain a first IR. For example, after decompiling the first code to obtain the second IR, the decompilation module 101 can also determine the semantics of each instruction string in the second IR, and based on the correspondence between the semantics and the compilation optimization rules, determine the compilation optimization rules corresponding to the semantics of each instruction string in the second IR, so as to optimize the second IR based on the determined compilation optimization rules to obtain a first IR. Thus, when the compilation module 102 performs compilation, it can compile the optimized first IR.

[0145] After compiling the first IR, a binary file in the ".obj" format or other formats can be obtained, and thus the assembly language code applied to the target platform can be generated according to this binary file. In this way, the assembly language code obtained through the above process can be applicable to the target platform, and thus can run successfully on the target platform, realizing the transplantation of the code of the source platform to the target platform.

[0146] In this embodiment, the target platform can be any platform different from the source platform, that is, the decompilation module 101 and the compilation module 102 can transform the first code into low-level language code applicable to any platform. Specifically, for the convenience of description, the above target platform can be referred to as the first target platform below. Then, the decompilation module 101 and the compilation module 102 can not only transform the first code of the source platform into low-level language code applicable to the first target platform based on the above process, but also transform the second code of the source platform into low-level language code applicable to the second target platform in a similar process, where the second target platform, the first target platform, and the source platform are different from each other, specifically having different instruction sets.

[0147] In practical applications, the target platform to which the software code is to be transplanted can be determined according to the user's needs. In an exemplary specific implementation, a target platform selection interface can be presented to the user. For example, a target platform selection interface as shown in Figure 7 can be presented on the display screen of the terminal. The target platform selection interface can provide multiple different candidate target platforms, such as Figure 7The target platforms 1, 2, …, N (N is a positive integer greater than 1) as shown. The user can perform a selection operation for the target platform on the target platform selection interface, and determine the first target platform or the second target platform from the presented multiple candidate platforms according to actual needs. Thus, through the above similar process, the first code can be transformed into the low-level language-based code applicable to the first target platform or the second target platform. For example, the user can click the drop-down menu button with the mouse to present multiple target platforms, and then move the cursor to the target platform to be selected and click the mouse to select the target platform, thereby determining the first target platform.

[0148] Further, based on the target platform selected by the user, relevant information of the target platform can also be presented on the target platform selection interface. For example, as Figure 7 shown, for the target platform selected by the user, the data processing capacity, applicable hardware type, required hardware environment, etc. of the target platform can also be presented.

[0149] In addition, after obtaining the low-level language-based code applicable to the first target platform through the above decompilation and compilation processes, corresponding prompt information can also be generated for the code. The prompt information can be used to indicate the differences between the source platform and the first target platform. Then, the prompt information can be presented on the interface where the user selects the target platform to prompt the user. For example, the user can be prompted about the correspondence between the instructions in the first code and the instructions in the code of the first target platform, such as highlighting the code instructions of the two platforms in a specific color. In practical applications, the prompt information can also be presented on other interfaces, not limited to the above target platform selection interface. This embodiment does not limit how to present the prompt information and the specific implementation manner of presenting the prompt information.

[0150] For ease of understanding, the technical solution of the embodiment of the present application will be described from the perspective of human-computer interaction. Refer to Figure 8 the schematic flowchart of the interaction between the user 801 and the computing device 802 as shown. Among them, the computing device 802 can specifically be a cloud device, such as a cloud server, or a local terminal / server, etc. As Figure 8 shown, the process can specifically include:

[0151] S801: The computing device 802 presents a target platform selection interface to the user 801, and multiple candidate target platforms are provided in the target platform selection interface.

[0152] In this embodiment, the computing device 802 can support porting software code to a variety of different target platforms. Then, the computing device 802 can first present a target platform selection interface to the user, and present on the target platform selection interface the target platforms that the computing device 802 supports for porting, for the user to select. Among them, there are at least different instruction sets between different platforms.

[0153] S802: The computing device 802 determines a first target platform from multiple target platforms according to the user's selection operation for the target platform.

[0154] S803: The user 801 sends a first code based on a low-level language applied to the source platform to the computing device 802.

[0155] Among them, the user 801 can specifically send the first code to the computing device 802 through media such as a terminal or a client.

[0156] In this embodiment, the first code can be a binary file based on the ".obj" format, or a binary file based on other formats. In other embodiments, the user can also send the assembly language code of the source platform to the computing device 802. In this way, after receiving the assembly language code, the computing device 802 can first perform assembly processing on the assembly language code to obtain a binary file in the ".obj" format or other formats.

[0157] S804: The computing device 802 presents a second code on the code display interface, and the second code is a code based on a low-level language applied to the first target platform.

[0158] Among them, the second code is obtained by processing the first code. Specifically, the second code can be obtained by the computing device 802 performing decompilation and compilation processing on the first code. For its specific implementation, reference can be made to the relevant descriptions in the foregoing embodiments, and details are not described here.

[0159] Moreover, after obtaining the second code, the computing device 802 can present the second code on the corresponding code display interface, so that the user 801 can view how the processed second code is specifically.

[0160] In this embodiment, the second code can be a binary file in the "obj" format or other formats. In other possible embodiments, the second code can also be a code based on other languages applicable to the first target platform. For example, it can be an assembly language code applicable to the first target platform, and the assembly language code can be obtained by converting the "obj" format file generated after decompiling and compiling the first code.

[0161] S805: The computing device 802 presents a prompt message.

[0162] In this embodiment, the computing device 802 can prompt the user 801, such as prompting the correspondence between the instructions in the first code and the instructions in the second code; or prompting the user 801 about the possible problems that may occur during the code transplantation process. In this embodiment, the content and specific implementation of the prompt information presented by the computing device 802 are not limited.

[0163] S806: The user 801 modifies the second code presented by the computing device 802.

[0164] Exemplarily, after viewing the second code or the prompt information, the user 801 can change the second code. For example, when a logical vulnerability is found in the second code through the prompt information or by viewing the second code, the second code can be modified to solve the logical vulnerability problem.

[0165] S807: The computing device 802 presents the modified second code.

[0166] In practical applications, the user 801 can also continue to further modify the modified second code until the code meets the user's expectations.

[0167] S808: The computing device 802 presents the first IR or the second IR.

[0168] During the decompilation process of the first code, the computing device 802 can first decompile the first code to obtain the second IR, and then, through further optimization of the second IR, obtain the first IR. Then, the computing device 802 can present the first IR or the second IR obtained during the decompilation process to the user 801.

[0169] In this way, the user 801 can view the first IR or the second IR presented by the computing device 802 and can debug the first IR or the second IR, so that the computing device 802 can obtain the corresponding low-level language-based code applied to the first target platform according to the debugged first IR or second IR.

[0170] In a further possible implementation manner, this embodiment may further include:

[0171] S809: The user 801 modifies the instruction semantic library corresponding to the source platform.

[0172] During the decompilation process, the computing device 802 generally decompiles the first code according to the instruction semantic library corresponding to the source platform, and the user can adjust the instruction semantic library used during the decompilation process by viewing the second code, the prompt information, or the IR. For example, add the semantics of the SIMD instructions corresponding to the source platform to the instruction semantic library, so that when the first code is decompiled based on the adjusted instruction semantic library, the obtained IR or the code applied to the first target platform can be better.

[0173] It should be noted that for this embodiment, the decompilation, compilation, and related processes of the first code by the computing device 802 can refer to the relevant descriptions in the foregoing embodiments and will not be elaborated here.

[0174] In the above text, in combination with Figures 1 to 8 , the code processing method provided by the present application has been described in detail. Next, in combination with Figures 9 to 10 , the computing device provided by the present application will be described.

[0175] With the same inventive concept as the above method, an embodiment of the present application also provides a computing device, which can implement the functions of the computing device in the embodiment shown in Figure 2 . Referring to Figure 9 shown, the device 900 may include:

[0176] An acquisition module 901, configured to acquire a first code based on a low-level language applied to a source platform;

[0177] A decompilation module 902, configured to decompile the first code to obtain a first intermediate representation IR;

[0178] A compilation module 903, configured to compile the first IR into a second code based on a low-level language applied to a first target platform, where the source platform and the first target platform have different instruction sets.

[0179] In a possible implementation manner, the decompilation module 902 is further configured to decompile the first code to obtain an IR corresponding to a second target platform, where the first IR is different from the IR corresponding to the second target platform, and the first target platform and the second target platform have different instruction sets.

[0180] In a possible implementation manner, the device 900 further includes:

[0181] A presentation module 904, configured to present a target platform selection interface;

[0182] A determination module 905, configured to determine the first target platform from multiple target platforms in response to a selection operation for the target platform.

[0183] In a possible implementation, the decompilation module 902 is specifically configured to:

[0184] Obtain the annotation information of the first code, where the annotation information includes any one or more of the type, quantity, and jump address type of the parameters in the first code;

[0185] Decompile the first code according to the annotation information to obtain the first IR.

[0186] In a possible implementation, the decompilation module 902 is specifically configured to:

[0187] Decompile the first code to obtain a second IR;

[0188] Optimize the second IR according to the first target platform to obtain the first IR.

[0189] In a possible implementation, the device 900 further includes:

[0190] A generation module 906, configured to generate a prompt message for prompting for an item to be checked, where the item to be checked is generated based on the difference between the first target platform and the source platform;

[0191] A presentation module 904, configured to present the prompt message.

[0192] In a possible implementation, the device is applied to the cloud, and the acquisition module 901 is specifically configured to receive the first code from the user;

[0193] The device 900 further includes: a communication module 907, configured to send a second code based on a low-level language applied to the first target platform to the user.

[0194] In a possible implementation, the decompilation module 902 is specifically configured to decompile the first code according to the instruction semantic library corresponding to the source platform.

[0195] In a possible implementation, the device 900 further includes:

[0196] A modification module 908, configured to modify the instruction semantic library in response to a modification operation on the instruction semantic library.

[0197] In a possible implementation, the acquisition module 901 is specifically configured to acquire a first code based on a low-level language and variables based on a high-level language applied to the source platform;

[0198] The decompilation module 902 is specifically configured to translate the first code into a first IR including functions, where the functions include formal parameters, and the actual parameters corresponding to the formal parameters are the variables.

[0199] In a possible implementation manner, the first IR includes a first variable and a second variable. The first variable has a first logical address, and the second variable has a second logical address, and the first logical address is different from the second logical address.

[0200] In a possible implementation manner, the decompilation module 902 is specifically configured to decompile the first code according to the target platform function call convention or the single instruction stream multiple data stream SIMD instruction.

[0201] The computing device 900 in this embodiment corresponds to Figure 2 the code processing method shown. Therefore, for the specific implementation of each functional module in the computing device 900 of this embodiment and the technical effects thereof, reference can be made to Figure 2 the relevant descriptions in the shown embodiment, which will not be elaborated here.

[0202] In addition, the embodiment of the present application further provides another computing device, and this device can implement the functions of the computing device 802 in the above Figure 8 shown embodiment. As shown in Figure 10 the figure, the device 1000 may include:

[0203] An acquisition module 1001, configured to acquire a first code based on a low-level language applied to a source platform;

[0204] An output module 1002, configured to output a second code, where the second code is a code based on a low-level language applied to a first target platform, and the second code is obtained by processing the first code, and the source platform and the first target platform have different instruction sets.

[0205] In a possible implementation manner, the output module 1002 is specifically configured to present the second code through a code display interface.

[0206] In a possible implementation manner, the device is applied to the cloud, and the acquisition module 1001 is specifically configured to receive the first code from a user;

[0207] The output module 1002 is specifically configured to output the second code to the user.

[0208] In a possible implementation manner, the device 1000 further includes:

[0209] A presentation module 1003, configured to present a target platform selection interface;

[0210] A determination module 1004, configured to determine the first target platform from multiple target platforms in response to a selection operation for a target platform.

[0211] In a possible implementation manner, the obtaining module 1001 is further configured to obtain an instruction semantic library corresponding to the source platform, where the instruction semantic library is used to process the first code.

[0212] In a possible implementation manner, the apparatus 1000 further includes:

[0213] A modification module 1005, configured to modify the instruction semantic library in response to a modification operation for the instruction semantic library.

[0214] In a possible implementation manner, the apparatus 1000 further includes:

[0215] A generation module 1006, configured to generate a prompt message, where the prompt message is used to prompt for an item to be checked, and the item to be checked is generated based on a difference between the first target platform and the source platform;

[0216] A presentation module 1003, configured to present the prompt message.

[0217] In a possible implementation manner, the apparatus 1000 further includes:

[0218] A presentation module 1003, configured to present a first intermediate representation IR, where the first IR is obtained by decompiling the first code, and the second code is obtained by compiling the first IR.

[0219] In a possible implementation manner, the presentation module 1003 is further configured to present a second IR, where the second IR is obtained by decompiling the first code, and the first IR is optimized according to the first target platform for the first IR.

[0220] In a possible implementation manner, the apparatus 1000 further includes:

[0221] A modification module 1005, configured to obtain a modified first IR in response to a modification operation for the first IR;

[0222] The presentation module 1003 is further configured to present a third code, where the third code is a low-level language-based code applied to the first target platform, and the third code is obtained by compiling the modified first IR.

[0223] In a possible implementation, the presentation module 1003 is further configured to present a third IR, which is obtained by decompiling the first code, and the third IR is used to generate code based on a low-level language applied to a second target platform, and the third IR is different from the first IR.

[0224] In a possible implementation, the apparatus 1000 further includes:

[0225] A modification module 1005, configured to obtain a modified second code in response to a modification operation on the second code;

[0226] The output module 1002 is further configured to output the modified second code.

[0227] In a possible implementation, the obtaining module 1001 is specifically configured to obtain a first code based on a low-level language and variables based on a high-level language applied to a source platform.

[0228] The computing apparatus 1000 in this embodiment corresponds to Figure 8 the code processing method shown. Therefore, for the specific implementation of each functional module in the computing apparatus 1000 in this embodiment and the technical effects thereof, reference may be made to Figure 8 the relevant descriptions in the shown embodiment, which will not be elaborated herein.

[0229] In addition, an embodiment of the present application further provides a computing apparatus, as Figure 11 shown. The apparatus 1100 may include a communication interface 1110 and a processor 1120. Optionally, the apparatus 1100 may further include a memory 1130. Among them, the memory 1130 may be disposed inside the apparatus 1100 or outside the apparatus 1100. Exemplarily, each of the actions in the above Figure 2 shown embodiments may be implemented by the processor 1120. The processor 1120 may obtain a first code applied to a source platform through the communication interface 1110 and is used to implement Figure 2 any method executed in. During the implementation process, each step of the processing flow may be completed by an integrated logic circuit in hardware or an instruction in software form in the processor 1120 Figure 2 the method executed in. For the sake of brevity, it will not be elaborated herein. The program code for the processor 1120 to implement the above method may be stored in the memory 1130. The memory 1130 and the processor 1120 are connected, such as coupled connection, etc.

[0230] Some features of the embodiments of the present application can be completed / supported by the processor 1120 executing program instructions or software code in the memory 1230. The software components loaded on the memory 1230 can be generalized functionally or logically. For example, Figure 9 the acquisition module 901, decompilation module 902, compilation module 903, presentation module 904, determination module 905, generation module 906, and modification module 908 shown. The function of the communication module 907 can be implemented by the communication interface 1110.

[0231] Any communication interface involved in the embodiments of the present application can be a circuit, bus, transceiver, or any other device that can be used for information interaction. For example, the communication interface 1110 in the device 1100. Exemplarily, the other device can be a device connected to the device 1100, such as a user terminal that provides the first code.

[0232] In addition, the embodiments of the present application also provide a computing device, such as Figure 12 shown, the device 1200 may include a communication interface 1210 and a processor 1220. Optionally, the device 1200 may further include a memory 1230. Among them, the memory 1230 can be set inside the device 1200 or outside the device 1200. Exemplarily, Figure 8 each action in the embodiments shown can be implemented by the processor 1220. The processor 1220 can obtain the first code applied to the source platform through the communication interface 1210 and be used to implement Figure 8 any method executed in. During the implementation process, each step of the processing flow can be completed by the integrated logic circuit of the hardware in the processor 1220 or the instructions in the form of software Figure 8 the method executed in. For the sake of brevity, it will not be elaborated here. The program code used by the processor 1220 to implement the above method can be stored in the memory 1230. The memory 1230 and the processor 1220 are connected, such as coupled connection, etc.

[0233] Some features of the embodiments of the present application can be completed / supported by the processor 1220 executing program instructions or software code in the memory 1230. The software components loaded on the memory 1230 can be generalized functionally or logically. For example, Figure 10 the acquisition module 1001, output module 1002, presentation module 1003, determination module 1004, modification module 1005, and generation module 1006 shown.

[0234] Any communication interface involved in the embodiments of the present application may be a circuit, a bus, a transceiver, or any other device that can be used for information interaction. For example, the communication interface 1210 in the device 1200. Exemplarily, the other device may be a device connected to the device 1200. For example, it may be a user terminal that provides the first code, etc.

[0235] The processor involved in the embodiments of the present application may be a general-purpose processor, a digital signal processor, an application-specific integrated circuit, a field-programmable gate array, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, and can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application may be directly embodied as being executed by a hardware processor, or executed by a combination of hardware and software modules in the processor.

[0236] The coupling in the embodiments of the present application is an indirect coupling or communication connection between devices, modules, or modules, and may be electrical, mechanical, or other forms, and is used for information interaction between devices, modules, or modules.

[0237] The processor may operate in cooperation with the memory. The memory may be a non-volatile memory, such as a hard disk drive (HDD) or a solid-state drive (SSD), etc., or may also be a volatile memory, such as a random-access memory (RAM). The memory is any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.

[0238] The embodiments of the present application do not limit the specific connection medium between the above-mentioned communication interface, processor, and memory. For example, the memory, processor, and communication interface may be connected through a bus. The bus may be divided into an address bus, a data bus, a control bus, etc.

[0239] Based on the above embodiments, the embodiments of the present application further provide a computer storage medium. The software program stored in the storage medium can implement the method executed by the proxy edge device or the edge device or the cloud center provided in any one or more of the above embodiments when being read and executed by one or more processors. The computer storage medium may include: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory, a random-access memory, a magnetic disk, or an optical disc.

[0240] Based on the above embodiments, an embodiment of the present application further provides a chip, which includes a processor for implementing the functions of the proxy edge device or the edge device or the cloud center involved in the above embodiments. For example, it is used to implement Figures 3 to 4 the method executed by the proxy edge device in Figures 3 to 4 the method executed by the edge device in Figures 3 to 4 or the method executed by the cloud center in

[0241] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0242] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of processes and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0243] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0244] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Therefore, the instructions executed on the computer or other programmable device provide for implementing the functions in the processFigure 1 one process or multiple processes and / or blocks Figure 1 steps of the functions specified in one block or multiple blocks

[0245] The terms "first", "second", etc. in the description and claims of this application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that these terms can be interchanged under appropriate circumstances, and this is only a way of distinguishing objects with the same attributes when describing the embodiments of this application.

[0246] Obviously, those skilled in the art can make various changes and modifications to the embodiments of this application without departing from the scope of the embodiments of this application. Thus, if these modifications and variations of the embodiments of this application fall within the scope of the claims of this application and their equivalent technologies, this application is also intended to include these changes and modifications.

Claims

1. A code processing method, characterized in that, the method includes: obtaining a first code based on a low-level language applied to a source platform; decompiling the first code to obtain a first intermediate representation IR; compiling the first intermediate representation IR into a second code based on a low-level language applied to a first target platform, where the source platform and the first target platform have different instruction sets; the obtaining of the first code based on a low-level language applied to the source platform includes: obtaining a first code based on a low-level language applied to the source platform and variables based on a high-level language; the decompiling of the first code includes: translating the first code into a first intermediate representation IR including functions, where the functions include formal parameters, and the semantics expressed by the functions are consistent with the semantics of the first code; establishing a correspondence between the variables and the formal parameters of the functions, and passing the variables as actual parameters to the formal parameters of the functions.

2. The method according to claim 1, characterized in that, the method further includes: decompiling the first code to obtain an IR corresponding to a second target platform, where the first intermediate representation IR is different from the IR corresponding to the second target platform, and the first target platform and the second target platform have different instruction sets.

3. The method according to claim 1, characterized in that, the method further includes: presenting a target platform selection interface; responding to a selection operation for the target platform, and determining the first target platform from multiple target platforms.

4. The method according to claim 1, characterized in that, the decompiling of the first code to obtain a first intermediate representation IR includes: obtaining annotation information of the first code, where the annotation information includes any one or more of the types of parameters, the number of parameters, and the types of jump addresses in the first code; decompiling the first code according to the annotation information to obtain the first intermediate representation IR.

5. The method according to claim 1, characterized in that, the decompiling of the first code to obtain a first intermediate representation IR includes: decompiling the first code to obtain a second IR; optimizing the second IR according to the first target platform to obtain the first intermediate representation IR.

6. The method according to claim 1, characterized in that, the method further includes: generating prompt information for prompting for items to be checked, where the items to be checked are generated based on the differences between the first target platform and the source platform; presenting the prompt information.

7. The method according to claim 1, characterized in that, the method is applied to the cloud, and the obtaining of the first code based on a low-level language applied to the source platform includes: receiving a first code from a user; the method further includes: sending the second code to the user.

8. The method according to claim 1, characterized in that, the decompiling of the first code includes: decompiling the first code according to the instruction semantic library corresponding to the source platform.

9. The method according to claim 8, It is characterized in that The method further includes: Modifying the instruction semantic library in response to a modification operation on the instruction semantic library.

10. The method according to claim 1, It is characterized in that The first intermediate representation IR includes a first variable and a second variable, the first variable has a first logical address, the second variable has a second logical address, and the first logical address is different from the second logical address.

11. The method according to any one of claims 1 to 10, It is characterized in that The decompiling of the first code includes: Decompiling the first code according to the target platform function call convention or single instruction stream multiple data stream SIMD instructions.

12. A code processing method, It is characterized in that The method includes: Obtaining a first code based on a low-level language and variables based on a high-level language applied to a source platform; Outputting a second code, where the second code is a code based on a low-level language applied to a first target platform, the second code is obtained based on a first intermediate representation IR including functions, the first intermediate representation IR is obtained by decompiling the first code, the functions include formal parameters, the semantics expressed by the functions are consistent with the semantics of the first code, the variables correspond to the formal parameters of the functions, the variables are passed as actual parameters to the formal parameters of the functions, and the source platform and the first target platform have different instruction sets.

13. The method according to claim 12, It is characterized in that The outputting of the second code includes: Presenting the second code through a code display interface.

14. The method according to claim 12, It is characterized in that The method is applied to the cloud, and the obtaining of the first code includes: Receiving the first code from a user; The outputting of the second code includes: Outputting the second code to the user.

15. The method according to claim 12, It is characterized in that The method further includes: Presenting a target platform selection interface; Responding to a selection operation for the target platform and determining the first target platform from multiple target platforms.

16. The method according to claim 12, It is characterized in that The method further includes: Obtaining an instruction semantic library corresponding to the source platform, where the instruction semantic library is used to process the first code.

17. The method according to claim 16, It is characterized in that The method further includes: Modifying the instruction semantic library in response to a modification operation on the instruction semantic library.

18. The method according to claim 12, It is characterized in that The method further includes: Generating a prompt message for prompting for an item to be checked, where the item to be checked is generated based on the difference between the first target platform and the source platform; Presenting the prompt message.

19. The method according to claim 12, It is characterized in that The method further includes: Presenting a first intermediate representation IR, where the first intermediate representation IR is obtained by decompiling the first code, and the second code is obtained by compiling the first intermediate representation IR.

20. The method according to claim 19, wherein, the method further comprises: presenting a second IR, which is obtained by decompiling the first code, and the first intermediate representation IR is optimized according to the first target platform.

21. The method according to claim 19, wherein, the method further comprises: responding to a modification operation on the first intermediate representation IR to obtain a modified first intermediate representation IR; presenting a third code, which is a low-level language-based code applied to the first target platform and is obtained by compiling the modified first intermediate representation IR.

22. The method according to claim 19, wherein, the method further comprises: presenting a third IR, which is obtained by decompiling the first code, and the third IR is used to generate a low-level language-based code applied to the second target platform, and the third IR is different from the first intermediate representation IR.

23. The method according to any one of claims 12 to 22, wherein, the method further comprises: responding to a modification operation on the second code to obtain a modified second code; outputting the modified second code.

24. A code processing device, wherein, the device comprises: an acquisition module, configured to acquire a first low-level language-based code applied to a source platform and variables based on a high-level language; a decompilation module, configured to decompile the first code to obtain a first intermediate representation IR; a compilation module, configured to compile the first intermediate representation IR into a second low-level language-based code applied to a first target platform, and the source platform and the first target platform have different instruction sets; the decompilation module is specifically configured to translate the first code into a first intermediate representation IR including functions, the functions include formal parameters, and the semantics expressed by the functions are consistent with the semantics of the first code; establish a correspondence between the variables and the formal parameters of the functions, and the variables are passed as actual parameters to the formal parameters of the functions.

25. The device according to claim 24, wherein, the decompilation module is further configured to decompile the first code to obtain an IR corresponding to a second target platform, the first intermediate representation IR is different from the IR corresponding to the second target platform, and the first target platform and the second target platform have different instruction sets.

26. The device according to claim 24, wherein, the device further comprises: a presentation module, configured to present a target platform selection interface; a determination module, configured to respond to a selection operation for the target platform and determine the first target platform from multiple target platforms.

27. The device according to claim 24, wherein, the decompilation module is specifically configured to: acquire annotation information of the first code, and the annotation information includes any one or more of the type, quantity, and jump address type of the parameters in the first code. Decompile the first code according to the annotation information to obtain the first intermediate representation IR.

28. The apparatus according to claim 24, wherein, the decompilation module is specifically configured to: decompile the first code to obtain a second IR; optimize the second IR according to the first target platform to obtain the first intermediate representation IR.

29. The apparatus according to claim 24, wherein, the apparatus further comprises: a generation module, configured to generate a prompt message for prompting a to-be-checked item, the to-be-checked item being generated based on a difference between the first target platform and the source platform; a presentation module, configured to present the prompt message.

30. The apparatus according to claim 24, wherein, the apparatus is applied to the cloud, and the acquisition module is specifically configured to receive a first code from a user; the apparatus further comprises: a communication module, configured to send a second code based on a low-level language applied to the first target platform to the user.

31. The apparatus according to claim 24, wherein, the decompilation module is specifically configured to decompile the first code according to an instruction semantic library corresponding to the source platform.

32. The apparatus according to claim 31, wherein, the apparatus further comprises: a modification module, configured to modify the instruction semantic library in response to a modification operation on the instruction semantic library.

33. The apparatus according to claim 24, wherein, the first intermediate representation IR includes a first variable and a second variable, the first variable has a first logical address, the second variable has a second logical address, and the first logical address is different from the second logical address.

34. The apparatus according to any one of claims 24 to 33, wherein, the decompilation module is specifically configured to decompile the first code according to a target platform function call convention or a single instruction stream multiple data stream SIMD instruction.

35. A code processing apparatus, wherein, the apparatus comprises: an acquisition module, configured to acquire a first code based on a low-level language applied to a source platform and a variable based on a high-level language; an output module, configured to output a second code, the second code being a code based on a low-level language applied to a first target platform, the second code being obtained based on a first intermediate representation IR including a function, the first intermediate representation IR being obtained by decompiling the first code, the function includes formal parameters, the semantics expressed by the function is consistent with the semantics of the first code, the variable corresponds to the formal parameter of the function, the variable is passed as an actual parameter to the formal parameter of the function, and the source platform and the first target platform have different instruction sets.

36. The apparatus according to claim 35, wherein, the output module is specifically configured to present the second code through a code display interface.

37. The apparatus according to claim 35, wherein, The device is applied to the cloud, and the obtaining module is specifically configured to receive a first code from a user; The output module is specifically configured to output the second code to the user.

38. The device according to claim 35, wherein, The device further includes: A presentation module, configured to present a target platform selection interface; A determination module, configured to determine the first target platform from multiple target platforms in response to a selection operation for the target platform.

39. The device according to claim 35, wherein, The obtaining module is further configured to obtain an instruction semantic library corresponding to the source platform, and the instruction semantic library is used to process the first code.

40. The device according to claim 39, wherein, The device further includes: A modification module, configured to modify the instruction semantic library in response to a modification operation for the instruction semantic library.

41. The device according to claim 35, wherein, The device further includes: A generation module, configured to generate a prompt message, and the prompt message is used to prompt for an item to be checked, and the item to be checked is generated based on the difference between the first target platform and the source platform; A presentation module, configured to present the prompt message.

42. The device according to claim 35, wherein, The device further includes: A presentation module, configured to present a first intermediate representation IR, and the first intermediate representation IR is obtained by decompiling the first code, and the second code is obtained by compiling the first intermediate representation IR.

43. The device according to claim 42, wherein, The presentation module is further configured to present a second IR, and the second IR is obtained by decompiling the first code, and the first intermediate representation IR is optimized according to the first target platform.

44. The device according to claim 42, wherein, The device further includes: A modification module, configured to obtain a modified first intermediate representation IR in response to a modification operation for the first intermediate representation IR; The presentation module is further configured to present a third code, and the third code is a low-level language-based code applied to the first target platform, and the third code is obtained by compiling the modified first intermediate representation IR.

45. The device according to claim 42, wherein, The presentation module is further configured to present a third IR, and the third IR is obtained by decompiling the first code, and the third IR is used to generate a low-level language-based code applied to the second target platform, and the third IR is different from the first intermediate representation IR.

46. The device according to any one of claims 35 to 45, wherein, The device further includes: A modification module, configured to obtain a modified second code in response to a modification operation for the second code; The output module is further configured to output the modified second code.

47. A computing device, wherein, The device includes a memory and a processor. The memory is used to store software instructions. The processor calls the software instructions stored in the memory to execute the method according to any one of claims 1 to 11 above.

48. A computing device, wherein, the device includes a memory and a processor. The memory is used to store software instructions. The processor calls the software instructions stored in the memory to execute the method according to any one of claims 12 to 23 above.

49. A computer-readable storage medium, wherein, it includes instructions for implementing the method according to any one of claims 1 to 11.

50. A computer-readable storage medium, wherein, it includes instructions for implementing the method according to any one of claims 12 to 23.

51. A computer program product, wherein, when it runs on a computer, it causes the computer to execute the method according to any one of claims 1 to 11.

52. A computer program product, wherein, when it runs on a computer, it causes the computer to execute the method according to any one of claims 12 to 23.

Citation Information

Patent Citations

  • Load module compiler

    US20190227779A1