Code optimization method, device and equipment

By inserting a non-dependent instruction set after the hardware call instruction, the processor idle problem caused by the hardware call instruction is solved, and more efficient code execution is achieved.

CN120848889APending Publication Date: 2025-10-28HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510789763.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-12
Publication Date
2025-10-28

Smart Images

  • Figure CN120848889A_ABST
    Figure CN120848889A_ABST
Patent Text Reader

Abstract

The invention discloses a code optimization method, device and equipment, and the method comprises the steps: positioning a function where a first hardware call instruction corresponding to a first hardware unit in at least one hardware unit in a code is located, and determining an instruction set which has no dependence relation with an embedded instruction executed by the first hardware unit in the function; after the code related to the instruction set is inserted into the first hardware calling instruction, the optimized code is obtained, so that when the optimized code is executed, when the data processing equipment executes the optimized code, the instruction set is executed by a processing unit of the data processing equipment while the first hardware unit is called to execute the embedded instruction, and the data processing equipment executes the optimized code. The code can be efficiently executed, and the execution efficiency of the code is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computing, and more particularly to a code optimization method, apparatus, and device. Background Technology

[0002] Data processing devices, such as data processing units (DPUs), contain a special type of instruction in their executed code—hardware call instructions. These instructions trigger the data processing device to perform hardware call operations. A hardware call operation refers to an operation that requires the participation of the data processing device's hardware unit (a hardware unit within the data processing device configured to perform specific functions (such as encryption), which embeds execution instructions. Triggered by the hardware call instruction, this hardware unit executes these embedded instructions to complete the designated function). While this hardware unit is executing its embedded instructions, the processor in the processing device stops executing code until the hardware unit has completed its operation. Thus, the processor is idle during the execution of these embedded instructions. The more hardware call instructions in the code executed by the data processing device, the longer the processor's idle time becomes, resulting in underutilization of the processor's computing resources and impacting execution efficiency. Summary of the Invention

[0003] This application provides a code optimization method, apparatus, and device for improving code execution efficiency.

[0004] Firstly, embodiments of this application provide a code optimization method. This method optimizes code executed by a data processing device, which includes at least one hardware unit with embedded execution instructions. Embodiments of this application do not limit the specific type of data processing device; for example, the data processing device could be a DPU, a graphics processing unit (GPU), or a chip with other data processing functions. This method can be executed by a code optimization device, which can be the data processing device itself or other devices besides the data processing device. In this method:

[0005] The code optimization device locates the function containing the first hardware call instruction corresponding to the first hardware unit in at least one hardware unit in the code, and determines the set of instructions in that function that have no dependency on the embedded instructions executed by the first hardware unit.

[0006] The code optimization device inserts instruction set-related code after the first hardware call instruction to obtain optimized code. This optimizes the code so that when the optimized code is executed, the instruction set is executed by the processing unit of the data processing device simultaneously with the execution of the embedded instructions by the first hardware unit. This application does not limit the specific form of the processing unit in the data processing device; the processing unit can be a processor core in the data processing device, or it can be a thread or process in the data processing device.

[0007] Using the above method, the code optimization device can analyze and locate the function containing the first hardware call instruction corresponding to the first hardware unit of the processing device in the code. It then inserts relevant code from the instruction set within the function containing the first hardware call instruction—code that has no dependency on the embedded instructions executed by the first hardware unit—after the first hardware call instruction, thus optimizing the code. In this way, during the execution of the optimized code, while the first hardware unit calls the embedded instructions, the data processing device (its processing unit) can also execute the instruction set. Therefore, idle processing units are effectively utilized, allowing the code to be executed efficiently and improving execution efficiency.

[0008] In one possible implementation, the code optimization device determines the execution time of the embedded instructions executed by the hardware units invoked by at least one hardware call instruction in the code; the code optimization device determines the first hardware unit from the hardware units based on the execution time of each hardware unit.

[0009] Using the above method, the first hardware unit is determined based on the execution time of each hardware unit. If the first hardware unit is a hardware unit with a longer execution time, then the first hardware call instruction is determined. The first hardware call instruction is a hardware call instruction with greater optimization space among all hardware call instructions. That is, the latency of the first hardware call instruction is longer, and optimizing the first hardware call instruction can achieve better optimization results.

[0010] In one possible implementation, the first hardware call instruction corresponds to a wait instruction. The wait instruction is used to prevent the processing unit from continuing to execute code. When the code optimization device inserts instruction set related code after the hardware call instruction to obtain optimized code, it can insert the instruction set related code between the first hardware call instruction and the wait instruction, so that the processing unit executes the wait instruction after executing the instruction set.

[0011] By using the above method, the instruction set-related code is located before the instruction waiting, which ensures that the instruction set can be executed before the first hardware unit finishes executing the embedded instructions during the execution of the optimized code, thereby improving the parallelism of code execution.

[0012] In one possible implementation, when the code optimization device determines the execution time of the embedded instruction executed by the hardware unit called by any hardware call instruction, it instrumentes the hardware call instruction and inserts duration statistics code at the hardware call instruction; the execution time of the embedded instruction is counted by the duration statistics code during the execution of the embedded instruction by the hardware unit instructed by the hardware call instruction.

[0013] Using the above method, the code optimization device can easily and quickly obtain the execution time of hardware units through instrumentation.

[0014] In one possible implementation, the first hardware unit satisfies some or all of the following:

[0015] Condition 1: The execution time of the first hardware unit executing the embedded instructions is the largest among the N hardware units whose execution times of the embedded instructions are the longest, where N is a positive integer.

[0016] Condition 2: The execution time of the embedded instructions executed by the first hardware unit is greater than the execution time threshold.

[0017] Using the above method, since the execution time of the first hardware unit is relatively long, more code can be inserted after the first hardware call instruction corresponding to the first hardware unit to greatly improve the execution efficiency of the code.

[0018] In one possible implementation, the code optimization device can output optimized code to the user, which is presented in part or all of the following forms: intermediate representations formed during compilation, source code, and assembly files.

[0019] By using the methods described above, displaying the optimized code allows users to understand the optimization results in a timely manner, thus improving the user experience.

[0020] In one possible implementation, the code optimization device can show the user the first hardware call instruction and the instruction set-related code.

[0021] By using the above method, by displaying the first hardware call instruction and the related code of the instruction set that can be inserted after the first hardware call instruction, the user can decide for themselves whether to optimize the first hardware call instruction, or which first hardware call instructions to optimize.

[0022] In one possible implementation, after the code optimization device determines the execution time of the embedded instructions executed by the hardware units called by at least one hardware call instruction in the code, it generates and displays a duration analysis report to the user. The duration analysis report records the execution time of the embedded instructions executed by the hardware units called by at least one hardware call instruction.

[0023] By using the above method and displaying a duration analysis report, users can clearly understand whether the code has potential for optimization and what hardware call instructions can be optimized.

[0024] Secondly, embodiments of this application also provide a code optimization apparatus. This apparatus has the function of implementing the behavior described in the first aspect or the method example of the first aspect. The beneficial effects can be found in the description of the first aspect and will not be repeated here. The function can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above-described functions. In one possible design, the code optimization apparatus includes an instruction determination module and an optimization module. Optionally, it may also include a duration determination module and an output module. These modules can execute the corresponding functions in the method example of the first aspect, as detailed in the method example, and will not be repeated here.

[0025] Thirdly, this application also provides a computing device, which includes a processor and a memory, and may further include a communication interface. The processor executes program instructions in the memory to perform the method provided in the first aspect or any possible implementation thereof. The memory is coupled to the processor and stores computer program instructions and data necessary for determining code optimization processes. The communication interface is used for communicating with other devices.

[0026] Fourthly, this application provides a computing device system including at least one computing device. Each computing device includes a memory and a processor. The processor of at least one computing device is used to access code in the memory to execute the methods provided in the first aspect or any possible implementation thereof.

[0027] Fifthly, this application provides a computer-readable storage medium that, when executed by a computing device, allows the computing device to perform the method provided in the first aspect or any possible implementation thereof. The storage medium stores computer program instructions. The storage medium includes, but is not limited to, volatile memory, such as random access memory, and non-volatile memory, such as flash memory, hard disk drive (HDD), and solid-state drive (SSD).

[0028] Sixthly, this application provides a computing device program product, which includes computer program instructions. When executed by a computing device, the computing device performs the methods provided in the first aspect or any possible implementation thereof. The computer program product can be a software installation package, and when it is necessary to use the methods provided in the first aspect or any possible implementation thereof, the computer program product can be downloaded and executed on the computing device.

[0029] In a seventh aspect, this application also provides a computer chip connected to a memory, the chip being used to read and execute computer program instructions stored in the memory, and to execute the methods described in the first aspect and various possible implementations of the first aspect.

[0030] For the technical effects that can be achieved in the second to seventh aspects mentioned above, please refer to the description of the technical effects that can be achieved by the corresponding design scheme in the first aspect mentioned above. This application will not repeat them here. Attached Figure Description

[0031] Figure 1A A schematic diagram of the structure of a code optimization system provided in this application;

[0032] Figure 1B This application provides a schematic diagram of a code optimization scenario;

[0033] Figure 2 A schematic diagram of a code optimization method provided in this application;

[0034] Figure 3 A visual interface diagram provided for this application;

[0035] Figure 4 A schematic diagram of a duration analysis report provided for this application;

[0036] Figure 5 A schematic diagram of a code segmentation device provided in this application;

[0037] Figures 6-7 A schematic diagram of the result of a computing device provided in this application. Detailed Implementation

[0038] Before describing the code optimization method provided in the embodiments of this application, the concepts involved in the embodiments of this application will be explained first:

[0039] (1) Source code, assembly files, intermediate representations, microcode, executable code.

[0040] Source code is written by programmers using a language supported by development tools. Source code consists of a set of numbers or letters that indicate meaning. It includes multiple lines of code, each with a line number to identify it.

[0041] Source code cannot be executed directly on a computer. It needs to be compiled, such as by calling a compiler, to compile the source code into binary code that the machine can recognize. This binary code is also called executable code.

[0042] It's important to note that source code can contain a special type of instruction—microcode. Microcode, also known as microinstructions, is a series of relatively simple instructions broken down from complex instructions within a Complex Instruction Set Computer (CISC) architecture. Microcode is a low-level instruction in computer engineering, existing between the processor and other hardware and the machine instruction set (i.e., binary code). Microcode can be recognized and executed by the hardware. The hardware call instructions mentioned later belong to microcode.

[0043] The execution of microcode requires the assistance of certain hardware units with specific functions. Taking hardware call instructions as an example, when the processing unit in the processing device executes the executable code compiled from the source code, upon encountering a hardware call instruction, it invokes the hardware unit to execute the embedded execution instructions on that hardware unit. At this time, the processing unit is in a low-power state. After the hardware unit finishes executing the embedded instructions, the processing unit continues to execute the instructions following the hardware call instruction.

[0044] Assembly language is a low-level language that facilitates direct communication with computer hardware. It uses instructions contained in a hardware instruction set to represent the operations the processor must perform. The hardware instruction set contains a fixed set of instructions. In other words, assembly language uses instructions from the hardware instruction set to describe the code logic in the source code. It is a language that lies between high-level languages ​​(such as the languages ​​used to write source code) and binary languages ​​(the languages ​​used to execute code). Typically, assembly language can be expressed using hexadecimal and binary numbers, making it readable and understandable by humans. In this embodiment, the file (which can also be understood as a type of code) created by describing the source code using assembly language is called an assembly file.

[0045] Intermediate representation (IR) is intermediate code generated during the compilation of source code. Based on the intermediate representation, optimization of the intermediate identifiers can ultimately yield executable code or assembly files.

[0046] Source code, assembly files, intermediate representations, and executable code can be understood as different forms of code representation. There are correspondences between source code, assembly files, intermediate representations, and executable code. In other words, taking executable code as an example, the instructions in the executable code can correspond to code segments in the source code, assembly files, or intermediate representations; these instructions can be understood as the representation of the corresponding code segments in the executable code. Similarly, taking assembly files as an example, one or more instructions in the assembly file can correspond to code segments in the source code, executable code, or intermediate representations.

[0047] (2) Class, inner class, package, function / method, field / variable.

[0048] Source code must adhere to certain coding standards during its development, such as object-oriented programming (OOP) and procedural programming (POP). Languages ​​for object-oriented programming include C++, Java, and C#, while languages ​​for procedural programming include Fortran and C.

[0049] Software code written according to certain coding standards can include code elements such as classes, functions / methods, and fields / variables. Classes can include inner classes and functions / methods; functions / methods can call fields / variables. Inner classes are similar to classes; inner classes can also contain functions / methods.

[0050] In object-oriented programming, a class is a construct in an object-oriented computer programming language that describes the common methods and fields of the objects created.

[0051] A function / method is a subroutine within a class. A method typically consists of a series of statements and performs a specific function. In object-oriented programming, it's called a method; in procedural programming, it's called a function.

[0052] Fields / variables store data such as integers and characters, strings, hash tables, pointers, etc.

[0053] (3) Hardware call instructions.

[0054] There is a special type of "instruction" in the code. The execution of this type of instruction requires calling a hardware unit. That is, when executing the "instruction", the underlying logic of the hardware unit needs to be called to implement some functions. For ease of distinction, this application embodiment refers to this type of "instruction" as hardware call instruction.

[0055] A "hardware call instruction" refers to a hardware call instruction that invokes a hardware unit in the data processing device to execute embedded instructions. These embedded instructions are execution instructions that are embedded within the hardware unit. "Embedded" means that the execution instructions are programmed into the hardware unit and cannot be altered.

[0056] Different hardware call instructions can call the same or different hardware units, and there is a correspondence between the hardware call instructions and the hardware units they call.

[0057] In practical applications, the code contains wait instructions corresponding to hardware call instructions. These wait instructions prevent the processing unit in the data processing device from executing instructions following the hardware call instruction until the hardware unit called by the hardware call instruction completes the execution of its embedded instructions. Therefore, these wait instructions have a blocking effect, preventing code execution.

[0058] The code executed by a data processing device may contain hardware call instructions. Specifically, the data processing device includes a processing unit that can execute other instructions in the code besides the hardware call instructions. The data processing device also includes a hardware unit. During the execution of the code, when the hardware call instruction is encountered, the processing unit is idle and executes embedded instructions. After the hardware unit finishes executing the embedded instructions, the processing unit executes the instructions following the hardware call instruction.

[0059] This application does not limit the specific form of the data processing device. The data processing device can be a device including a data processing unit (DPU) or a device including a graphics processing unit (GPU). Any device capable of processing is applicable to this application. Specifically, within the data processing device, it includes one or more units with processing capabilities—processing units. This application does not limit the specific form of the processing unit; for example, the processing unit can be a thread or process within the data processing device, or it can be a processor core or processor of the data processing device.

[0060] Taking a data processing device containing a Data Processing Unit (DPU) as an example, the DPU contains a hardware unit A with embedded execution instructions. These embedded instructions are the instructions required for data encryption. Correspondingly, the hardware call instructions for the DPU can include hardware call instructions for data encryption. These data encryption hardware call instructions can be set in the source code to utilize hardware unit A within the DPU to encrypt data. When the DPU executes the executable code compiled from this source code, upon encountering a data encryption hardware call instruction, it invokes hardware unit A to encrypt the data indicated by that instruction. Compared to the processing unit performing data encryption itself, the use of embedded instructions by hardware unit A results in better data encryption efficiency, achieving a "hardware acceleration" effect.

[0061] Here, we further explain the hardware units of the DPU. The "hardware units of the DPU" refer to the circuit logic within the DPU that has code (i.e., embedded execution instructions, simply called embedded instructions) already programmed into it. Hardware call instructions essentially invoke this circuit logic to implement the corresponding function. The embedded instructions of this circuit logic mean that its function is fixed and cannot be changed. The function that the embedded instructions can achieve is the function that the circuit logic can achieve. In contrast to hardware calls is "software" code. Software code is code stored in memory (such as the RAM of a computing device, or the flash memory of the DPU). When running this software code, the processor reads it from memory and executes it. Software code is not fixed; programmers can modify it and then store the modified software code in memory (in practical applications, the source code of the software code is compiled into executable code and stored in memory) for the processor to call.

[0062] As explained above regarding hardware call instructions, these instructions involve a certain time delay. When a data processing device (such as a DPU) executes a hardware call instruction, it is in an idle state, and the hardware unit of the data processing device executes embedded instructions. During the execution of the hardware call instruction, the data processing device (processing unit) must wait until the result of the embedded instructions in the hardware unit is obtained before it can continue executing instructions following the hardware call instruction. If the code executed by the data processing device contains multiple hardware call instructions, there will be significant delays during code execution, reducing execution efficiency. Therefore, code optimization can be performed to reduce potential delays during code execution.

[0063] In this embodiment, the code optimization device can optimize some or all of the hardware call instructions in the code executed by the data processing device. Taking the optimization of the first hardware call instruction corresponding to the first hardware unit as an example, the instruction set related code (referred to as the dependency-free instruction set in this embodiment for ease of explanation) that has no dependency relationship with the embedded instructions of the first hardware unit in the function where the first hardware call instruction is located is inserted after the hardware call instruction. This allows the instruction set to be executed by the processing unit of the data processing device at the same time that the first hardware unit is called to execute the embedded instructions. By inserting the dependency-free instruction set related code after the first hardware call instruction in this way, the dependency-free instruction set can be executed during the process of the first hardware unit executing the embedded instructions triggered by the first hardware call instruction, thereby improving the execution efficiency of the code.

[0064] like Figure 1A The diagram shown is a structural schematic of a code optimization system provided in an embodiment of this application. The code optimization system includes a client 100 and a code optimization device 200.

[0065] The client 100 is deployed close to the user and can interact with the user. In this embodiment, the user can provide code that the data processing device can execute, such as source code or assembly files, to the code optimization device 200 through the client 100. The user can also optimize the code for the code optimization device 200 through the client 100. In other words, the client 100 can be understood as a "front-end" device of the code optimization device 200 deployed on the user side.

[0066] The client 100 can display a visual interface to the user through a device such as a monitor. The user can operate on the client 100 through external input / output devices (such as a monitor, keyboard, and mouse). For example, the user can type source code into the client 100, or upload source code or assembly files. The user can also trigger the client 100 to send a code optimization request to the code optimization device 200 through the "code optimization" function option provided by the client 100. This code optimization request is used to request optimization of the code (such as source code or assembly files) provided by the user. For example, after the code optimization device 200 completes code optimization, the client 100 can display a scheduling report to the user indicating the hardware call instructions that can be optimized in the code (such as the first hardware call instruction mentioned in the embodiments of this application) and the code related to the dependent instruction set of the hardware call instruction. The client 100 can also display the optimized code to the user, that is, the code with the dependent instruction set related code inserted after the hardware call instruction. The client 100 can also display a duration analysis report to the user, which records the execution time of the embedded instructions executed by the hardware unit called by at least one hardware call instruction.

[0067] The code optimization device 200 can determine optimizable hardware call instructions (referred to as first hardware call instructions in this embodiment) from code containing at least one hardware call instruction, and determine the dependency-free instruction set of the first hardware call instruction. It then extracts the related code of the dependency-free instruction set into the first hardware call instruction to obtain optimized code. In this embodiment, the selection range of optimizable hardware call instructions, i.e., the at least one hardware call instruction, is referred to as candidate hardware call instructions. The code optimization device 200 can also display scheduling reports, optimized code, or duration analysis reports to the user through the client 100.

[0068] In this embodiment of the application, the code optimization device 200 has the following functions:

[0069] Function 1: First hardware call instruction selection function.

[0070] The code may contain one or more candidate hardware call instructions. The code optimization device 200 may use all candidate hardware call instructions as the first hardware call instruction, or it may use the candidate hardware call instruction with the longer execution time among the one or more candidate hardware call instructions as the first hardware call instruction.

[0071] It should be noted that the execution time of the candidate hardware call instruction refers to the execution time of the hardware unit called by the candidate hardware call instruction to execute the embedded instructions. In this embodiment of the application, for ease of explanation, the execution time of the hardware unit called by the candidate hardware call instruction is referred to as the execution time of the candidate hardware call instruction.

[0072] For example, the code optimization device 200 can obtain the execution time of each candidate hardware call instruction (that is, the execution time of the hardware unit called by the candidate hardware call instruction to execute the embedded instruction) by compiling the source code (instrumenting the candidate hardware call instruction during the compilation process) and running the compiled executable code. After obtaining the execution time of each candidate hardware call instruction, the first hardware call instruction is determined from the candidate hardware call instructions according to the execution time of each candidate hardware call instruction.

[0073] Function 2: Dependency-free instruction set analysis function.

[0074] The code optimization device 200 can analyze the dependencies between instructions in the function containing the first hardware call instruction based on source code, intermediate representation, or assembly file, and determine the dependency-free instruction set. This dependency-free instruction set has no dependency relationship with the embedded instructions of the first hardware unit called by the first hardware call instruction. In other words, the execution of this dependency-free instruction set is independent of the execution of the embedded instructions of the first hardware unit, and the execution result of the dependency-free instruction set and the execution result of the embedded instructions of the first hardware unit will not affect each other.

[0075] Since there is no dependency between the dependency-free instruction set and the embedded instructions of the first hardware unit, the dependency-free instruction set-related code is code that describes the dependency-free instruction set. This dependency-free instruction set-related code is schedulable, and its position in the code can be adjusted. For example, the dependency-free instruction set-related code can be inserted after the first hardware call instruction. In this way, while the data processing device executes the first hardware call instruction (that is, the first hardware unit executes the embedded instructions), the processing unit of the data processing device can also execute the dependency-free instruction set. The embodiments of this application do not limit the specific form of the data processing unit. The processing unit can be a processor core, a thread, or a process.

[0076] This application does not limit the specific form of the code optimization device 200. The code optimization device 200 can be a hardware device, such as a server or terminal computing device, or a software device, specifically a software system running on a hardware computing device. This application does not limit the deployment location of the code optimization device 200. The code optimization device 200 can run on a cloud computing device system (including at least one cloud computing device, such as a server), on an edge computing device system (including at least one edge computing device, such as a server or desktop computer), or on various terminal computing devices, such as laptops or personal desktop computers.

[0077] This application does not limit the deployment method of the code optimization device 200. For example, the code optimization device 200 can be an application deployed in the cloud, capable of providing cloud services to users. That is, the code optimization device 200 can be deployed in an edge computing device system or a cloud computing system, and users can obtain the cloud services from the code optimization device 200 through a client 100 deployed on the user side. The code optimization device 200 can also be deployed on a computing device close to the user. In this case, the code optimization device 200 can not only perform code optimization but also interact with the user, that is, the code optimization device 200 also has the functions of a client 100. For example, the code optimization device 200 can be an application software on the terminal computing device or a plugin for an application software on the terminal computing device.

[0078] The following is an example of a scenario in which the code optimization method in this application embodiment is applicable—the compilation scenario, such as... Figure 1B The diagram shown is a schematic representation of a code optimization scenario provided in an embodiment of this application. Figure 1B The document showcases a code optimization device 200, a burning device 300, a data processing device 400, and a collection device 500.

[0079] The functions of the code optimization device 200 and the data processing device 400 can be found in the foregoing description, and will not be repeated here.

[0080] The programming device 300 is connected to the code optimization device 200. The programming device 300 can program the binary code generated by the code optimization device 200 during code compilation into the data processing device 400. For example, during the code compilation process, the code optimization device 200 instrumentes hardware call instructions in the code to generate binary code. The code optimization device 200 can then transmit this binary code to the programming device 300, which programs it into the data processing device 400, such as into the flash memory of the processor in the data processing device 400.

[0081] After binary code is burned into the data processing device 400, the data processing device 400 can execute the binary code. Taking the data processing device 400 containing a DPU as an example, as a device with message processing function, the data processing device 400 can receive and send messages, and execute the binary code during the process of receiving or sending messages.

[0082] After the data processing device 400 executes the binary code for a set time, the collection device 500 can obtain the execution duration of the hardware call instruction from the data processing device 400. When instrumenting the hardware call instruction, duration statistics code is inserted. During the execution of the duration statistics code, the execution duration of the hardware call instruction (i.e., the execution duration of the embedded instruction executed by the hardware unit) is recorded, and the execution duration of the embedded instruction executed by the hardware unit is stored in a specified location (such as a register or memory). The collection device 500 can then obtain the number of times the function is executed and the execution duration of the hardware call instruction from this specified location.

[0083] After acquiring the execution time of the hardware call instruction, the collection device 500 transmits the execution time of the hardware call instruction to the code optimization device 200.

[0084] After obtaining the execution time of the hardware call instruction, the code optimization device 200 optimizes the code. The optimization method can be found in the foregoing description (hereinafter referred to as...). Figure 2(The embodiments will also be described, and will not be repeated here.) After the code optimization device 200 completes the code optimization, it can transmit the optimized code or scheduling report to the user, or it can directly transmit the optimized code (which can exist in the form of binary code) to the burning device.

[0085] For users, after obtaining the optimized code, if the optimized code exists in the form of an assembly file or source code, the user can analyze whether the optimized code is effective; if the optimized code exists in the form of binary code, the user can decide whether to burn the optimized code to the data processing device 400. If it is determined that it needs to be burned to the data processing device 400, the user can instruct the burning device 300 to burn the optimized code to the data processing device 400. The process of the burning device 300 burning the code will be described later and will not be detailed here. When the user obtains the scheduling report, the user can learn from the scheduling report which hardware call instructions can be optimized and which instruction sets can be scheduled.

[0086] For the programming device 300, the programming device 300 can program the optimized code into the data processing device 400. In the data processing device 400 with the optimized code programmed, the processing unit in the data processing device 400 executes the optimized code. When a hardware call instruction is executed, the hardware unit is invoked to execute the embedded instructions. If code related to a dependency-free instruction set is inserted between the hardware call instruction and its corresponding wait instruction, the processing unit is no longer in an idle state, but processes the dependency-free instruction set. That is, the hardware unit and the processing unit can run in parallel. While the hardware unit executes the embedded instructions, the processing unit executes the dependency-free instruction set, which can effectively improve the execution efficiency of the code.

[0087] This application does not limit the specific form of the programming device 300 and the collection device 500. The programming device 300 and the collection device 500 can be hardware devices, such as the programming device 300 being a device independent of the data processing device 400, and the collection device 500 being a hardware module deployed in the data processing device 400. The programming device 300 and the collection device 500 can also be software devices, such as the programming device 300 and the mobile device being software deployed in the data processing device 400.

[0088] like Figure 2 As shown, this application provides a code optimization method, which is executed by a code optimization device 200. The method includes:

[0089] Step 201: The code optimization device 200 acquires the code executed by the data processing device 400.

[0090] This application embodiment does not limit the execution method of step 201; the code can be provided by the user to the code optimization device 200. The code optimization device 200 can provide an interface for optimizing code to the user, such as providing the code optimization interface to the user through the client 100. The user can transmit the code to be optimized to the code optimization device 200 through this interface and request the code optimization device 200 to optimize the code.

[0091] The interface for optimizing code mentioned herein is essentially a function provided by the code optimization device 200 to the user. This application embodiment does not limit the way this interface is presented to the user. For example, the interface may be presented in the form of a visual interface. Figure 3 An example of a possible visual interface is shown. This interface provides a code upload screen and options to trigger code optimization. Users can upload the source code to be optimized through the code upload screen. After uploading the code, users can click the "Optimize" option to notify the code optimization device 200 to optimize the uploaded source code.

[0092] It is worth noting that the embodiments of this application do not limit the specific form of the code obtained by the code optimization device 200, such as the code to be optimized being source code, or the code to be optimized being an assembly file. Figure 3 The embodiment shown uses source code obtained by the code optimization device 200 as an example for illustration.

[0093] The code can also be transmitted from the data processing device 400 to the code optimization device 200. The code optimization device 200 is connected to the data processing device 400, and the data processing device 400 can transmit the code to the code optimization device 200.

[0094] Step 202: The code optimization device 200 compiles the source code and instrumentes one or more candidate hardware call instructions in the source code during the compilation process to obtain executable code. Candidate hardware call instructions are hardware call instructions contained in the source code. For an explanation of hardware call instructions, please refer to the foregoing description; further details will not be provided here.

[0095] Identification of candidate hardware call instructions:

[0096] During the coding process, the code optimization device 200 can identify hardware call instructions in the source code based on their characteristics. As explained above, hardware call instructions are essentially memory access instructions. Memory access instructions have a specific format, such as carrying an address range. During compilation, the source code can be compiled into an intermediate representation (IR), and memory access instructions with a specific format can be found in this intermediate representation; these access instructions are the hardware call instructions.

[0097] Generally speaking, source code may include multiple hardware call instructions. These hardware call instructions can all be used as candidate hardware call instructions, or only some of them can be used as candidate hardware call instructions.

[0098] Instrumentation of candidate hardware call instructions:

[0099] During the compilation process, instrumentation can be performed on the compiled source code. "Instrumentation" is a technique that inserts extra code into a program. This extra code is used to collect information during program runtime, such as function call counts, variable value changes, and program execution paths, without altering the program's original functional logic.

[0100] In this embodiment, the code optimization device 200 instrumentes candidate hardware call instructions, that is, inserts an additional piece of code at the candidate hardware call instruction. This code is used to calculate the execution time of the embedded instructions executed by the hardware unit called by the candidate hardware call instruction. In this embodiment, this code can be called the time-counting code.

[0101] For example, the code optimization device 200 can encapsulate hardware counters and the performance monitoring unit (PMU) usage methods into functions. During compilation, these functions are inserted before candidate hardware call instructions and after the corresponding wait instructions.

[0102] The PMU is a hardware unit integrated within the data processing device 400. The PMU is primarily used to monitor various performance-related events during processor operation. The PMU records the occurrence or state changes of performance-related events by calling counters. These performance-related events include instruction execution counts, cache hit / miss rates, and branch prediction success rates.

[0103] In this embodiment, the PMU usage method refers to the code that triggers the PMU to monitor performance-related events. The hardware counter is a counter that the PMU needs to call when monitoring performance-related events. The hardware counter can be some circuit logic with counting function pre-configured in the processor. During operation, the function formed by the hardware counter and the PMU usage method can trigger the PMU to call the hardware counter for timing.

[0104] This function is inserted before the candidate hardware call instruction and after the corresponding wait instruction. During code execution, the function preceding the candidate hardware call instruction is executed first, and its timing is started by a hardware counter. After the candidate hardware call instruction completes, the function following the wait instruction is executed, and its timing is also started by a hardware counter. The time difference between these two functions, recorded by the hardware counter, is the execution duration of the candidate hardware call instruction.

[0105] In practical applications, the basic unit of time for timing can be called a cycle. This function (by calling a hardware counter) can record the number of cycles, and the number of cycles represents time. The difference between the number of cycles recorded by the function before the candidate hardware call instruction and the function after the candidate hardware call instruction (this difference can be simply referred to as the cycle number of the candidate hardware call instruction) is the execution time of the candidate hardware call instruction. In other words, the cycle number of the candidate hardware call instruction represents the execution time of the candidate hardware call instruction; the larger the cycle number, the longer the execution time of the candidate hardware call instruction.

[0106] It should be noted that in practical applications, the code to be optimized may contain some already optimized hardware call instructions. These already optimized hardware call instructions may have code segments inserted into them, such as those optimized manually where code segments are manually inserted. For these already optimized hardware call instructions, the code optimization device 200 can still instrument them to determine whether further optimization is possible, that is, whether other code can be inserted after the hardware call instruction. Therefore, the code optimization device 200 can use these already optimized hardware call instructions as candidate hardware call instructions and instrument them accordingly. Since a code segment is already set between this type of candidate hardware call instruction and its corresponding wait instruction, the code optimization device 200 can insert functions at two positions: after the code segment already inserted in the candidate hardware call instruction and after the wait instruction corresponding to the candidate hardware call instruction. These two functions are used to determine the duration from the end of the code segment execution to the end of the candidate hardware call instruction execution. The duration determined by these two functions is essentially the execution duration of the candidate hardware call instruction after deducting the execution time of the inserted code segment during the execution of the candidate hardware call instruction. This duration is the duration during which the data processing device 400 is in a low-power waiting state for the candidate hardware call instruction to finish execution. In this embodiment, this duration can be used as the execution duration of such candidate hardware call instructions.

[0107] Step 203: During the execution of executable code, the code optimization device 200 obtains the execution time of candidate hardware call instructions, that is, the execution time of the embedded instructions executed by the hardware unit called by each candidate hardware call instruction.

[0108] By compiling the source code, machine-readable executable code can be obtained. The code optimization device 200 triggers the processor to execute the executable code. Taking the DPU as an example again, the code optimization device 200 stores the executable code in the memory within the DPU (such as the DPU's flash memory) and starts the DPU. The DPU reads the executable code from the memory and executes it.

[0109] After the executable code has run for a set time, the code optimization device 200 can acquire instrumentation information for each candidate hardware call instruction. The instrumentation information for each candidate hardware call instruction indicates its execution duration. As explained in the foregoing description of instrumentation, functions inserted into the source code (or IR) can record time. For any candidate hardware call instruction, the code optimization device 200 acquires the times recorded by functions before and after the instruction, and determines the execution duration of the candidate hardware call instruction based on the time difference between these two recorded times.

[0110] When the code optimization device 200 instrumentes candidate hardware call instructions, it can also record mapping information. This mapping information records the position of the candidate hardware call instruction in the source code and the instrumentation identifier of the candidate hardware call instruction. The instrumentation identifier of the candidate hardware call instruction is used to identify the hardware counter called by the function before the candidate hardware call instruction and after the wait instruction corresponding to the candidate hardware call instruction. Optionally, the mapping information can also record the function where the candidate hardware call instruction is located.

[0111] When the code priority device obtains the time recorded by the function before the candidate hardware call instruction and after the wait instruction corresponding to the candidate hardware call instruction, it determines the hardware counters of the function calls before the candidate hardware call instruction and after the wait instruction corresponding to the candidate hardware call instruction based on the mapping information, reads the number of cycles recorded by these two hardware counters, calculates the difference between the two number of cycles, and obtains the number of cycles (i.e., execution time) of the candidate hardware call instruction.

[0112] It should be noted that, for optimized candidate hardware call instructions in the code, when the code priority device obtains the time recorded by the function after the code segment in the candidate hardware call instruction and the function after the wait instruction corresponding to the candidate hardware call instruction, it determines the hardware counters called by the function after the code segment in the candidate hardware call instruction and the function after the wait instruction corresponding to the candidate hardware call instruction based on the mapping information, reads the number of cycles recorded by these two hardware counters, calculates the difference between the two number of cycles, and obtains the number of cycles (i.e., execution time) of the candidate hardware call instruction.

[0113] Step 204: The code optimization device 200 generates and displays a duration analysis report. The duration analysis report records the execution time of each candidate hardware call instruction, that is, the execution time of the embedded instructions executed by the hardware unit called by each candidate hardware call instruction. This step 204 is optional; that is, after obtaining the execution time of each candidate hardware call instruction, the code optimization device 200 can either execute step 204 or skip step 204 and proceed to step 205.

[0114] After obtaining the execution time of each candidate hardware call instruction, the code optimization device 200 can generate a duration analysis report based on the execution time of each candidate hardware call instruction. This duration analysis report records the execution time of each candidate hardware call instruction. Optionally, the duration analysis report can also record the position of each candidate hardware call instruction in the source code (such as the line number in the source code) and the function in which the candidate hardware call instruction is located (such as the function name).

[0115] like Figure 4The diagram shown is a schematic of the duration analysis report provided in an embodiment of this application. The duration analysis report records the line number of each candidate hardware call instruction in the source code, the cycle number of each candidate hardware call instruction, and the function name of the function to which each candidate hardware call instruction belongs.

[0116] The code optimization device 200 can display the duration analysis report to the user, enabling the user to promptly understand the execution time of each candidate hardware call instruction in the source code. This application embodiment does not limit the method in which the code optimization device 200 displays the duration analysis report. For example, the code optimization device 200 (via client 100) displays the source code in a visual interface and marks the execution time of each candidate hardware call instruction at its location in the source code. Alternatively, the code optimization device 200 (via client 100) displays the duration analysis report in text form in a visual interface. Yet another example is that the code optimization device 200 displays the duration analysis report to the user via email, SMS, or application notification (the application can be an application installed on the user's computing device).

[0117] Step 205: The code optimization device 200 determines the first hardware unit based on the execution time of the embedded instructions of each hardware unit, and the code optimization device 200 can determine the first hardware call instruction corresponding to the first hardware unit.

[0118] This application embodiment does not limit the manner in which the code optimization device 200 performs step 205. For example, the code optimization device 200 may select all candidate hardware call instructions as the first hardware call instruction from among the candidate hardware call instructions.

[0119] For example, the code optimization device 200 can determine a first hardware call instruction from the candidate hardware call instructions based on the execution time of each candidate hardware call instruction. The first hardware call instruction (or the first hardware unit) satisfies some or all of the following:

[0120] Condition 1: The execution time of the first hardware call instruction is the largest among the N candidate hardware call instructions, where N is a positive integer.

[0121] The execution time of the candidate hardware call instructions is sorted from longest to shortest, and the first hardware call instruction is among the top N candidate hardware call instructions.

[0122] From a hardware unit perspective, the execution time of each hardware unit executing the embedded instructions is sorted in descending order, with the first hardware unit being the one with the highest execution time among the top N hardware units. The execution time of the first hardware unit executing the embedded instructions is then determined by selecting the N hardware units with the longest execution times among all the hardware units executing the embedded instructions.

[0123] Condition 2: The execution time of the first hardware call instruction is greater than the execution time threshold.

[0124] From the perspective of the hardware unit, the execution time of the first hardware unit executing the embedded instructions is greater than the execution time threshold.

[0125] This application does not limit the setting method of the execution time threshold. The specific value of the execution time threshold can be an empirical value or it can be determined based on the performance of the data processing device 400 that executes the code.

[0126] As can be seen from conditions one and two, the code optimization device 200 preferentially selects the candidate hardware call instruction with a longer execution time as the first hardware call instruction. The longer execution time of the candidate hardware call instruction means that during the execution of the candidate hardware call instruction, it is necessary to wait for a long time to obtain the result of the processor's hardware feedback (such as the result of the data already stored mentioned above).

[0127] It should be noted that the two conditions mentioned above apply to the screening of unoptimized candidate hardware call instructions (i.e., candidate hardware call instructions without inserted code segments). For optimized candidate hardware call instructions (i.e., candidate hardware call instructions with inserted code segments), the code optimization device 200 can also use a similar method to determine the first hardware call instruction from the optimized candidate hardware call instructions. For example, it can determine the first hardware call instruction from the optimized candidate hardware call instructions based on their execution time. The first hardware call instruction can satisfy the above conditions, i.e., conditions one and / or conditions two. The difference lies in the specific values ​​of N and the duration threshold involved in the optimized candidate hardware call instructions and the unoptimized candidate hardware call instructions. These values ​​can be different or the same. For example, when selecting the first hardware call instruction from unoptimized candidate hardware call instructions, the duration threshold and the value of N in condition two are smaller when selecting the first hardware call instruction from optimized candidate hardware call instructions.

[0128] Step 206: The code optimization device 200 determines the set of instructions that have no dependencies on the first hardware call instruction from the function containing the first hardware call instruction. This set of instructions is referred to as the dependency-free instruction set. The execution result of the dependency-free instruction set will not affect the execution result of the first hardware call instruction, and the execution result of the first hardware call instruction will not affect the execution result of the dependency-free instruction set.

[0129] The source code is usually quite complex (the content of the source code is related to the writing style of the programmer who wrote it). Therefore, when executing step 206, the code optimization device 200 can analyze the dependencies between code segments based on other representations of the source code. For example, the code optimization device 200 can analyze the dependencies between the basic blocks in the source code based on the intermediate representation of the source code (that is, the IR formed during the compilation of the source code). Here, a basic block is a set of instructions consisting of a set of instructions executed sequentially in the function where the first hardware call instruction is located (that is, a set of instructions consisting of multiple instructions executed in sequence). The basic block has only one entry point (that is, it can only be entered from the single entry point) and one exit point (that is, there is only one branch after the basic block, and there are no multiple branches).

[0130] When analyzing dependencies between basic blocks (i.e. code segments), the code optimization device 200 can focus only on true dependencies, anti-dependencies, and output dependencies between code segments.

[0131] Among them, true dependency, anti-dependency, and output dependency are three types of dependency relationships determined based on the order of data reading and data writing during code execution.

[0132] A true dependency is a dependency formed when data is read after it has been written. For example, instruction set 1 is used to write data to address 1, and instruction set 2, which follows instruction set 1, is used to read data from address 1. The dependency between instruction set 1 and instruction set 2 is a true dependency.

[0133] Output dependency refers to the dependency formed when data is written again after it has been written. For example, instruction set 1 is used to write data to address 1, and instruction set 2, which follows instruction set 1, is used to read data from address 1. The dependency between instruction set 1 and instruction set 2 is an output dependency.

[0134] Anti-dependency refers to a dependency formed when data is written after data is read. For example, instruction set 1 is used to read data from address 1, and instruction set 2, which follows instruction set 1, is used to write data to address 1. The dependency relationship between instruction set 1 and instruction set 2 is anti-dependency.

[0135] The code optimization device 200 analyzes the dependencies between basic blocks to obtain the dependencies between various code segments. It finds that the basic block containing the first hardware call instruction contains other basic blocks that do not contain the first hardware call instruction. For ease of explanation, the basic block containing the first hardware call instruction is referred to as the first instruction set, and the basic blocks that do not contain the first hardware call instruction are referred to as the second instruction set.

[0136] For any first instruction set, a basic block is determined from the second instruction set that has no dependency on the first instruction set. This basic block is the instruction set that has no dependency on the first instruction set (referred to as the dependency-free instruction set). The dependency-free instruction set of the first instruction set has no dependency on the first instruction set itself, such as not constituting an output dependency, true dependency, or anti-dependency. Since there is no dependency between the dependency-free instruction set of the first instruction set and the first instruction set, there is also no dependency between the hardware call instructions in the first instruction set and the dependency-free instruction set.

[0137] In step 206, the code optimization device 200 can determine a plurality of first hardware call instructions, and for each first hardware call instruction, it can determine its corresponding dependency-free instruction set. After that, the code optimization device 200 can execute step 207.

[0138] Step 207: The code optimization device 200 inserts the code related to the dependency-free instruction set after the first hardware call instruction to obtain optimized code, so that when the optimized code is executed, the dependency-free instruction set is executed by the processing unit of the data processing device 400 at the same time as the first hardware unit is called to execute the embedded instruction.

[0139] Dependency-free instruction set-related code is the code within the code that describes the dependency-free instruction set. The dependency-free instruction set-related code may be the same as or different from the dependency-free instruction set itself. For example, by analyzing the IR (Instruction Reference) of the code, the dependency-free instruction set of the first hardware call instruction is determined. In the IR, the dependency-free instruction set-related code represents that dependency-free instruction set. In the source code, this dependency-free instruction set-related code is the representation of that dependency-free instruction set within the source code. Since IR or assembly files typically refer to a string of characters as an instruction, this dependency-free instruction set can represent an instruction set in the IR or assembly file that has no dependency relationship with the first hardware dependency. The representation of this dependency-free instruction set in the IR, assembly file, or source code is called the dependency-free instruction set-related code.

[0140] Since the execution result of the dependency-free instruction set of the first hardware call instruction does not affect the execution result of the first hardware call instruction itself, the code related to the dependency-free instruction set has the possibility of scheduling. This code can be inserted after the first hardware call instruction. In practical applications, the code related to the dependency-free instruction set can be inserted between the first hardware call instruction and its corresponding wait instruction. Thus, during the execution of the optimized code, while the first hardware call instruction calls the first hardware unit to execute the embedded instructions, the processing unit of the data processing device 400 can execute the dependency-free instruction set.

[0141] Through step 207, the code optimization device 200 can obtain optimized code. If the optimized code obtained by the code optimization device 200 exists in binary code form, then the code optimization device 200 can burn the optimized code into the data processing device 400. For example, the code optimization device 200 can transmit the optimized code to the burning device 300, which will then burn the optimized code into the data processing device 400. If the optimized code obtained by the code optimization device 200 exists in the form of IR, assembly file, or source code, the code optimization device 200 compiles the optimized code to generate binary code. The code optimization device 200 then burns this binary code into the data processing device 400. During the execution of the optimized code by the data processing device, when the first hardware call instruction is executed, the hardware unit executes the embedded instructions while the processing unit executes the dependency-free instruction set, reducing code execution time and improving code execution efficiency.

[0142] It should be noted that this explanation uses the example of the code optimization device 200 directly burning the optimized code (via the burning device 300) into the data processing device 400. In practical applications, the code optimization device 200 can also output the optimized code to the user (e.g., by executing subsequent step 208). After obtaining the optimized code, the user can decide whether to burn it into the data processing device 400. If the user decides to burn the optimized code into the data processing device 400, the user can interact with the burning device 300 to instruct the burning device 400 to burn the optimized code into the data processing device 400; the user can also interact with the code optimization device 200 to inform it to burn the optimized code into the data processing device 400. The code optimization device 200 can then transmit the optimized code to the burning device 300, instructing the burning device 400 to burn the optimized code into the data processing device 400.

[0143] In this embodiment, after determining the dependency-free instruction set of each first hardware call instruction, the code optimization device 200 can use steps 208 and / or 209 to display code optimization suggestions to the user. It should be noted that steps 208 and 209 are optional steps. The purpose of these two steps is to enable the user to understand the optimizable first hardware call instructions and the optimization scheme (i.e., the dependency-free instruction set whose position can be adjusted), or to enable the user to know the specific content of the optimized code, so that the user can decide whether to burn the optimized code to the data processing device 400.

[0144] Step 208: The code optimization device 200 generates and displays a scheduling report, which indicates the first hardware call instruction and its dependency-free instruction set (or related code). Optionally, the scheduling report may also include the execution time of the first hardware call instruction.

[0145] When multiple first hardware call instructions exist, the scheduling report indicates the multiple first hardware call instructions and the dependency-free instruction set (or related code for the dependency-free instruction set) for each first hardware call instruction. Specifically, when recording the dependency-free instruction set for each first hardware call instruction, the scheduling report can record the location of the related code in the source code (e.g., line number) or the location of the dependency-free instruction set in the intermediate representation or assembly file.

[0146] The code optimization device 200 informs the user through a scheduling report which code segments (i.e., dependent code segments) can be inserted into the first hardware call instruction. When optimizing the source code later, the user can choose which first hardware scheduling instructions to optimize. For example, the user can inform the code optimization device 200 of the first hardware call instruction that needs to be optimized. After learning the first hardware survey instruction that needs to be optimized, the code optimization device 200 inserts the relevant code of the dependent instruction set of the first hardware survey instruction after the first hardware call instruction to optimize the code.

[0147] It should be noted that in practical applications, when the code optimization device 200 executes step 206, for one or more first hardware call instructions, the first hardware call instruction may not have a dependency-free instruction set. In this case, the first hardware call instruction that does not have a dependency-free instruction set may not be recorded in the scheduling report.

[0148] Step 209: The code optimization device 200 outputs the optimized code to the user.

[0149] When executing step 208, the code optimization device 200 can use the source code as a base and insert the dependent instruction set-related code into the source code after the first hardware call instruction to obtain the optimized code. Thus, the optimized code is the optimized source code. Furthermore, when outputting the optimized code, the device can also annotate the dependent instruction set-related code inserted into the first hardware call instruction; optionally, it can also annotate the position of this dependent instruction set-related code within the source code.

[0150] When executing step 208, the code optimization device 200 can also use an intermediate representation as a basis, inserting the dependent instruction set-related code into the intermediate representation after the first hardware call instruction to obtain the optimized code. Thus, the optimized code becomes the optimized intermediate representation. Furthermore, when outputting the optimized code, the dependent instruction set-related code inserted into the first hardware call instruction can be marked in the optimized code; optionally, the location of the dependent instruction set-related code in the source code can also be marked.

[0151] When executing step 208, the code optimization device 200 can also use an assembly file as a base, inserting the code related to the dependency-free instruction set after the first hardware call instruction into the assembly file to obtain the optimized code. Thus, the optimized code becomes an optimized assembly file. Furthermore, when outputting the optimized code, the device can also annotate the code containing the dependency-free instruction set-related code inserted into the first hardware call instruction; optionally, it can also annotate the location of this dependency-free instruction set-related code within the source code.

[0152] Based on the same inventive concept as the method embodiments, this application also provides a code optimization apparatus for executing the method executed by the code optimization apparatus 200 in the above method embodiments, optimizing the code executed by the data processing device. The data processing device includes at least one hardware unit, and the hardware unit embeds execution instructions. For example... Figure 5 As shown, the code optimization device 500 includes an instruction determination module 501 and an optimization module 502. In the code optimization device 500, the modules are connected through a communication path.

[0153] The instruction determination module 501 is used to determine the instruction set in the function where the first hardware call instruction corresponding to the first hardware unit is located that has no dependency relationship with the embedded instructions executed by the first hardware unit.

[0154] The optimization module 502 is used to insert instruction set related code after the first hardware call instruction to obtain optimized code, so that when the optimized code is executed, the instruction set is executed by the processing unit of the data processing device at the same time the first hardware unit is called to execute the embedded instruction.

[0155] As one possible implementation, the code optimization device 500 further includes a duration determination module 503, which determines the execution duration of the embedded instructions executed by the hardware units called by at least one hardware call instruction in the code; and determines the first hardware unit based on the execution duration of each hardware unit.

[0156] As one possible implementation, the first hardware call instruction corresponds to a wait instruction. The wait instruction is used to prevent the processing unit from continuing to execute code. When optimizing the code, the instruction set-related code is inserted between the first hardware call instruction and the wait instruction, so that the processing unit executes the wait instruction after executing the instruction set.

[0157] As one possible implementation, the duration determination module 503 determines the execution duration of any hardware unit and inserts duration statistics code at the hardware call instruction corresponding to that hardware unit. The duration statistics code calculates the execution duration of the embedded instructions by the hardware unit during execution of the embedded instructions, as instructed by the hardware call instruction.

[0158] As one possible implementation, the first hardware unit satisfies some or all of the following:

[0159] Condition 1: The execution time of the first hardware unit executing the embedded instructions is the largest among the N hardware units whose execution times of the embedded instructions are the longest, where N is a positive integer.

[0160] Condition 2: The execution time of the embedded instructions executed by the first hardware unit is greater than the execution time threshold.

[0161] As one possible implementation, the code optimization device 500 also includes an output module 504, which can output optimized code to the user. The optimized code is presented in part or all of the following forms: intermediate representations formed during compilation, source code, and assembly files.

[0162] As one possible implementation, the output module 504 can also display the first hardware call instruction and instruction set related code to the user.

[0163] As one possible implementation, the output module 504 can also generate and display a duration analysis report to the user, which records the execution time of the embedded instructions executed by the hardware units called by at least one hardware call instruction.

[0164] The module division in this application embodiment is illustrative and only represents one logical functional division. In actual implementation, other division methods may be used. Furthermore, the functional modules in the various embodiments of this application can be integrated into a single processor, exist as separate physical entities, or be integrated into a single module. The integrated modules described above can be implemented in hardware or as software functional modules.

[0165] If the integrated module is implemented as a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a terminal device (which may be a personal computer, mobile phone, or network device, etc.) or processor to execute all or part of the steps of the methods in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0166] This application also provides, for example Figure 6 The computing device 600 shown includes a bus 601, a processor 602, a communication interface 603, and a memory 604. The processor 602, the memory 604, and the communication interface 603 communicate with each other via the bus 601.

[0167] The processor 602 can be a central processing unit (CPU) or other specific integrated circuits. The processor 602 can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.

[0168] The memory 604 can typically be dynamic random access memory (DRAM). Besides DRAM, the memory 604 can also be other types of random access memory, such as static random access memory (SRAM) or storage class memory (SCM). Additionally, the memory 604 can also be read-only memory (ROM). For example, read-only memory can be programmable read-only memory (PROM) or erasable programmable read-only memory (EPROM). The memory 604 can also be a dual in-line memory module (DIMM), flash memory, hard disk drive (HDD), or solid-state drive (SSD).

[0169] The memory 604 stores computer program instructions, and the processor 602 executes these computer program instructions to perform the aforementioned tasks. Figure 2 The steps performed by the code optimization device 200 in the described method. The memory 604 may also include other software modules required for running processes, such as an operating system (e.g., multiple modules in the code optimization device 500). The operating system may be LINUX. TM UNIX TM WINDOWS TM wait.

[0170] This application also provides a computing device system, the computing device system including at least one such as Figure 7 The computing device 700 shown includes a bus 701, a processor 702, a communication interface 703, and a memory 704. The processor 702, memory 704, and communication interface 703 communicate with each other via the bus 701. At least one computing device 700 in the computing device system communicates with each other via a communication path.

[0171] The specific types of processor 702 and memory 704 can be found in the relevant descriptions of processor 602 and memory 604, and will not be repeated here. Processor 702 executes the computer program instructions stored in memory 704 to perform the aforementioned tasks. Figure 2The described method includes some or all of the steps performed by the code optimization device 200. The memory may also include other software modules required for running processes, such as an operating system. The operating system may be Linux. TM UNIX TM WINDOWS TM wait.

[0172] At least one computing device 700 in the computing device system establishes communication with each other through a communication network, and each computing device 700 runs any one or any multiple modules of the code optimization device 500.

[0173] The descriptions of the processes corresponding to the above-mentioned figures each have their own emphasis. For parts of a process that are not described in detail, please refer to the relevant descriptions of other processes.

[0174] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented, in whole or in part, in the form of a computer program product. A computer program product includes computer program instructions, which, when loaded and executed on a computer, generate, in whole or in part, the product according to the embodiments of the present invention. Figure 2 The process or function described.

[0175] The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium may be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., SSD).

[0176] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.

Claims

1. A code optimization method, characterized in that, The method is used to optimize code executed by a data processing device, the data processing device including at least one hardware unit, the hardware unit embedding execution instructions, and the method comprising: Identify the set of instructions in the function containing the first hardware call instruction of the first hardware unit in the at least one hardware unit that has no dependency on the embedded instructions executed by the first hardware unit; The instruction set-related code is inserted after the first hardware call instruction to obtain optimized code, so that when the optimized code is executed, the instruction set is executed by the processing unit of the data processing device at the same time that the first hardware unit is called to execute the embedded instruction.

2. The method as described in claim 1, characterized in that, The method further includes: Determine the execution time of the embedded instructions executed by the hardware unit invoked by at least one hardware call instruction in the code; The first hardware unit is determined based on the execution time of each hardware unit.

3. The method as described in claim 1, characterized in that, The first hardware call instruction corresponds to a wait instruction, which is used to prevent the processing unit from continuing to execute the code. The step of inserting the instruction set-related code after the first hardware call instruction to obtain the optimized code includes: The instruction set-related code is inserted between the first hardware call instruction and the wait instruction, so that the processing unit executes the wait instruction after executing the instruction set.

4. The method as described in claim 2, characterized in that, The execution time of the embedded instructions executed by the hardware unit invoked by at least one hardware invocation instruction in the determination code includes: For any one of the at least one hardware call instructions, a duration statistics code is inserted at the hardware call instruction. The hardware call instruction instructs the hardware unit to execute the embedded instruction, and the duration statistics code is run to count the execution time of the embedded instruction.

5. The method as described in claim 4, characterized in that, The first hardware unit satisfies some or all of the following: The execution time of the first hardware unit executing the embedded instructions is the largest among the N hardware units whose execution times of the embedded instructions are the longest, where N is a positive integer; or The execution time of the first hardware unit executing the embedded instructions is greater than the execution time threshold.

6. The method according to any one of claims 1 to 5, characterized in that, The method further includes: Output the optimized code, which is presented in part or all of the following forms: intermediate representation (IR) formed during compilation, source code, and assembly file.

7. The method as described in claim 2, characterized in that, After determining the execution time of the embedded instructions executed by the hardware unit invoked by at least one hardware call instruction in the code, the method further includes: A duration analysis report is generated and displayed to the user, which records the execution time of each hardware unit.

8. A code optimization device, characterized in that, The apparatus is used to optimize code executed by a data processing device, the data processing device including at least one hardware unit, the hardware unit embedding execution instructions, and the apparatus comprising: The instruction determination module is used to determine the set of instructions in the function containing the first hardware call instruction corresponding to the first hardware unit in the at least one hardware unit that has no dependency on the embedded instructions executed by the first hardware unit. An optimization module is used to insert the instruction set-related code after the first hardware call instruction to obtain optimized code, so that when the optimized code is executed, the instruction set is executed by the processing unit of the data processing device at the same time the first hardware unit is called to execute the embedded instruction.

9. The apparatus as claimed in claim 8, characterized in that, The device also includes a duration determination module; The duration determination module is used to determine the execution duration of the embedded instructions executed by the hardware unit called by at least one hardware call instruction in the code; and to determine the first hardware unit based on the execution duration of each hardware unit.

10. The apparatus as claimed in claim 8, characterized in that, The first hardware call instruction corresponds to a wait instruction, which is used to prevent the processing unit from continuing to execute the code. The optimization code is specifically used for: The instruction set-related code is inserted between the first hardware call instruction and the wait instruction, so that the processing unit executes the wait instruction after executing the instruction set.

11. The apparatus as claimed in claim 9, characterized in that, The duration determination module is specifically used for: For any one of the at least one hardware call instructions, a duration statistics code is inserted at the hardware call instruction. The hardware call instruction instructs the hardware unit to execute the embedded instruction, and the duration statistics code is run to count the execution time of the embedded instruction.

12. The apparatus as claimed in claim 11, characterized in that, The first hardware unit satisfies some or all of the following: The execution time of the first hardware unit executing the embedded instructions is the largest among the N hardware units whose execution times of the embedded instructions are the longest, where N is a positive integer; or The execution time of the first hardware unit executing the embedded instructions is greater than the execution time threshold.

13. The apparatus according to any one of claims 8 to 12, characterized in that, The device further includes an output module, the output module being used for: Output the optimized code, which is presented in part or all of the following forms: intermediate representation (IR) formed during compilation, source code, and assembly file.

14. The apparatus as claimed in claim 9, characterized in that, The output module of the device is also used for: A duration analysis report is generated and displayed to the user, which records the execution time of each hardware unit.

15. A computer-readable storage medium, characterized in that, When the computer-readable storage medium is executed by a computing device, the computing device performs the method according to any one of claims 1 to 7.