Program slicing method and device and computing device cluster

Through the program slicing method, PT/CoreSight technology is used to collect and optimize system call information, which solves the problem of system calls frequently occupying CPU cycles in cloud computing environments, improves execution performance and reduces bottleneck problems.

CN120066703APending Publication Date: 2025-05-30SHENZHEN HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510015041.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-06
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

In a cloud computing environment, system calls frequently occupy CPU cycles, resulting in low execution efficiency and affecting program performance.

Method used

The program slice method is used to collect the memory address and access information of the target instruction using PT/CoreSight technology, optimize the system calls, and slice the target program with the optimized instructions as the segmentation point.

Benefits of technology

Through the program slice method, the execution performance of system calls is improved, the efficiency of processor execution of system calls is enhanced, and the bottleneck problems of DDR bandwidth and drop-off are avoided.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120066703A_ABST
    Figure CN120066703A_ABST
Patent Text Reader

Abstract

The program slicing method comprises the steps that a PC address of at least one target instruction is obtained, the PC address of the at least one target instruction is cached in a cache unit, in the process of executing a target program, a tracking module collects an accessed memory address and accessed data when the PC address of the at least one target instruction cached in the cache unit is executed through the PT / CoreSight technology, and the accessed memory address and the accessed data are stored in the cache unit. Obtaining a memory address and access information of at least one target instruction; and based on the at least one target instruction, the memory address of the at least one target instruction and the access information, slicing the target program to obtain a plurality of slices. According to the method, the memory address and the access information in the simulation process of the simulator are collected in a lightweight mode, the simulation process is prevented from being affected, and the bottleneck problem that a tracking module collects memory information corresponding to the PC address of any instruction indifferently, and consequently DDR bandwidth and disk falling occupation are large is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of cloud computing, and in particular, to a program slicing method, apparatus, and computing device cluster. Background Art

[0002] Linux system calls, as a bridge between the user space and the kernel space, allow user programs to request kernel services, including file operations, process management, memory allocation, and network communication, etc., and are extremely important functional modules in the operating system. In a cloud computing environment, system calls are particularly frequent, on average occupying 35% of the central processing unit (CPU) cycles, and can reach up to 92%. From the perspective of the time occupation ratio, the execution efficiency of system calls has a direct impact on the overall performance of the program. In most programs, the execution time occupied by system calls is a part that cannot be ignored.

[0003] Therefore, according to the system call frequency statistics of the application program, identifying the system calls that need to be focused on, and analyzing the execution instruction stream to provide a basis for CPU design optimization is the key to improving the execution efficiency of the CPU instruction stream and the program performance. This is an urgent problem that needs attention and solution. Summary of the Invention

[0004] To solve the above problems, an embodiment of the present application provides a program slicing method, which can avoid the bottleneck problems caused by large DDR bandwidth and disk occupancy. In addition, the present application also provides a program slicing apparatus and a computing device cluster corresponding to the program slicing method.

[0005] For this reason, the following technical solutions are adopted in the embodiments of the present application:

[0006] In a first aspect, an embodiment of the present application provides a program slicing method. The method is applied to a management platform, and the management platform is used to manage infrastructure. The infrastructure includes at least one node, and a tracing module is deployed on at least one node. The tracing module includes a cache unit. The method includes: obtaining the program counter (PC) addresses of at least one target instruction, and caching the PC addresses of at least one target instruction in the cache unit. The target program includes multiple instructions, and the target instruction is an instruction selected from the multiple instructions. The PC address is used to indicate the memory location for reading the instruction; during the execution of the target program, the tracing module uses the PT / CoreSight technology to collect the memory addresses and accessed data when accessing the PC addresses of at least one target instruction cached in the cache unit, so as to obtain the memory addresses and access information of at least one target instruction; based on at least one target instruction, the memory addresses of at least one target instruction, and the access information, slice the target program to obtain multiple slices.

[0007] In this embodiment, the method can configure a cache in the trace module of the PT / CoreSight technology in the emulator to cache the PC addresses of target instructions. When the emulator performs simulation, the method can use the PT / CoreSight technology to collect the memory addresses and access information when the emulator executes the PC addresses of the target instructions cached locally, which can achieve lightweight collection of the memory addresses and access information during the emulator simulation process, avoid affecting the simulation process, and avoid the bottleneck problems of large DDR bandwidth and disk occupancy caused by the trace module collecting the memory information corresponding to the PC addresses of any instructions without discrimination.

[0008] In one embodiment, before obtaining the program counter PC addresses of at least one target instruction, the method further includes: using the PT / CoreSight technology to collect multiple instructions in the target program; analyzing the multiple instructions to obtain the PC addresses of the multiple instructions; and selecting the PC addresses of at least one target instruction from the PC addresses of the multiple instructions according to user requirements or set rules.

[0009] In this embodiment, the method can use the PT / CoreSight technology to collect the instructions in the target program, which can achieve lossless collection of the instructions in the program and avoid affecting the actual execution behavior of the program.

[0010] In one embodiment, based on at least one target instruction, the memory addresses and access information of at least one target instruction, slicing the target program to obtain multiple slices, including: optimizing at least one target instruction based on the memory addresses and access information of at least one target instruction to obtain optimized instructions; and slicing the target program according to the optimized instructions to obtain multiple slices.

[0011] In this embodiment, after obtaining the memory addresses and access information of the target instructions, the method can use the memory addresses and access information to optimize the target instructions. By slicing the program based on these optimized instructions, the obtained slices will be able to improve the execution performance of the instructions and enhance the efficiency of the processor in executing these instructions.

[0012] In one embodiment, during the execution of the target program, before the trace module uses the PT / CoreSight technology to collect the memory addresses and accessed data when accessing the PC addresses of at least one target instruction cached in the execution cache unit to obtain the memory addresses and access information of at least one target instruction, the method further includes: sequentially inputting the PC addresses of at least one target instruction into the cache unit according to a set input quantity; and the set quantity is the maximum quantity of PC addresses cached by the cache unit.

[0013] In this embodiment, when configuring the cache unit, the method can set the memory size of the cache unit and limit the number of PC addresses that the cache unit can store. By setting like this, it is possible to prevent an excessive number of PC addresses from being stored in the cache unit, thereby avoiding affecting the bandwidth required for the tracking module to collect data.

[0014] In one embodiment, the target instruction is a system call.

[0015] In this embodiment, since the system call is a hot spot in most programs, the method can optimize the execution performance of the system call, and then use the optimized system call instruction as a split point to perform program slicing on the instruction stream of the target program. The obtained slices will be able to improve the execution performance of the system call and enhance the efficiency of the processor in executing the system call.

[0016] In a second aspect, an embodiment of the present application provides a program slicing device, including: a first processing module, configured to obtain the program counter PC addresses of at least one target instruction and cache the PC addresses of at least one target instruction in a cache unit. The target program includes multiple instructions, the target instruction is an instruction selected from the multiple instructions, the PC address is used to indicate the memory location for reading the instruction, and the tracking module includes a cache unit; a second processing module, configured to, during the execution of the target program, the tracking module uses the PT / CoreSight technology to collect the memory addresses and accessed data when accessing the PC addresses of at least one target instruction cached in the cache unit, and obtain the memory addresses and access information of at least one target instruction; a third processing module, configured to slice the target program based on at least one target instruction, the memory addresses and access information of at least one target instruction, and obtain multiple slices.

[0017] In one embodiment, before obtaining the PC addresses of at least one target instruction, the first processing module is further configured to use the PT / CoreSight technology to collect multiple instructions in the target program; analyze the multiple instructions to obtain the PC addresses of the multiple instructions; and select the PC addresses of at least one target instruction from the PC addresses of the multiple instructions according to user requirements or set rules.

[0018] In one embodiment, the third processing module is configured to optimize at least one target instruction based on the memory addresses and access information of at least one target instruction to obtain optimized instructions; and slice the target program according to the optimized instructions to obtain multiple slices.

[0019] In one embodiment, before obtaining the memory addresses and access information of at least one target instruction, the first processing module is further configured to sequentially input the PC addresses of at least one target instruction into the cache unit according to a set input quantity; the set quantity is the maximum quantity of PC addresses cached by the cache unit.

[0020] In one embodiment, the target instruction is a system call.

[0021] In a third aspect, an embodiment of the present application provides a computing device, including: at least one memory; at least one processor, where the processor is configured to execute instructions stored in the memory, so that the computing device executes the embodiments of all possible implementations in the first aspect.

[0022] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, including computer program instructions, when the computer program instructions are executed by a computing device, the computing device executes the embodiments of all possible implementations in the first aspect.

[0023] In a fifth aspect, an embodiment of the present application provides a computer program product including instructions, characterized in that the computer program product stores instructions, and when the instructions are executed by a computing device, the computing device implements the embodiments of all possible implementations in the first aspect.

[0024] In a sixth aspect, an embodiment of the present application provides a computing device cluster, including at least one computing device, each computing device including a processor and a memory; the processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster executes the embodiments of all possible implementations in the first aspect.

[0025] In a seventh aspect, an embodiment of the present application provides a computer-readable storage medium, including computer program instructions, when the computer program instructions are executed by a computing device cluster, the computing device cluster executes the embodiments of all possible implementations in the first aspect.

[0026] In an eighth aspect, an embodiment of the present application provides a computer program product including instructions, characterized in that the computer program product stores instructions, and when the instructions are executed by a computing device cluster, the computing device cluster implements the embodiments of all possible implementations in the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] The following briefly introduces the drawings required for the description of the embodiments or the prior art.

[0028] Figure 1 It is a schematic structural diagram of a program slicing system provided in an embodiment of the present application;

[0029] Figure 2 It is a schematic diagram of a scenario where a user uses a program slicing system provided in an embodiment of the present application;

[0030] Figure 3Flowchart of a program slicing method provided in an embodiment of the present application;

[0031] Figure 4 Structural schematic diagram of a program slicing device provided in an embodiment of the present application;

[0032] Figure 5 Structural schematic diagram of a computing device provided in an embodiment of the present application;

[0033] Figure 6 Architectural schematic diagram of a computing device cluster provided in an embodiment of the present application;

[0034] Figure 7 Another architectural schematic diagram of a computing device cluster provided in an embodiment of the present application. Detailed implementation manners

[0035] Next, the technical solutions in the embodiments of the present application will be described with reference to the accompanying drawings in the embodiments of the present application.

[0036] The term "and / or" in this document is an association relationship describing associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The symbol " / " in this document represents an "or" relationship between associated objects. For example, A / B represents A or B.

[0037] The terms "first", "second", etc. in the description and claims of this document are used to distinguish different objects, rather than to describe a specific order of objects. For example, the first response message and the second response message are used to distinguish different response messages, rather than to describe the specific order of response messages.

[0038] In the embodiments of the present application, words such as "exemplary" or "for example" are used to represent examples, illustrations, or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Exactly, using words such as "exemplary" or "for example" is intended to present related concepts in a specific manner.

[0039] In the description of the embodiments of the present application, unless otherwise specified, the meaning of "a plurality" refers to two or more. For example, a plurality of processing units refers to two or more processing units; a plurality of elements refers to two or more elements.

[0040] Before introducing the technical solutions protected by the present application, several professional terms related to the technical solutions protected by the present application are explained in advance, which are respectively:

[0041] Program slicing is a program analysis technique that can extract code fragments related to specific concerns (such as a variable, function, or statement) from the entire program. This technique analyzes the control flow and data flow of the program to generate the smallest subset of code that affects a specific variable or program point. By extracting the complete execution instruction stream related to a specific part of the program and restoring environmental variables such as the register state and values at the execution time, a file that can be executed on an emulator can be generated.

[0042] Processor trace (PT) is an extended function technology that uses specialized hardware to capture information related to software execution and has a very small impact on software performance. These information are collected in a packaged manner, and PT technology can perform control flow tracing and generate various data packets for software decoders to process. The information contained in the data packets includes timing data, program flow information (such as branch targets, branch execution conditions), and information related to the patterns triggered by the program (such as state changes, changes in the control register 3 (CR3)).

[0043] On-chip debugging and tracing (CoreSight) is a debugging and tracing technology integrated in advanced reduced instruction set computing (RISC) machines (ARM) processors, specifically for processor debugging and performance analysis. By integrating a series of hardware components inside the processor, CoreSight technology can provide in-depth monitoring and analysis capabilities to enable detailed observation and analysis of the internal state and execution information of the processor.

[0044] The program counter (PC) is a register inside the CPU, and its main function is to store the address of the instruction to be executed next. The PC register points to the memory location of the next instruction that the CPU will read, guiding the CPU to obtain instructions from where. Usually, instructions are executed sequentially, so after each instruction is executed, the PC register is updated to point to the address of the next instruction. When encountering a branch or jump instruction, the PC register is updated to a new non-sequential address, enabling the program to perform non-linear flow control. In loops and function calls, the value of the PC register changes accordingly so that the program can return to the loop start point or the start position of the function. When the CPU processes an interruption, the current PC value is saved so that after the interruption service program is executed, the CPU can return to the interrupted instruction and continue execution.

[0045] The symbol table is a data structure generated by the compiler. It contains the mapping of the names and related information of all symbols (such as variables, functions, classes, etc.) in the program, as well as the names of each symbol and their addresses (or offsets) in memory. The main role of the symbol table is to help the compiler resolve and link symbol references during the compilation process.

[0046] Next, the technical solutions provided by this application will be introduced.

[0047] Generally, in the stage of CPU simulation design, the input program or slice is usually simulated based on the complete program flow, without specifically analyzing individual system calls to obtain the local characteristics of instructions. Such analysis can help CPU designers focus on the instruction behavior characteristics unique to system calls and perform targeted optimizations. Since the execution efficiency of the emulator is very low during the simulation stage, the smaller the input fragment provided to the emulator, the better. Currently, we need an analysis method that can identify key instruction patterns and common access patterns by analyzing the system calls of the program and provide this information to the CPU chip design team. These instruction patterns can be used to generate slices of the program for simulation and iterative optimization in the CPU design stage. By controlling the program instruction slices at the tens of millions level, for example, generating 30 million instruction slices, the simulation can be completed on the CPU emulator in approximately 6000 seconds (about 1.66 hours), which will greatly improve the speed of CPU design optimization.

[0048] When executing a program on a server and collecting CPU instructions or memory information, the industry generally relies on invasive tools such as Pin or DynamoRIO, or conducts data collection in a simulation environment such as QEMU. These methods cannot achieve online lossless data collection and may affect the actual execution behavior of the software, resulting in inconsistencies with the actual program execution results.

[0049] The PT supported by X86 CPUs and the CoreSight hardware instruction collection scheme supported by ARM CPUs, although they can collect the memory access information of any instruction without discrimination, doing so will consume a large amount of memory bandwidth and storage space. Currently, the hardware collection functions of CPUs such as PT and CoreSight mainly support the collection of instruction information, and the support for the collection of detailed memory read and write information is limited or restricted, which makes the collection function quite limited.

[0050] In view of this, an embodiment of the present application provides a method for program slicing. This method can use PT / CoreSight technology to collect instructions in the target program, and can achieve lossless collection of instructions in the program, avoiding affecting the actual execution behavior of the program. This method can configure a cache in the trace module of PT / CoreSight technology in the emulator to cache the PC addresses of the target instructions. When the emulator performs simulation, this method can use PT / CoreSight technology to collect the memory addresses and access information generated when collecting the PC addresses of the target instructions cached in the local cache unit of the emulator, and can achieve lightweight collection of the memory addresses and access information during the emulator simulation process, avoiding affecting the simulation process.

[0051] After obtaining the memory addresses and access information of the target instructions, this method can optimize the target instructions using the memory addresses and access information. By performing program slicing based on these optimized instructions, the obtained slices will be able to improve the execution performance of the instructions and enhance the efficiency of the processor in executing these instructions.

[0052] Figure 1 It is a schematic structural diagram of a program slicing system provided in an embodiment of the present application. As Figure 1 shown, the program slicing system 100 can be divided into a collection module 110, an analysis module 120, an emulator 130, and a generation module 140 according to the executed functions.

[0053] The collection module 110 is used to collect instruction information in the target program when receiving instructions input by the user or according to built-in logical rules. Among them, a program is generally a set of a series of instructions, and these instructions define the logic and behavior of the program. Each instruction is a basic operation, such as arithmetic operations, logical operations, data transmission, etc. The instructions in the program are arranged in a specific order, and this order determines the execution logic of the program. The target program refers to the program running on the client used by the user. The instructions input by the user are used to instruct the program slicing system 100 to slice the target instructions in the target program.

[0054] In the embodiments of the present application, the acquisition module 110 may collect instruction information in the target program by using the PT / CoreSight technology. Taking the PT technology as an example, when the target program runs on an operating system that supports the PT technology, the acquisition module 110 may start processor tracing through software tools. The acquisition module 110 uses the PT technology to collect instruction branch information during the execution of the target program, including the type of branch (such as conditional branch, unconditional jump, etc.), the target address of the branch, and whether the branch is executed, etc. The PT technology can encode the instruction information into data packets and store them in a dynamic random access memory (DRAM). The acquisition module 110 records and stores the branch information through the PT technology, and then reads these data packets from the DRAM to perform a detailed trace of the target program.

[0055] Since the PT / CoreSight technology provides hardware-level support in data acquisition and has advantages such as high efficiency, high real-time performance, and high accuracy, compared with software-level tools such as Pin or DynamoRIO, or data acquisition in an emulation environment such as QEMU, it can achieve lossless data acquisition and will not affect the actual execution behavior of the software.

[0056] The analysis module 120 is used to analyze multiple instructions in the target program to obtain the PC addresses of at least one target instruction in the target program. The target instruction refers to a type of instruction selected by the user or filtered by preset rules, such as system calls, pure user-mode execution instructions, etc. Exemplarily, after collecting the data packets, the analysis module 120 may parse these data packets to restore each instruction of the target program. After parsing the instructions, the analysis module 120 may select the target instructions from multiple instructions according to the user's needs or preset rules. Then, the analysis module 120 may combine the processor trace data with the binary file of the application program. The binary file of the application program is used by a software decoder to reconstruct the execution flow of the program. The analysis module 120 may map the executed instructions to the binary file of the target program and determine the instruction flow of the target program according to the execution situation of the instructions, so as to obtain the PC addresses of the target instructions.

[0057] Based on the symbol table, the analysis module 120 can find out the specific function names, symbol types, address ranges, module information, offsets, and other address information of each PC address. After obtaining the address information of each PC address, the analysis module 120 filters out the PC addresses that perform specific operations from the PC addresses of at least one target instruction. The specific operation may be a read operation, a write operation, etc.

[0058] Take the target instruction as a system call as an example. The analysis module 120 first identifies the system call numbers corresponding to read and write operations. In the Linux system, these system call numbers are defined in the unistd.h or asm / unistd.h files. For example, in the x86 architecture, SYS_read and SYS_write are specific system call numbers. Subsequently, the analysis module 120 obtains the address of the system call table (sys_call_table), which is an array containing function pointers, and each pointer points to a kernel service routine. The address of the sys_call_table can be obtained through the / proc / kallsyms or System.map files. After obtaining the system call number and the address of the system call table, the analysis module 120 can locate the specific system call function through the system call table. The system call number is used as an array subscript to find the corresponding kernel service routine. Next, the analysis module 120 can use GDB or a disassembler tool (such as objdump) to further analyze and confirm the exact address and code of the system call function. Finally, the analysis module 120 can use the kallsyms_lookup_name() function to find the symbol address to obtain the PC addresses of the system calls corresponding to read and write operations.

[0059] After obtaining the specific PC address, the analysis module 120 can input the specific PC address into the emulator 130 and instruct the emulator 130 to perform corresponding operations based on the specific PC address. During the simulation process of the emulator 130, the analysis module 120 can obtain the memory addresses and access information accessed by the emulator 130. Assume that the access operation of the emulator 130 refers to a read operation or a write operation, the accessed memory address refers to the memory address from which data is read or to which data is written, and the access information refers to the data read or the data written.

[0060] The emulator 130 can be a processor managed by the program slicing system 100, such as a CPU, GPU, etc., for simulating the PC address of the received instruction. The emulator 130 can include a tracing module 131 for tracing the execution of corresponding instructions by the upper-layer operating system during the simulation process of the emulator 130 to obtain the memory addresses and access information of the system executing the instructions. An additional cache unit can be added inside the tracing module 131 to cache the PC address of the target instruction, so that the tracing module 131 only collects the memory addresses and access information generated by the PC addresses in the cache unit, avoiding the bottleneck problems of high double data rate (DDR) bandwidth and large disk occupancy caused by the tracing module 131 collecting the memory information corresponding to the PC addresses of any instructions without discrimination.

[0061] In the embodiments of the present application, the tracing module 131 can provide the PT / CoreSight technology to trace the PC address. When the emulator 130 performs simulation, the tracing module 131 can trace the PC address cached by the local cache unit of the upper operating system through the PT / CoreSight technology, so as to realize the memory address and access information when the lightweight acquisition system executes instructions, without affecting the simulation process of the system.

[0062] The cache unit can be an independent register, a random access memory (RAM), etc., or can be a specified cache space divided from the caches configured by each core in the CPU to form a cache. In the embodiments of the present application, when the emulator 130 receives the PC address input by the analysis module 120, it can write the PC address into the cache unit. During the process of simulating the PC address where the emulator 130 executes the instruction, the tracing module 131 can only trace the process of the emulator 130 executing the PC address stored in the cache unit, and obtain the memory address and access information accessed by this process.

[0063] When configuring the cache unit, the tracing module 131 can set the memory size of the cache unit and limit the number of PC addresses that the cache unit can store. By setting like this, it can prevent too many PC addresses from being stored in the cache unit, thereby avoiding affecting the bandwidth required for the tracing module 131 to collect data. Since the number of PC addresses cached by the cache unit each time is limited, the analysis module 120 sends a set input number of PC addresses to the emulator 130 each time, and stores the set input number of PC addresses in the cache unit.

[0064] The analysis module 120 can collect the memory address and access information accessed by the emulator 130, so as to obtain all the memory addresses and access information of the instruction corresponding to the PC address where the emulator 130 executes the target instruction. Optionally, the set input number of PC addresses written into the cache unit at one time is less than or equal to 1000.

[0065] When configuring the cache unit, the tracing module 131 can limit the number of PC addresses output by the cache unit. By setting like this, it can prevent the bandwidth of the PC addresses output by the cache unit from being too large, thereby avoiding affecting the bandwidth required for the tracing module 131 to collect data. When the emulator 130 performs simulation and encounters the PC address cached by the cache unit, it will export the PC address in sequence. When the number of PC addresses exported by the cache unit exceeds the set output number, the tracing module 131 can pause exporting the PC addresses in the cache unit. Subsequently, the analysis module 120 can instruct the generation module 140 to execute the program slicing function. Optionally, the set output number of PC addresses output by the cache unit is less than or equal to one billion.

[0066] The generation module 140 is used to slice at least one target program according to at least one target instruction, the memory address corresponding to at least one target instruction, and access information, to obtain multiple slices. Still taking the target instruction as a system call as an example. The generation module 140 can utilize the memory address and access information corresponding to the system call to optimize the system call, which can improve the performance and efficiency of the system or program, and at the same time reduce resource consumption. The optimization measures are as follows:

[0067] The generation module 140 can merge multiple system calls with the same memory address and access information into one, so as to reduce the total number of system calls. For example, the generation module 140 can use the writev system call to replace multiple write system calls, which can reduce the number of switches between the user mode and the kernel mode, thereby improving the system performance. In addition, the generation module 140 can optimize the system call table based on the memory address and access information of the system call. The system call table is an array of function pointers containing the kernel function addresses corresponding to each system call. Optimizing the implementation of the system call table can speed up the lookup and execution speed of system calls. The generation module 140 can utilize the memory address and access information of the system call and adopt the system call multiplexing technology to process multiple eligible system calls simultaneously. This method can improve the processing efficiency of system calls. And other optimization strategies.

[0068] The generation module 140 can use the optimized system call as the cutting point and as the entry point for program slicing. The generation module 140 determines which instruction execution results have a direct or indirect impact on the system call by analyzing the control flow and data flow of the target program, including all instructions that affect the input and output of the system call. To perform program slicing, the generation module 140 creates a control flow graph (CFG) to show all possible execution paths in the program. The CFG is the basis for program slicing. At the same time, the generation module 140 also creates a data flow graph (DFG) to show the flow of variables and values in the program to help determine the variables and values related to the system call entry point. Using data flow analysis techniques such as reaching definitions or active variables analysis, the generation module 140 identifies all instructions that may affect the system call entry point. According to the results of the control flow and data flow analysis, the generation module 140 extracts all instructions related to the system call entry point to form a slice, and removes all instructions that do not affect the system call entry point from the slice, only retaining the code necessary for understanding and modifying the program behavior related to the system call entry point.

[0069] Since system calls are hotspots in most programs, the program slicing system 100 can optimize the execution performance of system calls and then perform program slicing on the instruction stream of the target program with the instructions of the optimized system calls as the splitting points. The obtained slices will be able to improve the execution performance of system calls and enhance the efficiency of the processor in executing system calls.

[0070] It should be understood that the functional modules, functional devices, etc. involved in the above program slicing system 100 can also be implemented in software or hardware, which can be determined according to the actual situation and will not be limited here. In addition, the functional modules, functional devices, etc. involved in the above program slicing system 100 can be arranged separately or integratedly, which will not be limited here.

[0071] The above is the introduction to the program slicing system 100 provided in the embodiments of the present application. It can be understood that the above program slicing system 100 can be configured on a cloud management platform. For example, it can be deployed on at least one virtual machine or container instance so that the cloud management platform can provide program slicing services. Of course, the program slicing system 100 can also be configured on nodes other than the cloud management platform. For example, it can be deployed in at least one data center or on at least one server, which can be determined according to the actual situation and will not be limited here. Among them, the cloud management platform can provide pages related to public cloud services for users to remotely access public cloud services. In this embodiment, users can purchase the program slicing services provided by the program slicing system 100 on the cloud management platform in advance. For the convenience of understanding, the interaction form between users and the cloud management platform will be described below.

[0072] As Figure 2As shown in the figure, the interaction between the user and the cloud management platform mainly includes: the user logs in to the cloud management platform 200 through the web page of the client, selects and purchases the cloud service (i.e., the program slicing service) related to the program slicing system 100 in the cloud management platform 200. After the purchase, the user can generate the program slicing system 100 on the cloud management platform 200 based on the functions provided by the program slicing service. Among them, the cloud management platform 200 is mainly used to manage the infrastructure for running the program slicing service. Exemplarily, the infrastructure of the program slicing service may include multiple data centers set in different regions, and each data center includes multiple servers. The data center can provide basic resources for the program slicing service, such as computing resources, storage resources, etc. Therefore, when the user purchases and uses the program slicing service, the user mainly pays for the used resources. When the user uses the program slicing service, after inputting the requirements for the program slicing service through the configuration interface, application program interface (API), or the interface interacting with the user provided by the cloud management platform 200, the cloud management platform 200 can generate the program slicing service that matches the user's requirements according to the requirements input by the user (or other software / hardware, etc.).

[0073] In addition, a part of the modules in the program slicing system 100 can also be configured on the cloud side, and another part can be configured on the end side, so as to implement the program slicing service in a way of end-cloud collaboration. In addition, the program slicing system 100 can also be all configured on the end side, which can be determined according to the actual situation and will not be limited here.

[0074] The above is the introduction to the program slicing system provided by the embodiments of the present application. Next, based on the above content, the simulation method provided by the embodiments of the present application will be introduced.

[0075] Exemplarily, Figure 3 The flowchart of a program slicing method provided by the embodiments of the present application is shown. It can be understood that this program slicing method is applied to the management platform, and the management platform is used to manage the infrastructure. The infrastructure includes at least one node, and a tracing module is deployed on at least one node. The tracing module includes a cache unit. The specific implementation process of this method is as follows:

[0076] Step S301, obtain the PC address of at least one target instruction. The target instruction refers to a type of instruction selected by the user or filtered out by a preset rule, such as a system call, a pure user-mode execution instruction, etc.

[0077] Upon receiving an instruction input by a user or according to built-in logic rules, the program slicing system 100 can collect instruction information in a target program by using the PT / CoreSight technology. Since the PT / CoreSight technology provides hardware-level support in data collection and has advantages such as high efficiency, high real-time performance, and high accuracy, compared with software-level tools such as Pin or DynamoRIO, or data collection in a simulation environment such as QEMU, it can achieve lossless data collection and will not affect the actual execution behavior of the software.

[0078] Step S302, during the execution of the target program, the tracing module uses the PT / CoreSight technology to collect the memory addresses and accessed data accessed when collecting the PC addresses of at least one target instruction cached in the execution cache unit, and obtains the memory addresses and access information of at least one target instruction.

[0079] After collecting the data packets, the program slicing system 100 can parse these data packets to restore each instruction of the target program. After parsing the instructions, the program slicing system 100 can combine the processor trace data with the binary file of the application program. The binary file of the application program is used by a software decoder to reconstruct the execution flow of the program. The program slicing system 100 can select target instructions from multiple instructions according to user requirements or set rules. Then, the program slicing system 100 can map the executed instructions to the binary file of the target program and determine the instruction flow of the target program according to the execution situation of the instructions, so as to obtain the PC addresses of the target instructions. Based on the symbol table, the program slicing system 100 can find out the specific function names, symbol types, address ranges, module information, offsets and other address information of each PC address. After obtaining the address information of each PC address, the program slicing system 100 filters out the PC addresses that perform specific operations from the PC addresses of at least one target instruction. The specific operation can be a read operation, a write operation, etc.

[0080] After obtaining a specific PC address, the program slicing system 100 can input the specific PC address into the emulator and instruct the emulator to perform corresponding operations based on the specific PC address. During the simulation process of the emulator, the program slicing system 100 can obtain the memory addresses and access information accessed by the emulator.

[0081] The tracing module can provide the PT / CoreSight technology to trace the PC address. When the emulator is performing simulation, the tracing module can use the PT / CoreSight technology to trace the PC addresses cached in the local cache unit of the upper-layer operating system, so as to achieve lightweight collection of the memory addresses and access information when the system executes instructions without affecting the simulation process of the system.

[0082] When the emulator receives the PC address input by the program slicing system 100, it can write the PC address into the cache unit. During the process of the emulator executing the PC address for simulation, the tracing module can only trace the process of the emulator executing the PC address stored in the cache unit, and obtain the memory addresses and access information accessed by this process.

[0083] When configuring the cache unit, the tracing module can set the memory size of the cache unit and limit the number of PC addresses that the cache unit can store. By setting like this, it can prevent too many PC addresses from being stored in the cache unit, thus avoiding affecting the bandwidth required for the tracing module to collect data. Since the number of PC addresses cached by the cache unit each time is limited, the program slicing system 100 sends a set number of input PC addresses to the emulator each time, and stores the set number of input PC addresses in the cache unit. The program slicing system 100 can collect the memory addresses and access information accessed by the emulator, so as to obtain all the memory addresses and access information of the instructions corresponding to the PC address of the emulator executing the target instruction.

[0084] When configuring the cache unit, the tracing module can limit the number of PC addresses output by the cache unit. By setting like this, it can prevent the bandwidth of the PC addresses output by the cache unit from being too large, thus avoiding affecting the bandwidth required for the tracing module to collect data. When the emulator performs simulation and encounters the PC address cached by the cache unit, it will export the PC address in sequence. When the number of PC addresses exported from the cache unit exceeds the set output number, the tracing module can pause exporting the PC addresses in the cache unit.

[0085] Step S303, slice the target program based on at least one target instruction, the memory addresses and access information of at least one target instruction, to obtain multiple slices.

[0086] Taking the target instruction as a system call as an example. The program slicing system 100 can utilize the memory address and access information corresponding to the system call to optimize the system call, which can improve the performance and efficiency of the system or program while reducing resource consumption. The program slicing system 100 can use the optimized system call as a cut point and serve as the entry point for program slicing. The program slicing system 100 determines which instruction execution results have a direct or indirect impact on the system call by analyzing the control flow and data flow of the target program, including all instructions that affect the input and output of the system call. To perform program slicing, the program slicing system 100 creates a CFG to show all possible execution paths in the program. At the same time, the program slicing system 100 also creates a DFG to show the flow of variables and values in the program, helping to determine the variables and values related to the system call entry point. Using data flow analysis techniques such as reachable definition or live variable analysis, the program slicing system 100 identifies all instructions that may affect the system call entry point. According to the results of the control flow and data flow analysis, the program slicing system 100 extracts all instructions related to the system call entry point to form a slice, and removes all instructions that do not affect the system call entry point from the slice, only retaining the code necessary for understanding and modifying the program behavior related to the system call entry point.

[0087] Since the system call is a hot spot in most programs, the program slicing system 100 can optimize the execution performance of the system call, and then perform program slicing on the instruction stream of the target program with the instructions of the optimized system call as the split point. The obtained slice will be able to improve the execution performance of the system call and enhance the efficiency of the processor to execute the system call.

[0088] In the embodiments of the present application, after obtaining the memory address and access information of the target instruction, the method can utilize the memory address and access information to optimize the target instruction. By performing program slicing based on these optimized instructions, the obtained slice will be able to improve the execution performance of the instructions and enhance the efficiency of the processor to execute these instructions.

[0089] Based on the content described above, the embodiments of the present application provide a program slicing device 400. As Figure 4 shown, the device 400 includes:

[0090] The first processing module 410 is used to obtain the PC addresses of at least one target instruction and cache the PC addresses of the at least one target instruction in the cache unit. The target program includes multiple instructions, the target instruction is an instruction selected from the multiple instructions, the PC address is used to indicate the memory location for reading the instruction, and the tracing module includes the cache unit. The second processing module 420 is used to, during the execution of the target program, the tracing module uses the PT / CoreSight technology to collect the memory addresses and the accessed data when accessing the PC addresses of the at least one target instruction cached in the cache unit, so as to obtain the memory addresses and access information of the at least one target instruction. The third processing module 430 is used to slice the target program based on the at least one target instruction, the memory addresses of the at least one target instruction, and the access information, so as to obtain multiple slices.

[0091] In one implementation, before obtaining the PC addresses of the at least one target instruction, the first processing module 410 is further used to collect multiple instructions in the target program by using the PT / CoreSight technology; analyze the multiple instructions to obtain the PC addresses of the multiple instructions; and select the PC addresses of the at least one target instruction from the PC addresses of the multiple instructions according to user requirements or set rules.

[0092] In one implementation, the third processing module 430 is used to optimize the at least one target instruction based on the memory addresses and access information of the at least one target instruction to obtain optimized instructions; and slice the target program according to the optimized instructions to obtain multiple slices.

[0093] In one implementation, before obtaining the memory addresses and access information of the at least one target instruction, the first processing module 410 is further used to sequentially input the PC addresses of the at least one target instruction into the cache unit according to a set input quantity; the set quantity is the maximum quantity for the cache unit to cache the PC addresses.

[0094] In one implementation, the target instruction is a system call.

[0095] Among them, the first processing module 410, the second processing module 420, and the third processing module 430 can all be implemented by software or can be implemented by hardware. Exemplarily, next, taking the first processing module 410 as an example, the implementation manner of the first processing module 410 is introduced. Similarly, the implementation manners of the second processing module 420 and the third processing module 430 can refer to the implementation manner of the first processing module 410.

[0096] As an example of a software functional unit, the first processing module 410 may include code running on a computing instance. The computing instance may include at least one of a physical host (computing device), a virtual machine, and a container. Further, the computing instance may be one or more. For example, the first processing module 410 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers for running the code may be distributed in the same region or in different regions. Further, the multiple hosts / virtual machines / containers for running the code may be distributed in the same availability zone (AZ) or in different AZs, and each AZ includes one data center or multiple geographically proximate data centers. Usually, one region may include multiple AZs.

[0097] Similarly, the multiple hosts / virtual machines / containers for running the code may be distributed in the same virtual private cloud (VPC) or in multiple VPCs. Usually, one VPC is set within one region. For cross-region communication between two VPCs within the same region and between VPCs in different regions, a communication gateway needs to be set in each VPC, and the interconnection between VPCs is achieved through the communication gateway.

[0098] As an example of a hardware functional unit, the first processing module 410 may include at least one computing device, such as a server. Alternatively, the first processing module 410 may also be a device implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be implemented using a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.

[0099] The multiple computing devices included in the first processing module 410 may be distributed in the same region or in different regions. The multiple computing devices included in the first processing module 410 may be distributed in the same availability zone (AZ) or in different AZs. Similarly, the multiple computing devices included in the first processing module 410 may be distributed in the same virtual private cloud (VPC) or in multiple VPCs. Among them, the multiple computing devices may be any combination of computing devices such as servers, application-specific integrated circuits (ASICs), programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), and generic array logic (GALs).

[0100] It should be noted that in other embodiments, the first processing module 410 may be used to execute any step in the method as Figure 3 shown, the second processing module 420 may be used to execute any step in the method as Figure 3 shown, the third processing module 430 may be used to execute any step in the method as Figure 3 shown. The steps to be implemented by the first processing module 410, the second processing module 420, and the third processing module 430 can be specified as needed. The first processing module 410, the second processing module 420, and the third processing module 430 respectively implement different steps in the method as Figure 3 shown to implement all the functions of the apparatus 400.

[0101] Figure 5 This is a schematic structural diagram of a computing device provided in an embodiment of the present application. As Figure 5 shown, the computing device 500 includes a bus 510, a processor 520, a memory 530, and a communication interface 540. The processor 520, the memory 530, and the communication interface 540 communicate with each other through the bus 510. The computing device 500 may be a server, a computer, a portable notebook, a cabinet, etc. It should be understood that the present application does not limit the number of processors and memories in the computing device 500.

[0102] The bus 510 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, Figure 5 only one line is shown in, but it does not mean that there is only one bus or one type of bus. The bus 510 may include a path for transmitting information between various components (such as the processor 520, the memory 530, and the communication interface 540) of the computing device 500.

[0103] The processor 520 can be any one or more of processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).

[0104] The memory 530 can include volatile memory, such as random access memory (RAM). The memory 530 can also include non-volatile memory, such as read-only memory (ROM), flash memory, a hard disk drive (HDD), or a solid state drive (SSD).

[0105] The memory 530 stores executable program code, and the processor 520 executes the executable program code to implement the functions of the foregoing multiple modules respectively, such as the functions of the first processing module 410, the second processing module 420, and the third processing module 430, etc., so as to implement the method as Figure 3 shown. That is, the memory 530 stores instructions for executing the method as Figure 3 shown.

[0106] Alternatively, the memory 530 stores executable code, and the processor 520 executes the executable code to implement the functions of the foregoing modules respectively, so as to implement the method as Figure 3 shown. That is, the memory 530 stores instructions for executing the method as Figure 3 shown.

[0107] The communication interface 540 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement the communication between the computing device 500 and other devices or a communication network.

[0108] The embodiments of the present application further provide a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smart phone.

[0109] As Figure 6As shown, the computing device cluster includes at least one computing device 500. Instructions for executing the method as shown in Figure 3 may be stored in the memory 530 of one or more of the computing devices 500 in the computing device cluster.

[0110] In some possible implementation manners, instructions for executing parts of the method as shown in Figure 3 may also be stored separately in the memory 530 of one or more of the computing devices 500 in the computing device cluster. In other words, a combination of one or more computing devices 500 can jointly execute the instructions for executing the method as shown in Figure 3 shown.

[0111] It should be noted that the memories 530 in different computing devices 500 in the computing device cluster may store different instructions, respectively for executing partial functions of the above-mentioned first processing module 410, second processing module 420, and third processing module 430. That is, the instructions stored in the memories 530 of different computing devices 500 can implement the functions of one or more of the above-mentioned first processing module 410, second processing module 420, and third processing module 430.

[0112] In some possible implementation manners, one or more computing devices in the computing device cluster can be connected through a network. Among them, the network can be a wide area network or a local area network, etc. Figure 7 shows a possible implementation manner. As shown in Figure 7 shown, two computing devices, namely computing device 500A and computing device 500B, are connected through a network. Specifically, they are connected to the network through the communication interfaces in each computing device. In this type of possible implementation manner, the memory 530 in computing device 500A stores instructions for executing partial functions of the above-mentioned first processing module 410. At the same time, the memory 530 in computing device 500B stores instructions for executing other partial functions of the above-mentioned second processing module 420 and third processing module 430.

[0113] Figure 7 The connection manner between the computing device clusters as shown in Figure 3 shown may be considered because the method provided in this application as shown in

[0114] shown requires a large amount of data storage. Therefore, it is considered to hand over the functions implemented by other partial modules of the above-mentioned first processing module 410, second processing module 420, and third processing module 430 to computing device 500B for execution. Figure 7The functions of the computing device 500A shown can also be completed by multiple computing devices 500. Similarly, the functions of the computing device 500B can also be completed by multiple computing devices 500.

[0115] Embodiments of the present application also provide another computing device cluster. The connection relationships among the computing devices in this computing device cluster can be similarly referred to Figure 5 and Figure 6 the connection manner of the computing device cluster. The difference is that in the memory 530 of one or more computing devices 500 in this computing device cluster, there may be stored the same instructions for executing the method as Figure 3 shown.

[0116] In some possible implementation manners, in the memory 530 of one or more computing devices 500 in this computing device cluster, there may also be respectively stored partial instructions for executing the method as Figure 3 shown. In other words, the combination of one or more computing devices 500 can jointly execute the instructions for executing the method as Figure 3 shown.

[0117] It should be noted that the memories 530 in different computing devices 500 in the computing device cluster can store different instructions for executing partial functions of the computing device 500. That is, the instructions stored in the memories 530 of different computing devices 500 can implement the functions of one or more of the above-mentioned first processing module 410, second processing module 420, and third processing module 430.

[0118] Embodiments of the present application also provide a computer program product containing instructions. The computer program product can be software or a program product containing instructions that can run on a computing device or be stored in any available medium. When the computer program product runs on at least one computing device, it causes at least one computing device to execute the method as Figure 3 shown.

[0119] Embodiments of the present application also provide a computer-readable storage medium. The computer-readable storage medium can be any available medium that a computing device can store or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid-state drive), etc. The computer-readable storage medium includes instructions, and the instructions instruct the computing device to execute the method as Figure 3 shown.

[0120] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the protection scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A program slicing method, characterized in that: The method is applied to a management platform, the management platform is used to manage infrastructure, the infrastructure includes at least one node, a tracking module is deployed on the at least one node, the tracking module includes a cache unit, and the method includes: Obtaining a program counter PC address of at least one target instruction, and caching the PC address of the at least one target instruction in the cache unit, wherein the target program includes a plurality of instructions, the target instruction is an instruction selected from the plurality of instructions, and the PC address is used to indicate a memory location of the read instruction; During the execution of the target program, the tracking module uses PT / CoreSight technology to collect memory addresses and accessed data accessed when executing the PC address of the at least one target instruction cached by the cache unit, and obtains the memory address and access information of the at least one target instruction; Based on the at least one target instruction, the memory address and access information of the at least one target instruction, the target program is sliced ​​to obtain a plurality of slices.

2. The method according to claim 1, characterized in that Before obtaining the program counter PC address of at least one target instruction, the method further includes: Collecting the plurality of instructions in the target program using the PT / CoreSight technology; Analyze the multiple instructions to obtain PC addresses of the multiple instructions; According to user requirements or set rules, the PC address of the at least one target instruction is selected from the PC addresses of the multiple instructions.

3. The method according to claim 1 or 2, characterized in that: The step of slicing the target program based on the at least one target instruction, the memory address and the access information of the at least one target instruction to obtain a plurality of slices includes: Based on the memory address and access information of the at least one target instruction, optimizing the at least one target instruction to obtain an optimized instruction; The target program is sliced ​​according to the optimized instructions to obtain the multiple slices.

4. The method according to any one of claims 1 to 3, characterized in that: In the process of executing the target program, the tracking module uses PT / CoreSight technology to collect the memory address and accessed data accessed when executing the PC address of the at least one target instruction cached by the cache unit, and before obtaining the memory address and access information of the at least one target instruction, the method also includes: The PC address of the at least one target instruction is input into the cache unit in sequence according to a set input quantity; the set quantity is the maximum quantity of PC addresses cached by the cache unit.

5. The method according to any one of claims 1 to 4, characterized in that: The target instruction is a system call.

6. A program slicing device, characterized in that: include: A first processing module is used to obtain a program counter PC address of at least one target instruction, and cache the PC address of the at least one target instruction in a cache unit, the target program includes a plurality of instructions, the target instruction is an instruction selected from the plurality of instructions, the PC address is used to indicate a memory location of a read instruction, and the tracking module includes a cache unit; A second processing module is used for, during the execution of the target program, the tracking module uses PT / CoreSight technology to collect the memory address and accessed data when executing the PC address of the at least one target instruction cached by the cache unit, so as to obtain the memory address and access information of the at least one target instruction; The third processing module is used to slice the target program based on the at least one target instruction, the memory address and access information of the at least one target instruction to obtain multiple slices.

7. A computing device cluster, characterized in that: include: at least one computing device, each computing device comprising a processor and a memory; The processor of the at least one computing device is used to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster executes the method according to any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that: The method comprises computer program instructions. When the computer program instructions are executed by a computing device cluster, the computing device performs the method according to any one of claims 1 to 5.

9. A computer program product comprising instructions, characterized in that The computer program product stores instructions, which, when executed by a computing device cluster, enable the computing device cluster to implement the method according to any one of claims 1 to 5.