Method for operating information technology system, information technology system and vehicle

By dividing program code into functional blocks and processing using single-instruction multi-data technology, the space and cost challenges of integrating high-performance computing components in the vehicle are solved, achieving higher computing efficiency and performance improvements.

CN120153349APending Publication Date: 2025-06-13MERCEDES BENZ GRP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202380077621.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-11-21
Filing Date
2023-10-26
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

Integrating high-performance computing components in vehicles has challenges of space limitations and increased costs, and prior art is difficult to effectively improve the performance and efficiency of computing components.

Method used

The compiler divides the program code into functional blocks and arranges it in the data flow diagram, and uses single-instruction multi-data (SIMD) technology to allocate the functional blocks that can be executed in parallel to the same execution unit of the processor for synchronous processing.

Benefits of technology

It realizes a higher degree of program parallelization, saves processor cycle time, improves computing efficiency, and effectively improves the performance of computing components in the vehicle, while saving installation space.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120153349A_ABST
    Figure CN120153349A_ABST
Patent Text Reader

Abstract

The invention relates to a method for operating an information technology system. In order to process at least one task (1.1, 1.2), program code usable by a processor is divided into at least two functional blocks (2) by a compiler in a compilation step and arranged in a dataflow graph (3). A compiler analyzes the dataflow graph (3) in order to determine an execution order of the processor on the respective functional blocks (2). The method according to the invention is characterized in that when processing at least one task (1.1, 1.2), the compiler determines the function blocks (2) that can be executed in parallel by means of the processor, checks for each group (4) of function blocks (2) that can be executed in parallel whether a program code part consisting of at least two function blocks (2) at least needs to be executed by the processor (5), and the program code parts are distributed to the same execution unit of the processor for synchronous processing according to the single-instruction multi-data.
Need to check novelty before this filing date? Find Prior Art

Description

Field of the Invention

[0001] The present invention relates to a method for operating an information technology system, a corresponding information technology system, and a vehicle having such an information technology system, which are more specifically defined in the preamble of claim 1. Background Art

[0002] Information technology systems, such as computers, embedded systems, or similar computing devices, have become indispensable in solving various daily problems. Here, a development goal is to improve the performance of hardware components while achieving miniaturization. Thus, complex tasks can be solved, including in the context of mobile applications.

[0003] Thereby, with the increasing degree of digitization, the importance of computer systems in vehicles has also been increasing. Modern vehicles have a variety of different control units, for example, for powertrain control, navigation route calculation, or for communication connections between the vehicle and the Internet. Powerful computing hardware components are indispensable here, so that, for example, time-sensitive control signals can be calculated quickly enough to ensure the user experience. For this purpose, user inputs must be processed quickly enough and corresponding responses must be output quickly enough, etc. In addition, only limited installation space is available in the vehicle for integrating computing components. Moreover, integrating such computing devices must not lead to a disproportionate increase in the vehicle manufacturing cost. This makes the design extremely challenging for computer systems supplied to the automotive industry. In summary, there is a continuous need to improve the performance and efficiency of computing components.

[0004] For computer-aided problem-solving, a well-known method for speeding up is parallelization. Modern processors can have multiple execution units, such as computing cores, which allows so-called multi-threading, that is, different tasks can be processed simultaneously by the same processor.

[0005] Another option for improving computational efficiency is Single Instruction Multiple Data (SIMD). SIMD enables a computer system to apply the same computational operation synchronously at multiple data points. For example, for a personal computer, this can reduce the necessary number of read and write accesses to memory, as well as the number of computational operations to be performed by the processor (CPU). Accordingly, the time required for the program to process tasks can be shortened. For example, D. Nuzman et al., “VaporSIMD: Auto-vectorize once, run everywhere,” International Symposium on Code Generation and Optimization (CGO 2011), 2011, pp. 151-160, doi:10.1109 / CGO.2011.5764683 describes how a compiler transforms loops in order to encode multiple iterations of the same loop in SIMD instructions in parallel.

[0006] In addition, WO 2007 / 113369 A1 discloses a method for generating a program that can be executed in parallel. This method describes compiling the source code of a computer program into executable code, where a data flow graph is generated based on the corresponding source code and studied to investigate data dependencies. The executable program thus generated is divided into specific program packages and assigned to the respective execution units of the processor of the computing device used. In this process, the execution order is determined as late as possible in order to take into account the computational architecture of the currently used computing device. In this way, the load can be distributed to the respective execution units of the processor particularly efficiently, regardless of the diversity of variants of the computing device available for processing the program. Here too, the increase in computational speed is also based on distributing the individual program parts to multiple execution units of one or more processors, thereby parallelizing the program sequence. Summary of the Invention

[0007] The object of the present invention is to provide an improved method for operating an information technology system by means of which the performance of the information technology system can be further enhanced.

[0008] According to the present invention, this object is achieved by a method for operating an information technology system having the features of claim 1. Advantageous designs and improvements, as well as an information technology system suitable for carrying out the method according to the present invention and a vehicle having such an information technology system, are given in the dependent claims.

[0009] A method for operating an information technology system according to the present invention, wherein, in order to process / perform at least one task, in a compilation step, a compiler divides program code that can be used by a processor into at least two functional blocks and arranges them in a data flow graph, and the compiler analyzes the data flow graph to determine the execution order of the processor for each functional block. According to the present invention, the method is improved such that when processing at least one task, the compiler determines, through the processor, functional blocks that can be executed in parallel. For each group of functional blocks that can be executed in parallel, it is checked whether there is a program code portion composed of at least two functional blocks that requires the processor to execute the same instruction at least, and these program code portions are assigned to the same execution unit of the processor for synchronous processing according to single instruction multiple data.

[0010] By means of the method according to the present invention, the degree of parallelization of the program can be further enhanced. In this way, not only can a processor with multiple execution units be used to process different programs simultaneously, but also the same execution unit of the processor can process different parts of the same program within one computing cycle. Since partial operations of the same task can be executed simultaneously, the cycle time of the processor can be saved, and thus the underlying tasks, that is, the corresponding programs, can be processed more quickly.

[0011] Here, the processor can have multiple execution units, such as processor cores. The processor does not necessarily have to have multiple physical execution units. The processor can also be designed to execute multiple threads. The processor can be a central processing unit, more precisely a central processor (CPU), a graphics processing unit, more precisely a graphics processing unit (GPU), or other dedicated programmable arithmetic logic units, such as those integrated on a system on a chip (SoC).

[0012] In order to determine functional blocks that can be executed in parallel, the compiler checks the mutual dependencies of each functional block in the data flow graph. Typically, a functional block has an input interface and an output interface. The required data is provided to the functional block through the input interface, and the result generated by the corresponding functional block is provided through the output interface. If a functional block requires a specific result from another functional block, or if a functional block provides its calculation result to another functional block, then these functional blocks cannot be executed in parallel. However, if two functional blocks have no dependencies in terms of their respective input and output data, they can generally be executed in parallel.

[0013] For the functional blocks included in a loop, if these functional blocks or each loop have the same depth, and particularly preferably the same upper limit, then the method according to the present invention can be applied particularly efficiently next.

[0014] According to an advantageous improved design of the present invention, the compiler determines the number of instructions to be processed by the processor for each functional block that can be executed in parallel, and for those functional blocks that contain fewer instructions than the functional block with the largest number of instructions in the corresponding group, for the corresponding group of functional blocks that can be executed in parallel, the compiler embeds placeholder instructions in the corresponding functional blocks, so that all functional blocks in the corresponding group contain the same number of instructions. If each functional block in the functional block group that can be executed in parallel contains the same number of instructions that should be executed by the processor, then according to single instruction multiple data, the program code part can be particularly conveniently and efficiently allocated to the same execution unit of the processor. Here, "instruction" refers to the same instruction that exists in all functional blocks of the functional block group that can be executed in parallel. For example, the instruction can be an "addition" or "multiplication" instruction. If a specific program code part deviates from the instruction or is missing, for example, a "multiplication" instruction is used instead of an "addition", or the instruction line is completely missing, then a placeholder instruction should be introduced at the corresponding position of the specific functional block so that the allocation of the program code part can be performed according to single instruction multiple data.

[0015] Furthermore, in another advantageous embodiment of the method according to the present invention, if the program code portion contained in each function block in the function block group that can be executed in parallel includes a loop:

[0016] -The compiler will determine the number of iterations corresponding to the loop;

[0017] - for program code portions that can be assigned to the same execution unit of the processor according to SIMD, the compiler checks whether the function block loops containing said program code portions have different numbers of iterations; and if this is the case:

[0018] The compiler processes the various program code parts of the functional blocks simultaneously by the same execution unit of the processor according to the minimum number of iterations, and processes the program code parts of the functional blocks with a greater number of iterations sequentially / successively by the processor according to the remaining number of iterations.

[0019] If the program code portion that can be assigned to the same execution unit of the processor according to SIMD consists only of the same type of instructions, the method according to the present invention can be executed particularly conveniently and efficiently. However, the program code portion corresponding to the function block may also contain loops. If multiple function blocks have different numbers of iterations of corresponding loops, only those program code portions with the same instructions corresponding to the number of iterations of the loop with the least number of iterations in the group of function blocks that can be executed in parallel can be parallelized according to SIMD. In other words, for the program code portion of the function block with a large number of iterations, there are no corresponding instructions in the remaining function blocks that can be simultaneously assigned to the execution units of the processor. The corresponding remaining portions will then be processed sequentially / successively by the processor.

[0020] Preferably, for a processor with multi-threading capabilities, at least two different tasks are processed in parallel on at least two different execution units of the processor. This further improves the degree of parallelization. Here, it is also possible to consider applying the method according to the present invention individually for each task on each execution unit of the processor according to single instruction multiple data to allocate program code portions. That is to say, multiple tasks can be processed simultaneously on multiple execution units, and parallelization can be achieved on each execution unit according to the method according to the present invention during this process.

[0021] An information technology system for performing the foregoing method is provided according to the present invention. The information technology system can be a known computing device, such as a personal computer, an embedded system, a SoC, etc. Especially when combined with an embedded system, the method according to the present invention can be particularly advantageously applied because relatively weak hardware components are installed therein, which requires taking other measures to improve computing efficiency.

[0022] Such an information technology system is particularly preferably integrated into a vehicle according to the present invention. The vehicle can be any vehicle, such as a car, a truck, a van, a bus, etc. Here, the information technology system can be designed as a central in-vehicle computer, a control unit of a vehicle subsystem, etc. It is also possible to integrate multiple sets of such information technology systems into the vehicle. As described above, there are limitations in integrating high-performance computing components in a vehicle. By applying the method according to the present invention and integrating such an information technology system according to the present invention, it is possible to improve the computing efficiency of the installed computing device in the vehicle, and then improve its performance, while saving the existing installation space and implementing high cost performance in vehicle manufacturing. Description of the Drawings

[0023] Other advantageous designs of the method for operating an information technology system according to the present invention are also given by the embodiments described in more detail hereinafter with reference to the drawings.

[0024] Wherein:

[0025] Figure 1 As a data flow diagram, the program code divided into specific functional blocks is shown; and

[0026] Figure 2 is shown Figure 1 the virtual code of functional blocks C and D shown in, and the instructions contained in the virtual code are allocated to the same execution unit of the processor according to the method according to the present invention. Detailed Description of the Invention

[0027] Figure 1Program code for solving one or more tasks 1.1, 1.2 and executing on an information technology system is shown in the form of an abstract representation of data flow diagram 3, which is divided into specific functional blocks 2. The functional blocks 2 here correspond to sub-problems that can be calculated separately for each of the tasks 1.1, 1.2. To execute multi-threading, the processor of the information technology system can be set to allow parallel processing of multiple tasks 1.1, 1.2, i.e., multiple execution threads. The execution threads can have the same or different execution frequencies. Here, for each task 1.1, 1.2, a separate execution unit is assigned to the processor, i.e., for example, a separate processor core. This implements the first program parallelization option. For clarity, only some of the like reference objects in the drawings are provided with reference numerals.

[0028] Each of the functional blocks 2 has an input interface 2.E for reading input data and an output interface 2.A for outputting calculation results. Considering the corresponding data dependencies, the execution order of the functional blocks 2 can be derived. In an information technology system known in the prior art, the functional blocks 2 of the specific tasks 1.1 and 1.2 are executed sequentially in order. However, this wastes computing power because each of the functional blocks 2 or the part of the program code contained in the corresponding functional block 2 can usually also be executed in parallel.

[0029] To further improve the computing efficiency, the method according to the present invention is adopted, which can distribute the instructions 5 contained in each of the functional blocks 2 and shown in more detail in Figure 2 to the processor such that the components of a certain task 1.1 can also be parallelized. Here, the method according to the present invention is designed to allocate the like executions 5 of the functional blocks 2 that can be executed in parallel to the same execution unit of the processor according to single instruction multiple data (SIMD). That is, the same instruction is applied to different register regions of the same register in one computing clock cycle. That is, in this context, "like" refers to the same type of instructions, i.e., for example, all the code lines that need to be parallelized contain "addition" instructions, which is a basic prerequisite for single instruction multiple data.

[0030] For this purpose, the method according to the present invention is designed to be analyzed by the corresponding compiler Figure 1 the data flow diagram 3 shown in, and determine the execution order of the corresponding functional blocks 2 in the process. If there is no dependency between the data exchanged through the input interface 2.E and the output interface 2.A, the functional blocks 2 can be executed in parallel. For example, in Figure 1 the cases of functional blocks C and D and C and B are like this. In addition, the parallel processing of functional blocks A and H can also be achieved. For functional blocks E and F and the functional blocks M, N, O, and Q of task 1.2, parallelization cannot be achieved due to dependencies.

[0031] Next, assume that functional blocks C and D can be executed in parallel, and they then form the shown group 4 of functional blocks 2 that can be executed in parallel in Figure 2 According to the complexity of the program code, such a group 4 can also include three, four or more functional blocks 2 that can be executed in parallel.

[0032] Figure 2 As virtual code, the program code portions contained in functional blocks C and D are abstractly shown here. According to the specific design of the information technology system and the specific problems to be processed, any possible programming language can generally be used to implement the method according to the present invention.

[0033] The instruction "Mov" describes the operation of moving data from one register to another register, or from memory to a register. Accordingly, the value 42 or 30 is assigned to register r0.

[0034] ":Loop" describes the start of a loop, and "Bnzloop" describes the end of a loop, where it is assumed that r0 = 0.

[0035] The instructions "Mov" and "Add" are instructions 5 that need to be executed by the processor of the information technology system. In addition to addition, multiplication, etc. can also be requested. For example, the instruction "Add r1,[r2],+r0" describes assigning the sum of the content of register r2 and the value of r0 to register r1. During this process, the loop shown in functional block C counts down from 42 to 0. During this process, the countdown is achieved through the code line "Add r0,r0,-1".

[0036] For "Add r1,[r2],+r0" in functional block C and "Add r1,[r2],0" in functional block D, these two code lines correspond to the same instruction 5 for the processor and access the same register area. Therefore, according to the method of the present invention, they are suitable for being assigned to the same execution unit of the information technology system processor in the single instruction multiple data manner according to the present invention.

[0037] However, the loops contained in functional blocks C and D have different numbers of iterations and different numbers of instructions 5. To accommodate this situation, so-called placeholder instructions 6 are first inserted into functional block C. Next, functional blocks C and D have the same number of instructions 5. For example, the name "nop" is used as the placeholder instruction 6, and it causes the processor not to perform any calculation operations when reading in the placeholder instruction 6.

[0038] Since the number of iterations of the loops in functional blocks C and D is different, the program code portion of functional block C is first executed sequentially on the processor, and then the program code portions of functional blocks C and D are executed in parallel. Here, the loop in functional block C has 42 iterations, while the loop in functional block D has 30 iterations. In this way, the loop in functional block C is first executed sequentially 12 times, and then the program codes of functional blocks C and D are jointly executed 30 times on the same execution unit of the processor. Depending on the design of the program code and the underlying programming language, it is also possible to first execute in parallel and then sequentially.

[0039] Here, Figure 2 The program code "Add r1a,[r2a],0&r1b,[r2b],+r0" marked by the dashed line in describes the parallel program code portion that actually achieves an improvement in computational efficiency. Additionally, it is also possible to reduce the number of executions of the program code portion "Add r0,r0,-1", that is, 12 successive loops and 30 parallel loops (for a total of 42 times), instead of 42 successive times and 30 successive times (for a total of 72 times) according to the prior art.

[0040] Register r1 and register r2 are respectively divided into register regions a and b here. That is to say, the program code is designed to access the register region b of the first register and the second register according to functional block C, and to access the register region a of the first register and the second register according to the program code of functional block D.

Claims

1. A method for operating an information technology system, wherein, in order to process at least one task (1.1, 1.2), in a compilation step, a compiler divides program code that can be used by a processor into at least two functional blocks (2) and arranges the at least two functional blocks in a data flow graph (3), and the compiler analyzes the data flow graph (3) to determine the execution order of the processor for each functional block (2). It is characterized in that the compiler determines functional blocks (2) that can be executed in parallel by the processor when processing the at least one task (1.1, 1.2). For each group (4) of functional blocks (2) that can be executed in parallel, the compiler checks whether the program code portions of at least two functional blocks (2) need to execute the same instruction (5) by the processor, and allocates these program code portions to the same execution unit of the processor for synchronous processing according to single instruction multiple data.

2. The method according to claim 1, It is characterized in that for each of the functional blocks (2) that can be executed in parallel, the compiler separately determines the number of instructions (5) to be processed by the processor. For the corresponding groups (4) of functional blocks (2) that can be executed in parallel, for those functional blocks (C) whose contained number of instructions (5) is less than that of the functional block (D) with the most instructions in the corresponding group (4), the compiler embeds placeholder instructions (6) in each functional block (C) so that all functional blocks (2) of the corresponding group (4) contain the same number of instructions (5).

3. The method according to claim 1 or 2, It is characterized in that if the program code portions contained in the functional blocks (2) of the group (4) of functional blocks (2) that can be executed in parallel contain loops: - then the compiler determines the number of iterations corresponding to the loop; - for the program code portions that can be allocated to the same execution unit of the processor according to single instruction multiple data, the compiler checks whether the loops of the functional blocks (2) containing the program code portions have different numbers of iterations; and if this is the case: - the compiler simultaneously processes each program code portion of each functional block (C, D) through the same execution unit of the processor according to the minimum number of iterations, and for the functional block (C) with more iterations, processes its program code portion successively by the processor according to the remaining number of iterations.

4. The method according to any one of claims 1 to 3, It is characterized in that for a processor with multi-threading capability, at least two different tasks (1.1, 1.2) are processed in parallel on at least two different execution units of the processor.

5. An information technology system, It is characterized in that it includes means for executing the method according to any one of claims 1 to 4.

6. A vehicle, It is characterized in that it includes the information technology system according to claim 5.

Citation Information

Patent Citations

  • Parallel program generation method

    WO2007113369A1