Multi-core heterogeneous system and data processing method thereof

By reserved dedicated registers for kernel functions to store parameter address pointers in multi-core heterogeneous systems, and directly read parameters from the public memory on the device side, the problem of long startup delay of kernel functions is solved and data processing efficiency is improved.

CN120429262APending Publication Date: 2025-08-05STREAM COMPUTING INC
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202410156819.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-02-02
Publication Date
2025-08-05

AI Technical Summary

Technical Problem

In the prior art, multi-core heterogeneous systems have cumbersome processes in the process of transferring kernel function parameters, resulting in a long delay in starting kernel function and affecting data processing efficiency.

Method used

In a multi-core heterogeneous system, dedicated registers are reserved for kernel functions, and parameter access pointers are stored, and kernel function parameters stored in public memory on the device side are initialized through the compilation process, and parameters are read directly from public memory to avoid accessing the private stack area through the stack pointer.

Benefits of technology

The kernel function parameter transfer process is simplified, the kernel function start delay is shortened, and data processing efficiency is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120429262A_ABST
    Figure CN120429262A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a multi-core heterogeneous system and a data processing method thereof. The method comprises the steps that a plurality of kernel functions corresponding to a device end in the multi-kernel heterogeneous system are configured, a corresponding special register is reserved for each kernel function, the special registers store parameter taking address pointers, the parameter taking address pointers point to kernel function parameters stored in a public memory of the device end, and in the kernel function compiling process, the parameter taking address pointers point to the kernel function parameters stored in the public memory of the device end; according to the method, the parameter fetching address of the kernel function is initialized into the parameter fetching address pointer stored in the special register, so that when the target kernel function is started, the target parameter fetching address pointer can be determined according to the target special register corresponding to the target kernel function, then the parameter of the target kernel function is fetched, and finally the target kernel function is executed according to the fetched parameter. And obtaining a corresponding data processing result. The parameters are directly read from the public memory of the device end through the special register of the kernel function, the kernel function parameter transmission process is simplified, the starting time delay of the kernel function is shortened, and the data processing efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing, and more particularly, to a multi-core heterogeneous system and a data processing method thereof. Background Art

[0002] In the field of artificial intelligence, multi-core or many-core architectures are often used to execute kernel functions in parallel to achieve high-performance computing. How to correctly and efficiently support the deployment of heterogeneous systems on multi-core architectures has become an important problem to be solved urgently. Summary of the Invention

[0003] In view of this, embodiments of the present invention provide a multi-core heterogeneous system and a data processing method thereof to simplify the kernel function parameter passing process, shorten the startup latency of the kernel function, and thus reduce the total running time of the heterogeneous system and improve data processing efficiency.

[0004] In a first aspect, a data processing method for a multi-core heterogeneous system is provided. The method includes:

[0005] Configuring a plurality of kernel functions corresponding to the device side in the multi-core heterogeneous system;

[0006] Determining dedicated registers corresponding to each of the kernel functions, where the dedicated registers are used to store the parameter fetch address pointers of the kernel functions, and the parameter fetch address pointers point to the parameters corresponding to the kernel functions stored in the public memory of the device side;

[0007] Compiling each of the kernel functions to initialize the parameter fetch address of the kernel function as the parameter fetch address pointer stored in the dedicated register;

[0008] Starting a target kernel function and determining a target dedicated register corresponding to the target kernel function;

[0009] Obtaining the parameters of the target kernel function according to the target parameter fetch address pointer stored in the target dedicated register;

[0010] Executing the target kernel function according to the parameters of the target kernel function to obtain a corresponding data processing result.

[0011] In a second aspect, a data processing apparatus for a multi-core heterogeneous system is provided. The apparatus includes:

[0012] A configuration module for configuring a plurality of kernel functions corresponding to the device side in the multi-core heterogeneous system;

[0013] A determination module for determining dedicated registers corresponding to each of the kernel functions, where the dedicated registers are used to store the parameter fetch address pointers of the kernel functions, and the parameter fetch address pointers point to the parameters corresponding to the kernel functions stored in the public memory of the device side;

[0014] A compilation module, configured to compile each of the kernel functions to initialize the parameter fetching address of the kernel function as the parameter fetching address pointer stored in the dedicated register;

[0015] A start-up module, configured to start a target kernel function and determine the target dedicated register corresponding to the target kernel function;

[0016] A fetching module, configured to fetch the parameters of the target kernel function according to the target parameter fetching address pointer stored in the target dedicated register;

[0017] An execution module, configured to execute the target kernel function according to the parameters of the target kernel function to obtain the corresponding data processing result.

[0018] In a third aspect, a heterogeneous multi-core system is provided, the system comprising:

[0019] A host side, configured to compile the source code corresponding to the heterogeneous multi-core system to configure a plurality of kernel functions corresponding to a device side, specify dedicated registers corresponding to each of the kernel functions, and obtain kernel function parameters required for executing the kernel functions, wherein the dedicated register is used to store the parameter fetching address pointer of the kernel function, and the parameter fetching address pointer points to the kernel function corresponding parameter stored in the public memory of the device side;

[0020] A device side, configured to copy and load the kernel function parameters in the host side memory to the public memory of the device side, and execute the corresponding kernel function according to the kernel function parameters to obtain the corresponding data processing result, wherein the device side includes a plurality of cores, and each core is configured to obtain the corresponding kernel function parameters from the public memory of the device side according to the parameter fetching address pointer stored in the corresponding dedicated register, and execute the kernel function corresponding to each core according to the kernel function parameters.

[0021] In a fourth aspect, an electronic device is provided, comprising a memory and a processor, the memory being configured to store one or more computer program instructions, wherein the one or more computer program instructions are executed by the processor to implement the method as described in the first aspect above.

[0022] In a fifth aspect, a computer-readable storage medium is provided, in which a computer program is stored, and the computer program implements the method as described in the first aspect above when being executed by a processor.

[0023] In an embodiment of the present invention, multiple kernel functions corresponding to the device side in a heterogeneous multi-core system are configured, and dedicated registers corresponding to each kernel function are reserved. The dedicated registers store parameter fetching address pointers, and the parameter fetching address pointers point to the kernel function parameters stored in the public memory of the device side. During the compilation process of the kernel function, the parameter fetching address of the kernel function is initialized to the parameter fetching address pointer stored in the dedicated register. Thus, when starting the target kernel function, the target parameter fetching address pointer can be determined according to the target dedicated register corresponding to the target kernel function, and then the parameters of the target kernel function can be fetched. Finally, the target kernel function is executed according to the fetched parameters to obtain the corresponding data processing result. In the embodiment of the present invention, the required parameters are directly read from the public memory of the device side through the dedicated register corresponding to the kernel function, avoiding the cumbersome process of accessing the private stack area through the stack pointer when fetching parameters, simplifying the kernel function parameter passing process, shortening the startup delay of the kernel function, and improving the data processing efficiency. Description of the Drawings

[0024] Through the following description of the embodiments of the present invention with reference to the drawings, the above and other objects, features, and advantages of the present invention will become clearer. In the drawings:

[0025] Figure 1 It is a flowchart of the data processing method for the heterogeneous multi-core system in the embodiment of the present invention;

[0026] Figure 2 It is a flowchart of the method for configuring multiple kernel functions corresponding to the device side in the heterogeneous multi-core system in the embodiment of the present invention;

[0027] Figure 3 It is a flowchart of the method for initializing the parameter fetching address at the front end of the compiler in the embodiment of the present invention;

[0028] Figure 4 It is a schematic diagram of semantic replacement at the front end of the compiler in the embodiment of the present invention;

[0029] Figure 5 It is a flowchart of the method for initializing the parameter fetching address at the back end of the compiler in the embodiment of the present invention;

[0030] Figure 6 It is a schematic diagram of passing kernel function parameters in the heterogeneous multi-core system in the embodiment of the present invention;

[0031] Figure 7 It is a schematic diagram of the data processing device for the heterogeneous multi-core system in the embodiment of the present invention;

[0032] Figure 8 It is a schematic diagram of the heterogeneous multi-core system in the embodiment of the present invention;

[0033] Figure 9 It is a schematic diagram of the electronic device in the embodiment of the present invention. Detailed Embodiments

[0034] The present application will be described based on embodiments, but the present application is not limited to these embodiments. In the following detailed description of the present application, some specific details are described in detail. Those skilled in the art can fully understand the present application without the description of these details. In order to avoid obscuring the essence of the present application, well-known methods, processes, procedures, components, and circuits are not described in detail.

[0035] In addition, those of ordinary skill in the art should understand that the accompanying drawings provided herein are for illustrative purposes only, and the drawings are not necessarily drawn to scale.

[0036] Unless the context clearly requires otherwise, the words "including", "comprising", and the like throughout the specification of the application should be construed in an inclusive sense rather than an exclusive or exhaustive sense; that is, in the sense of "including but not limited to".

[0037] In the description of the present application, it should be understood that the terms "first", "second", etc. are only used for descriptive purposes and cannot be construed as indicating or implying relative importance. In addition, in the description of the present application, unless otherwise specified, the meaning of "plurality" is two or more.

[0038] In neural network computing, a multi-core or many-core architecture is often used to execute kernel functions in parallel to achieve high-performance computing. How to correctly and efficiently support the deployment of heterogeneous programs to a multi-core architecture has become an important issue for compilers to solve.

[0039] In the existing solutions, the compiler compiles the source code corresponding to the multi-core heterogeneous system according to the RISC-V standard ABI specification. Before the kernel function starts, the kernel arguments are defaultly stored in the private stack area (NPC Private Stack) pointed to by the stack pointer. That is, before the kernel function starts, the kernel arguments need to be copied and loaded from the host memory to the device side and stored in the public memory space on the device side, and then the kernel arguments are copied and distributed to the private stack areas corresponding to each core to wait for the corresponding core to take the kernel arguments. After the kernel function starts, each core on the device side takes the kernel arguments from the corresponding private stack area on the device side according to the stack pointer. In the above method, the kernel arguments need to be copied and passed multiple times.

[0040] The embodiments of the present invention focus on optimizing the transfer process of kernel arguments in a multi-core heterogeneous system to shorten the startup latency of the kernel function, thereby reducing the total running time of the heterogeneous system and improving the data processing efficiency.

[0041] Figure 1The flowchart of the data processing method for the multi-core heterogeneous system according to the embodiment of the present invention. As Figure 1 shown, the data processing method for the multi-core heterogeneous system includes the following steps:

[0042] Step S100, configure multiple kernel functions corresponding to the device side in the multi-core heterogeneous system.

[0043] Among them, the multi-core heterogeneous system includes two parts: a host side and a device side. One host side can correspond to one or more device sides. The host side is usually a central processing unit (CPU), and the device side is a coprocessor. A coprocessor is a processor developed and applied to assist the central processing unit to complete processing tasks that it cannot execute or execute with low efficiency or poor effect, such as a graphics processing unit (GPU), a neural processing unit (NPU), or a field programmable gate array (FPGA), etc. When the host side and the device side cooperate to process data, it is necessary to divide the corresponding calculations of the multi-core heterogeneous system between the host side and the device side. The function executed on the coprocessor is the kernel function.

[0044] Among them, for step S100, for example, it can be implemented through the steps as Figure 2 shown. Figure 2 The flowchart of the method for configuring multiple kernel functions corresponding to the device side in the multi-core heterogeneous system according to the embodiment of the present invention. As Figure 2 shown, the method for configuring multiple kernel functions corresponding to the device side in the multi-core heterogeneous system includes the following steps:

[0045] Step S110, obtain the source code corresponding to the multi-core heterogeneous system.

[0046] Specifically, the specific form of the source code can be selected according to actual needs, as long as it can meet the calculations required by the multi-core heterogeneous system.

[0047] Step S120, configure multiple kernel functions corresponding to each device side in the source code.

[0048] Specifically, a kernel function is defined and called in the source code. The method of defining and calling the kernel function is the same as the existing method. Exemplarily, a function modified by the keyword __global__ is the kernel function, that is, the kernel function is declared and defined with the _global_ symbol. When calling, the execution configuration using <<>> can specify the way the threads are scheduled to run on the GPU, and <<<grid_size, block_size>>> is used to specify the number of threads that the kernel function is to execute. Moreover, the kernel function is the call execution entry of the device-side function. The parameter type of the kernel function is related to the type of calculation performed by the coprocessor, that is, the parameter type of the kernel function is related to the type of the coprocessor. The parameter type of the kernel function supports basic data types, such as integer int or floating point float, etc. The supported basic data types are subject to the calculations of the multi-core heterogeneous system and are not limited herein.

[0049] After step S100, it includes step S200: determining the dedicated register corresponding to each kernel function, where the dedicated register is used to store the parameter fetch address pointer of the kernel function, and the parameter fetch address pointer points to the parameter corresponding to the kernel function stored in the device-side public memory.

[0050] Specifically, the dedicated register is a preset register corresponding to each kernel function, and the dedicated register can be a general-purpose register, such as GR, X5, X6, X7, etc.

[0051] In a possible implementation manner, for each core on the device side, any one general-purpose register is selected from the multiple registers corresponding to each core as the dedicated register corresponding to that core.

[0052] In a possible implementation manner, for each core on the device side, the general-purpose register with the largest free storage space among the multiple registers corresponding to each core is used as the dedicated register corresponding to that core.

[0053] Step S300, compiling each kernel function to initialize the parameter fetch address of the kernel function as the parameter fetch address pointer stored in the dedicated register.

[0054] Specifically, the compiler is used to compile the kernel function. During the compilation process, different processing methods can be adopted according to different compilation positions inside the compiler to initialize the parameter fetching address of the kernel function as the parameter fetching address pointer stored in the dedicated register. The function of the compiler is to translate the high-level program code that is easy to understand into equivalent machine instructions executable by the computer. The source code written in a high-level programming language is used as the input, and the machine instruction program executable by the computer is used as the output. The traditional compiler architecture generally adopts a separate design of the front end, optimizer, and back end. The front end mainly performs lexical, syntactic, and semantic analysis to decompose the program string stream into intermediate code recognizable by the computer language. The optimizer mainly optimizes the intermediate code and eliminates redundant calculations to make the code run faster and have a smaller volume. The back end is responsible for generating machine code executable by different architectures from the optimized intermediate code, such as Advanced RISC Machine (ARM) or X86 architecture, etc.

[0055] In a possible implementation manner, when the compiler front end performs semantic analysis, the parameter fetching address of the kernel function is initialized as the parameter fetching address pointer stored in the dedicated register through the method of semantic replacement. Specifically, for example Figure 3 as shown.

[0056] Figure 3 The flowchart of the method for initializing the parameter fetching address at the compiler front end according to the embodiment of the present invention. The method for initializing the parameter fetching address at the compiler front end includes the following steps:

[0057] Step S311, use the compiler front end to perform semantic analysis on each of the kernel functions to obtain the corresponding semantic analysis result.

[0058] Step S312, determine the variables corresponding to the kernel function parameters according to the semantic analysis result.

[0059] Step S313, initialize the variable as the parameter pointed to by the parameter fetching address pointer stored in the dedicated register.

[0060] Specifically, when the compiler front end performs semantic analysis, each variable can be associated with its usage, check whether each expression has the correct type, and translate the abstract syntax into a simpler form to facilitate the generation of machine language. After obtaining the semantic analysis result, find and determine the variables corresponding to the kernel function parameters in the semantic analysis result, and then initialize the variables corresponding to the kernel function parameters as the parameters pointed to by the parameter fetching address pointer stored in the dedicated register. Specifically, for example Figure 4 as shown.

[0061] Figure 4Schematic diagram of semantic replacement in the compiler front end according to an embodiment of the present invention. Among them, subfigure 41 shows the variables corresponding to the kernel function parameters before semantic replacement, and subfigure 42 shows the variables corresponding to the kernel function parameters after semantic replacement. Figure 4 The semantic replacement shown is taken as an example where the compiler fully complies with the Reduced Instruction Set Computer–Five (RISCV) architecture, the standard Application Binary Interface (ABI) specification, and is written in the C language.

[0062] In subfigure 41, the variables a and b corresponding to the kernel function parameters before semantic replacement are passed in the form of function parameters.

[0063] In subfigure 42, the variables a and b corresponding to the kernel function parameters after semantic replacement are obtained from the array pointed to by the parameter fetching address pointer args. The parameter fetching address pointer args is the address pointer stored in the dedicated register, and the array pointed to by args is the parameter array stored in the public memory on the device side. That is, the variables corresponding to the kernel function parameters after semantic replacement are obtained by accessing the parameter fetching address pointer in the dedicated register.

[0064] In a possible implementation manner, when downgrading the kernel function in the compiler backend, the parameter fetching address of the kernel function is initialized to the parameter fetching address pointer stored in the dedicated register, specifically as follows Figure 5 shown.

[0065] Figure 5 Flowchart of the method for initializing the parameter fetching address in the compiler backend according to an embodiment of the present invention. The method for initializing the parameter fetching address in the compiler backend includes the following steps:

[0066] Step S321, use the compiler backend to determine the original registers corresponding to each kernel function, and the parameter fetching address stored in the original registers points to the private stack area corresponding to each kernel function on the device side.

[0067] Step S322, replace the original register with the dedicated register to initialize the parameter fetching address of the kernel function to the parameter fetching address pointer stored in the dedicated register.

[0068] Specifically, when performing kernel function downgrading in the compiler backend, determine the parameters required to execute the kernel function and the original register storing the parameter fetching address. The parameter fetching address stored in the original register points to the private stack area corresponding to each kernel function on the device side. In the prior art, when kernel function parameters need to be fetched from the private stack area, it is required that after the driver copies and loads the kernel function parameters from the host side to the public memory space on the device side, it also needs to copy and distribute them to the private stack areas corresponding to each core on the device side again. In this application, the original register is replaced with the dedicated register, so that the address pointer in the original register that points to the private stack area corresponding to each core on the device side is replaced with an address pointer for fetching parameters that points to the public memory on the device side. After the replacement, the parameters of the kernel function can be directly fetched from the corresponding position in the public memory space on the device side according to the parameter fetching address pointer in the dedicated register. At this time, the driver only needs to copy and load the kernel function parameters from the host side to the public memory space on the device side.

[0069] Step S400: Start the target kernel function and determine the target dedicated register corresponding to the target kernel function.

[0070] The target kernel function is a kernel function run by any core on the device side.

[0071] In a possible implementation, before starting the target kernel function, it is also necessary to determine the compilation result corresponding to the target kernel function, and then generate an executable file according to the compilation result. The executable file is used to start the target kernel function.

[0072] Step S500: Obtain the parameters of the target kernel function according to the target parameter fetching address pointer stored in the target dedicated register.

[0073] Specifically, after receiving the target kernel function start instruction, start the target kernel function according to the executable file, determine the dedicated register corresponding to the target kernel function, and fetch the parameters required to execute the target kernel function from the public memory on the device side according to the parameter fetching address pointer stored in the dedicated register.

[0074] Step S600: Execute the target kernel function according to the parameters of the target kernel function to obtain the corresponding data processing result.

[0075] In the data processing process of this embodiment, start and execute the target kernel function to process the input data and obtain the data processing result.

[0076] The method of the embodiment of the present invention includes configuring multiple kernel functions corresponding to the device side in a heterogeneous multi-core system, and reserving corresponding dedicated registers for each kernel function. The dedicated register stores a parameter fetch address pointer, and the parameter fetch address pointer points to the kernel function parameters stored in the public memory of the device side. During the compilation of the kernel function, the parameter fetch address of the kernel function is initialized to the parameter fetch address pointer stored in the dedicated register. Thus, when starting the target kernel function, the target parameter fetch address pointer can be determined according to the target dedicated register corresponding to the target kernel function, and then the parameters of the target kernel function can be fetched. Finally, the target kernel function is executed according to the fetched parameters to obtain the corresponding data processing result. The embodiment of the present invention directly reads the required parameters from the public memory of the device side through the dedicated register corresponding to the kernel function, avoiding the cumbersome process of accessing the private stack area through the stack pointer when fetching parameters, simplifying the kernel function parameter passing process, shortening the startup delay of the kernel function, and improving the data processing efficiency.

[0077] Figure 6 It is a schematic diagram of passing kernel function parameters in the heterogeneous multi-core system of the embodiment of the present invention.

[0078] As Figure 6 shown, the steps of passing kernel function parameters in the heterogeneous multi-core system are as follows:

[0079] Step S601, the host side obtains the parameters required to run the target kernel function through human interaction and stores them in the host side memory.

[0080] Among them, the target kernel function is the kernel function run by any core on the device side.

[0081] Step S602, transfer the kernel function parameters from the host side memory to the public memory of the device side.

[0082] Specifically, the driver running on the device side obtains the parameters of the target kernel function on the device side from the host side, and then loads them into the public memory of the device side.

[0083] Step S603, after the target kernel function on the device side is started, the corresponding core reads the parameters required to execute the target kernel function from the public memory of the device side.

[0084] Specifically, after the target kernel function is started, each core on the device side reads the parameter fetch address pointer from its corresponding dedicated register, and then reads the corresponding kernel function parameters from the corresponding position in the public memory space of the device side according to the parameter fetch address pointer.

[0085] In the embodiments of the present invention, the storage positions corresponding to the cores in the public memory at the device end are stored in dedicated registers, that is, the parameter fetching address pointers. When each core at the device end runs the corresponding kernel function, it can directly fetch parameters from the public memory space at the device end, eliminating the process of distributing the kernel function parameters stored in the public memory at the device end to the private stack areas corresponding to each core. This simplifies the process of kernel function parameter passing, reduces the startup latency of the kernel function, and further reduces the total execution time of the heterogeneous program. It also reduces the design difficulty of the driver.

[0086] Figure 7 It is a schematic diagram of the data processing device of the multi-core heterogeneous system according to the embodiments of the present invention. As Figure 7 shown, the data processing device of the multi-core heterogeneous system includes:

[0087] A configuration module 701, configured to configure a plurality of kernel functions corresponding to the device end in the multi-core heterogeneous system.

[0088] A determination module 702, configured to determine the dedicated registers corresponding to each of the kernel functions, where the dedicated registers are used to store the parameter fetching address pointers of the kernel functions, and the parameter fetching address pointers point to the parameters corresponding to the kernel functions stored in the public memory at the device end.

[0089] A compilation module 703, configured to compile each of the kernel functions to initialize the parameter fetching address of the kernel function as the parameter fetching address pointer stored in the dedicated register.

[0090] A startup module 704, configured to start a target kernel function and determine the target dedicated register corresponding to the target kernel function.

[0091] An acquisition module 705, configured to acquire the parameters of the target kernel function according to the target parameter fetching address pointer stored in the target dedicated register.

[0092] An execution module 706, configured to execute the target kernel function according to the parameters of the target kernel function to obtain the corresponding data processing result.

[0093] The device according to the embodiment of the present invention configures multiple kernel functions corresponding to the device side in the heterogeneous multi-core system, reserves corresponding dedicated registers for each kernel function, wherein the dedicated register stores the parameter fetching address pointer, and the parameter fetching address pointer points to the kernel function parameters stored in the public memory of the device side. During the compilation process of the kernel function, the parameter fetching address of the kernel function is initialized to the parameter fetching address pointer stored in the dedicated register. Thus, when starting the target kernel function, the target parameter fetching address pointer can be determined according to the target dedicated register corresponding to the target kernel function, and then the parameters of the target kernel function can be fetched. Finally, the target kernel function is executed according to the fetched parameters to obtain the corresponding data processing result. The embodiment of the present invention directly reads the required parameters from the public memory of the device side through the dedicated register corresponding to the kernel function, avoiding the cumbersome process of accessing the private stack area through the stack pointer when fetching parameters, simplifying the kernel function parameter passing process, shortening the startup delay of the kernel function, and improving the data processing efficiency.

[0094] Figure 8 It is a schematic diagram of the heterogeneous multi-core system according to the embodiment of the present invention. As Figure 8 shown, the heterogeneous multi-core system includes:

[0095] A host side 801, which is used to compile the source code corresponding to the heterogeneous multi-core system to configure multiple kernel functions corresponding to the device side, specify the dedicated registers corresponding to each of the kernel functions, and obtain the kernel function parameters required for executing the kernel function. The dedicated register is used to store the parameter fetching address pointer of the kernel function, and the parameter fetching address pointer points to the kernel function corresponding parameters stored in the public memory of the device side;

[0096] A device side 802, which is used to copy and load the kernel function parameters in the host side memory to the public memory of the device side, and execute the corresponding kernel function according to the kernel function parameters to obtain the corresponding data processing result. The device side includes multiple cores, and each core is used to obtain the corresponding kernel function parameters from the public memory of the device side according to the parameter fetching address pointer stored in the corresponding dedicated register, and execute the kernel function corresponding to each core according to the kernel function parameters.

[0097] The system according to the embodiments of the present invention configures multiple kernel functions corresponding to the device side in a heterogeneous multi-core system, reserves corresponding dedicated registers for each kernel function, where the dedicated registers store parameter fetching address pointers, and the parameter fetching address pointers point to the kernel function parameters stored in the shared memory of the device side. During the compilation of the kernel function, the parameter fetching address of the kernel function is initialized to the parameter fetching address pointer stored in the dedicated register. Thus, when starting the target kernel function, the target parameter fetching address pointer can be determined according to the target dedicated register corresponding to the target kernel function, and then the parameters of the target kernel function can be fetched. Finally, the target kernel function is executed according to the fetched parameters to obtain the corresponding data processing result. The embodiments of the present invention directly read the required parameters from the shared memory of the device side through the dedicated register corresponding to the kernel function, avoiding the cumbersome process of accessing the private stack area through the stack pointer during parameter fetching, simplifying the kernel function parameter passing process, shortening the startup latency of the kernel function, and improving the data processing efficiency.

[0098] Figure 9 is a schematic diagram of an electronic device according to an embodiment of the present invention. As Figure 9 shown, Figure 9 the electronic device shown is a general address query device, which includes a general computer hardware structure and at least includes a processor 901 and a memory 902. The processor 901 and the memory 902 are connected through a bus 903. The memory 902 is used to store instructions or programs executable by the processor 901. The processor 901 includes a main processor and a coprocessor. The main processor is the host side, and the coprocessor is the device side. Thus, the processor 901 executes the instructions stored in the memory 902 to implement the method flow of the embodiments of the present invention as described above to process data and control other devices. The bus 903 connects the above-mentioned multiple components together and also connects the above-mentioned components to a display controller 904, a display device, and an input / output (I / O) device 905. The input / output (I / O) device 905 can be a mouse, a keyboard, a modem, a network interface, a touch input device, a body sensing input device, a printer, and other devices well known in the art. Typically, the input / output (I / O) device 905 is connected to the system through an input / output (I / O) controller 906.

[0099] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a device (equipment), or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can be implemented as a computer program product on one or more computer-readable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.

[0100] This application is described by referring to the flowcharts of methods, apparatuses (devices), and computer program products according to the embodiments of this application. It should be understood that each process in the flowchart can be implemented by computer program instructions.

[0101] These computer program instructions can be stored in a computer-readable memory that can guide a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured product including an instruction device, and the instruction device implements the processes Figure 1 specified functions in one process or multiple processes.

[0102] These computer program instructions can also be provided to the processor of a general computer, a special computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the Figure 1 specified functions in one process or multiple processes.

[0103] Another embodiment of the present invention relates to a non-volatile storage medium for storing a computer-readable program, and the computer-readable program is used for a computer to execute the above-mentioned partial or all method embodiments.

[0104] In the embodiments of the present invention, by configuring multiple kernel functions corresponding to the device side in a multi-core heterogeneous system, dedicated registers are reserved for each kernel function. Among them, the dedicated register stores a parameter fetch address pointer, and the parameter fetch address pointer points to the kernel function parameters stored in the public memory of the device side. During the compilation process of the kernel function, the parameter fetch address of the kernel function is initialized to the parameter fetch address pointer stored in the dedicated register. Thus, when starting the target kernel function, the target parameter fetch address pointer can be determined according to the target dedicated register corresponding to the target kernel function, and then the target parameter fetch address pointer can be obtained, and finally the parameters of the target kernel function can be fetched, and the target kernel function can be executed according to the fetched parameters to obtain the corresponding data processing result. In the embodiments of the present invention, the required parameters are directly read from the public memory of the device side through the dedicated register corresponding to the kernel function, avoiding the cumbersome process of accessing the private stack area through the stack pointer during parameter fetching, simplifying the kernel function parameter passing process, shortening the startup delay of the kernel function, and improving the data processing efficiency.

[0105] That is, those skilled in the art can understand that all or part of the steps in the methods of the above embodiments can be completed by specifying relevant hardware through a program. The program is stored in a storage medium, including several instructions to enable a device (such as a single-chip microcomputer, chip, etc.) or a processor to execute all or part of the steps of the methods described in the embodiments of the present application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs.

[0106] The above are only the preferred embodiments of the present application and are not used to limit the present application. For those skilled in the art, various modifications and changes can be made to the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.

Claims

1. A data processing method for a multi-core heterogeneous system, characterized in that: The method comprises: Configure multiple kernel functions corresponding to the device side in a multi-core heterogeneous system; Determine a dedicated register corresponding to each kernel function, where the dedicated register is used to store a parameter address pointer of the kernel function, and the parameter address pointer points to a parameter corresponding to the kernel function stored in the public memory of the device end; Compiling each of the kernel functions to initialize the parameter address of the kernel function to the parameter address pointer stored in the dedicated register; Starting a target kernel function and determining a target dedicated register corresponding to the target kernel function; Obtaining the parameters of the target kernel function according to the target parameter address pointer stored in the target dedicated register; The target kernel function is executed according to the parameters of the target kernel function to obtain a corresponding data processing result.

2. The method according to claim 1, characterized in that The configuration of multiple kernel functions corresponding to the device side in the multi-core heterogeneous system includes: Obtaining source code corresponding to the multi-core heterogeneous system; A plurality of kernel functions corresponding to each of the device ends are configured in the source code.

3. The method according to claim 1 or 2, characterized in that The compiling of each kernel function to initialize the parameter address of the kernel function to the parameter address pointer stored in the dedicated register includes: Perform semantic analysis on each of the kernel functions using a compiler front end to obtain corresponding semantic analysis results; Determine variables corresponding to the parameters of the kernel function according to the semantic analysis result; The variable is initialized to the parameter pointed to by the parameter address pointer stored in the dedicated register.

4. The method according to claim 1 or 2, characterized in that The compiling of each kernel function to initialize the parameter address of the kernel function to the parameter address pointer stored in the dedicated register includes: Using the compiler backend to determine the original registers corresponding to the kernel functions, the parameter addresses stored in the original registers point to the private stack areas corresponding to the kernel functions on the device side; The original register is replaced by the special register to initialize the parameter address of the kernel function to the parameter address pointer stored in the special register.

5. The method according to claim 1 or 2, characterized in that Before starting the target kernel function and determining the target dedicated register corresponding to the target kernel function, the method further includes: Determining a compilation result corresponding to the target kernel function; An executable file is generated according to the compilation result, and the executable file is used to start the target kernel function.

6. The method according to any one of claims 1 to 5, characterized in that The method is performed based on the RISC-V standard ABI specification.

7. A data processing device for a multi-core heterogeneous system, characterized in that: The device comprises: Configuration module, used to configure multiple kernel functions corresponding to the device side in a multi-core heterogeneous system; A determination module, configured to determine a dedicated register corresponding to each kernel function, wherein the dedicated register is configured to store a parameter address pointer of the kernel function, and the parameter address pointer points to a parameter corresponding to the kernel function stored in a public memory on the device side; A compiling module, configured to compile each of the kernel functions to initialize the parameter address of the kernel function to the parameter address pointer stored in the dedicated register; A startup module, configured to start a target kernel function and determine a target dedicated register corresponding to the target kernel function; An acquisition module, configured to acquire the parameters of the target kernel function according to the target parameter address pointer stored in the target dedicated register; An execution module is used to execute the target kernel function according to the parameters of the target kernel function to obtain a corresponding data processing result.

8. A multi-core heterogeneous system, characterized in that: The system comprises: The host side is used to compile the source code corresponding to the multi-core heterogeneous system to configure multiple kernel functions corresponding to the device side, specify the dedicated registers corresponding to each kernel function, and obtain the kernel function parameters required to execute the kernel function, wherein the dedicated registers are used to store the parameter address pointers of the kernel function, and the parameter address pointers point to the parameters corresponding to the kernel function stored in the public memory of the device side; The device side is used to copy the kernel function parameters in the host side memory and load them to the public memory of the device side, and execute the corresponding kernel function according to the kernel function parameters to obtain the corresponding data processing results, wherein the device side includes multiple cores, each core is used to obtain the corresponding kernel function parameters from the public memory of the device side according to the parameter address pointer stored in the corresponding dedicated register, and execute the kernel function corresponding to each core according to the kernel function parameters.

9. An electronic device comprising a memory and a processor, characterized in that: The memory is configured to store one or more computer program instructions, wherein the one or more computer program instructions are executed by the processor to implement the method according to any one of claims 1 to 6.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.

Citation Information

Cited By

  • Performance optimization method and device of graphics processor set communication operator

    CN121210143A

  • A performance optimization method and device for a graphics processor set communication operator

    CN121210143B

  • Large model reasoning method and device, equipment, medium and product

    CN122412113A

  • Inference method, device, equipment, medium and product of large model

    CN122412113B

  • Kernel function debugging method and device and storage medium

    CN122432022A