Code processing method and device, electronic equipment and storage medium

By compiling and linking the objective function code of heterogeneous computing devices, heterogeneous program executable files are generated, which solves the problem of low computing resource utilization in heterogeneous computing devices and realizes efficient collaborative work of different types of artificial intelligence acceleration devices.

CN120045189APending Publication Date: 2025-05-27KUNLUNXIN TECHNOLOGY (BEIJING) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510174043.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-17
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

In existing heterogeneous computing devices, it is difficult to effectively utilize the computing resources of different types of artificial intelligence acceleration devices, and the threshold for writing kernel functions is high, resulting in low computing resource utilization.

Method used

By compiling the object function code of multiple artificial intelligence acceleration devices, a relocatable object file is generated and linked to the host-side relocatable object file to generate heterogeneous program executable files to realize the collaborative work of different types of artificial intelligence acceleration devices.

Benefits of technology

It improves the computing resource utilization rate of heterogeneous computing devices, reduces the difficulty of writing kernel functions, and realizes the unified startup and management of different types of artificial intelligence acceleration devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045189A_ABST
    Figure CN120045189A_ABST
Patent Text Reader

Abstract

The invention provides a code processing method, and relates to the technical field of artificial intelligence, in particular to the technical field of chips and the technical field of heterogeneous computing. According to the specific implementation scheme, target function codes of multiple pieces of artificial intelligence acceleration equipment are compiled, multiple to-be-processed target function relocatable target files are obtained, and the multiple pieces of artificial intelligence acceleration equipment correspond to the host end relocatable target files; obtaining a plurality of processed target function relocatable target files according to a plurality of device end executable files and a plurality of to-be-processed target function relocatable target files used for a plurality of artificial intelligence acceleration devices; and linking the host end relocatable target file and the plurality of processed target function relocatable target files to obtain a heterogeneous program executable file. The invention further provides a code processing device, electronic equipment and a storage medium.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, in particular to the field of chip technology and heterogeneous computing technology. More specifically, the present disclosure provides a code processing method, device, electronic device and storage medium. Background Art

[0002] With the development of artificial intelligence technology, the application of heterogeneous computing is increasing. Heterogeneous computing can use different types of processors to jointly perform computing tasks, which can give full play to the characteristics of different processors and is widely used in fields such as high-performance computing. Summary of the invention

[0003] The present disclosure provides a code processing method, apparatus, device and storage medium.

[0004] According to one aspect of the present disclosure, a code processing method is provided, the method comprising: compiling target function codes of respective multiple artificial intelligence acceleration devices to obtain multiple target function relocatable target files to be processed, wherein the multiple artificial intelligence acceleration devices correspond to host-side relocatable target files; obtaining multiple processed target function relocatable target files according to multiple device-side executable files for multiple artificial intelligence acceleration devices and multiple target function relocatable target files to be processed; linking the host-side relocatable target file and the multiple processed target function relocatable target files to obtain a heterogeneous program executable file.

[0005] According to another aspect of the present disclosure, a code processing device is provided, which includes: a first compilation module, used to compile the target function code of each of multiple artificial intelligence acceleration devices to obtain multiple target function relocatable target files to be processed, wherein the multiple artificial intelligence acceleration devices correspond to the host-side relocatable target files; an acquisition module, used to obtain multiple processed target function relocatable target files according to multiple device-side executable files for multiple artificial intelligence acceleration devices and multiple target function relocatable target files to be processed; a linking module, used to link the host-side relocatable target file and the multiple processed target function relocatable target files to obtain a heterogeneous program executable file.

[0006] According to another aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method provided according to the present disclosure.

[0007] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided. The computer instructions are used to cause a computer to execute the method provided according to the present disclosure.

[0008] According to another aspect of the present disclosure, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the method provided according to the present disclosure is implemented.

[0009] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] The accompanying drawings are used to better understand the present solution and do not constitute a limitation of the present disclosure.

[0011] Figure 1 is a schematic diagram of an exemplary system architecture to which a code processing method and apparatus can be applied according to an embodiment of the present disclosure;

[0012] Figure 2 is a flowchart of a code processing method according to an embodiment of the present disclosure;

[0013] Figure 3 is an execution flow chart of a code processing method according to an embodiment of the present disclosure;

[0014] Figure 4 is a block diagram of a code processing device according to an embodiment of the present disclosure; and

[0015] Figure 5 is a block diagram of an electronic device to which a code processing method can be applied according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0016] The following is a description of exemplary embodiments of the present disclosure in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0017] Different types of processors used for heterogeneous computing may include a central processing unit (CPU) and an artificial intelligence processor. An artificial intelligence processor may be a general purpose graphics processing unit (GPGPU), a tensor processing unit (TPU), a neural network processing unit (NPU), and other processors. A heterogeneous computing device may include a host side and a device side. The host side may include a central processing unit. The device side may include an artificial intelligence processor. In a heterogeneous computing device, the device side may be one or more. One or more device sides may be one or more artificial intelligence acceleration devices. An artificial intelligence acceleration device may also be referred to as an artificial intelligence acceleration board.

[0018] In the field of heterogeneous computing, kernel functions can be codes for computing tasks to be performed on the device side, including data processing logic that needs to be executed by the device side. Different AI accelerators have different architectures and are suitable for performing different types of tasks. With the continuous development of domain-specific architectures, the types of heterogeneous computing devices are increasing. Writing kernel functions for different types of AI accelerators and enabling multiple different types of AI accelerators to collaborate to complete specific computing tasks are important requirements for heterogeneous program writing.

[0019] Different types of AI acceleration devices have different hardware architectures. The programming methods based on different hardware architectures are also different. It is difficult to start the kernel functions of different types of AI acceleration devices in the same way, which makes it difficult to effectively improve the computing resource utilization of heterogeneous computing devices and also leads to a high threshold for writing kernel functions for different types of AI acceleration devices.

[0020] Therefore, in order to fully utilize the computing resources of heterogeneous computing devices, the present disclosure provides a code processing method, which will be described below.

[0021] Figure 1 is a schematic diagram of an exemplary system architecture to which a code testing method and apparatus can be applied according to an embodiment of the present disclosure. It should be noted that: Figure 1 What is shown is merely an example of a system architecture to which the embodiments of the present disclosure can be applied, in order to help those skilled in the art understand the technical content of the present disclosure, but it does not mean that the embodiments of the present disclosure cannot be used in other devices, systems, environments or scenarios.

[0022] like Figure 1 As shown, the system architecture 10 according to this embodiment may include a host end 110, a device end 121 ... and a device end 122. Data and signals may be transmitted between the device end 121 and the host end 110 via an interface 131. Data and signals may be transmitted between the device end 122 and the host end 110 via an interface 132. The system architecture 10 may be, for example, the system architecture of the above-mentioned heterogeneous computing device.

[0023] The host end 110 may include a first processor 111. The device end 121 may include a second processor 1211. The device end 122 may include a second processor 1221. The instruction set of the second processor 1211 may be different from the instruction set of the first processor 111. The instruction set of the second processor 1221 may be different from the instruction set of the first processor 111. The second processor 1211 and the second processor 1221 may be different types of artificial intelligence processors. The second processor 1211 may be a general-purpose graphics processing unit (GPU). The second processor 1221 may be a neural network processor (NPU). The first processor 111 may be a central processing unit. It can be understood that the number of multiple device ends may be greater than 1. Figure 1 The 2 device ends shown are just examples.

[0024] like Figure 1 As shown, the interface 131 and the interface 132 may be, for example, a peripheral component interconnect express (PCIe) interface. The interfaces 131, ..., and 132 are hardware interfaces. Via the interface 131, the first processor 111 and the second processor 121 may be connected via a high-speed serial computer expansion bus. Via the interface 132, the first processor 111 and the second processor 1221 may be connected via a high-speed serial computer expansion bus. It is understood that the interface 131 and the interface 132 may be hardware interfaces supporting various protocols, and the above-mentioned peripheral component interconnect express interface is only an example.

[0025] It should be noted that the code testing method provided in the embodiment of the present disclosure can generally be executed by the host end 110. Accordingly, the code testing device provided in the embodiment of the present disclosure can generally be set in the host end 110.

[0026] Understandably, Figure 1 The present disclosure is described by taking the example that the second processors of two device ends are of different types, but the present disclosure is not limited thereto, and among the multiple device ends, there may be two or more device ends with the same type of second processors.

[0027] It can be understood that the above describes the system architecture of the present disclosure, and the following describes the method of the present disclosure.

[0028] Figure 2 is a flowchart of a code processing method according to an embodiment of the present disclosure.

[0029] like Figure 2 As shown, the method 200 may include operations S210 to S230.

[0030] In operation S210, the target function codes of the respective multiple artificial intelligence acceleration devices are compiled to obtain multiple relocatable target files of the target functions to be processed.

[0031] In the disclosed embodiment, the target function code can be compiled using a compilation tool on the host side. The host side can be the above-mentioned host side 110. The target function relocatable target file to be processed is a relocatable target file (object file). The target function code can also be called a stub function code.

[0032] In the embodiments of the present disclosure, an artificial intelligence acceleration device may serve as the above-mentioned device end. Multiple artificial intelligence acceleration devices may include different types of artificial intelligence processors. For example, multiple artificial intelligence acceleration devices may include a first artificial intelligence acceleration device and a second artificial intelligence acceleration device. The first artificial intelligence acceleration device may include a general-purpose graphics processor. The second artificial intelligence acceleration device may include a neural network processor. The objective function code of the artificial intelligence acceleration device may indicate the processor type of the artificial intelligence acceleration device.

[0033] In the embodiment of the present disclosure, multiple artificial intelligence acceleration devices correspond to host-side relocatable target files. For example, the host-side relocatable target file involves multiple artificial intelligence acceleration devices.

[0034] In operation S220, a plurality of processed target function relocatable target files are obtained according to a plurality of device-side executable files for a plurality of artificial intelligence acceleration devices and a plurality of target function relocatable target files to be processed.

[0035] In the disclosed embodiment, each device-side executable file can be used for one artificial intelligence acceleration device.

[0036] In an embodiment of the present disclosure, a target function relocatable target file to be processed and a device-side executable file for the same artificial intelligence acceleration device may be merged to obtain a processed target function relocatable target file. The processed target function relocatable target file may include a device-side executable file.

[0037] In operation S230, the host-side relocatable object file and the plurality of processed target function relocatable object files are linked to obtain a heterogeneous program executable file.

[0038] In the embodiment of the present disclosure, the heterogeneous program executable file can be executed by a host end and multiple device ends. The heterogeneous program executable file includes a target function executable file corresponding to the target function code.

[0039] Through the disclosed embodiments, the target function code of the artificial intelligence acceleration device is compiled to obtain a target function relocatable target file to be processed that includes the type information of the artificial intelligence acceleration device. Then, the device-side executable file for the same artificial intelligence acceleration device is merged with the target function relocatable target file to be processed to obtain a processed target function relocatable target file containing the device-side executable file. Thus, the device-side executable file of the artificial intelligence acceleration device is abstractly encapsulated using the target function code. The heterogeneous program executable file may include a target function executable file, which may include the type information of the artificial intelligence acceleration device and may be executed by the host side. After executing multiple target function executable files, multiple device-side executable files may be executed separately, and the launch of multiple different types of artificial intelligence acceleration devices may be implemented in a unified manner.

[0040] It can be understood that the method of the present disclosure is described above. Figure 3 The method of the present disclosure is further described.

[0041] Figure 3 It is a schematic diagram of the execution flow of a code processing method according to an embodiment of the present disclosure.

[0042] like Figure 3 As shown, multiple host-side codes host_code30 and multiple device-side codes device_code31 can be obtained. For example, taking the multiple artificial intelligence acceleration devices as 3 as an example, the multiple artificial intelligence acceleration devices may include a first artificial intelligence acceleration device, a second artificial intelligence acceleration device, and a third artificial intelligence acceleration device. The first artificial intelligence acceleration device may include a general-purpose graphics processor. The second artificial intelligence acceleration device may include a neural network processor. The third artificial intelligence accelerator may include a tensor processor. The multiple host-side codes host_code30 may include: a first host-side code for the first artificial intelligence acceleration device, a second host-side code for the second artificial intelligence acceleration device, and a third host-side code for the third artificial intelligence acceleration device. The multiple device-side codes device_code30 may include: a first device-side code for the first artificial intelligence acceleration device, a second device-side code for the second artificial intelligence acceleration device, and a third device-side code for the third artificial intelligence acceleration device. Through the embodiments of the present disclosure, different host-side codes and device-side codes are written for different artificial intelligence acceleration devices, respectively, to achieve modular code development, and the code logic of different artificial intelligence acceleration devices can be separated for subsequent compilation. It can be understood that the host-side code can be a host-side code file, and the device-side code can be a device-side code file.

[0043] Next, operation S301 can be performed to compile multiple host-side codes for multiple artificial intelligence acceleration devices respectively to obtain a host-side relocatable target file. For example, the first host-side code, the second host-side code, and the third host-side code can be compiled using a host-side compilation tool to obtain a host-side relocatable target file file_30. It can be understood that the processor included in the host side can be a central processing unit. The host-side code for different artificial intelligence processors can be compiled by the same host-side compilation tool. The host-side code can implement functions such as initialization of the operating environment of the artificial intelligence acceleration device, storage management of the artificial intelligence acceleration device, and processing of device-side execution results.

[0044] Before executing operation S301, after executing operation S301, or when executing operation S301, operation S302 may be executed to compile multiple device-side codes for multiple artificial intelligence acceleration devices respectively to obtain multiple device-side executable files. The device-side code for the artificial intelligence acceleration device may be compiled using a device-side compilation tool for the artificial intelligence acceleration device. For example, a first device-side compilation tool for a first artificial intelligence acceleration device may be used to compile a first device-side code to obtain a first device-side executable file. A second device-side compilation tool for a second artificial intelligence acceleration device may be used to compile a second device-side code to obtain a second device-side executable file. A third device-side compilation tool for a third artificial intelligence acceleration device may be used to compile a third device-side code to obtain a third device-side executable file. It can be understood that the programming models of different artificial intelligence acceleration devices are different, and the corresponding device-side code may be compiled using a device-side compilation tool for an artificial intelligence acceleration device. Through the embodiments of the present disclosure, different host-side codes are compiled using a host-side compilation tool, and different device-side compilation tools are used to compile different device-side codes, thereby achieving separate compilation of host-side codes and different device-side codes, which can effectively reduce the huge overhead required for developing a unified device-side compilation.

[0045] It can be understood that the above describes the compilation process of the device-side code and the host-side code. The following will further describe the process of compiling the target function codes of multiple artificial intelligence acceleration devices in combination with operation S311 and operation S312.

[0046] In operation S311, target function codes of respective multiple artificial intelligence acceleration devices are generated according to respective kernel information of the multiple device-side codes.

[0047] In some embodiments, the kernel information may include kernel type sub-information. The kernel type sub-information may indicate an artificial intelligence acceleration device corresponding to the device-side code. The target function code generated according to the kernel information may include kernel type sub-information. For example. As described above, multiple device-side codes may be compiled using a first device-side compilation tool, a second device-side compilation tool, and a third device-side compilation tool, respectively. When compiling the first device-side code, the first device-side compilation tool may extract the first kernel information from the first device-side code. The first kernel information may include the first kernel type sub-information. The first kernel type sub-information may indicate that the device executing the first device-side executable file is the first artificial intelligence acceleration device. The first target function code of the first artificial intelligence acceleration device may include the first kernel type sub-information. When compiling the second device-side code, the second device-side compilation tool may extract the second kernel information from the second device-side code. The second kernel information may include the second kernel type sub-information. The second kernel type sub-information may indicate that the device executing the second device-side executable file is the second artificial intelligence acceleration device. The second target function code of the second artificial intelligence acceleration device may include the second kernel type sub-information. When compiling the third device-side code, the third device-side compilation tool may extract the third kernel information from the third device-side code. The third kernel information may include the third kernel type sub-information. The third kernel type sub-information may indicate that the device executing the third device-side executable file is a third artificial intelligence acceleration device. The third target function code of the third artificial intelligence acceleration device may include the third kernel type sub-information.

[0048] In some embodiments, the objective function code may also include a startup interface for an artificial intelligence acceleration device. The startup interface may be a software interface based on the underlying driver setting of the artificial intelligence acceleration device, or it may be a startup interface of a runtime application program interface. For example, a first objective function code may include a first startup interface for a first artificial intelligence acceleration device. A second objective function code may include a second startup interface for a second artificial intelligence acceleration device. A third objective function code may include a third startup interface for a third artificial intelligence acceleration device.

[0049] In operation S312, the host-side compilation tool is used to compile the target function codes of the multiple artificial intelligence acceleration devices to obtain multiple relocatable target files of the target functions to be processed.

[0050] In some embodiments, the target function code is a function code on the host side, and the target function code can be compiled using a host-side compilation tool. For example, by compiling the first target function code using a host-side compilation tool, a first target function relocatable target file to be processed can be obtained. By compiling the second target function code using a host-side compilation tool, a second target function relocatable target file to be processed can be obtained. By compiling the third target function code using a host-side compilation tool, a third target function relocatable target file to be processed can be obtained. Through the embodiments of the present disclosure, different target function codes include startup interfaces of different artificial intelligence acceleration devices, which realizes the encapsulation of the startup interface by the target function code. In addition, the target function code is compiled by the host-side compilation tool, and the host side can implement the call of different startup interfaces based on different target function codes.

[0051] It can be understood that the above describes the target function code and the target function relocatable target file to be processed of the present disclosure. Next, multiple processed target function relocatable target files file_31 can be obtained based on multiple device-side executable files for multiple artificial intelligence acceleration devices and multiple target function relocatable target files to be processed. The multiple processed target function relocatable target files file_31 may include a first processed target function relocatable target file, a second processed target function relocatable target file, and a third processed target function relocatable target file, which will be described in detail below in conjunction with operation S320.

[0052] In operation S320, the data segment to be embedded for the artificial intelligence acceleration device is embedded in the relocatable target file of the target function to be processed corresponding to the artificial intelligence acceleration device, so as to obtain the relocatable target file of the processed target function corresponding to the artificial intelligence acceleration device.

[0053] In some embodiments, the data segment to be embedded is obtained according to the device-side executable file for the artificial intelligence acceleration device. For example, the first device-side executable file can be used as the first data segment to be embedded. The first data segment to be embedded can be embedded in the first target function to be processed relocatable target file to obtain the first processed target function relocatable target file. For another example, the second device-side executable file can be used as the second data segment to be embedded. The second data segment to be embedded can be embedded in the second target function to be processed relocatable target file to obtain the second processed target function relocatable target file. For another example, the third device-side executable file can be used as the third data segment to be embedded. The third data segment to be embedded can be embedded in the third target function to be processed relocatable target file to obtain the third processed target function relocatable target file. Through the embodiment of the present disclosure, the device-side executable file is used as a data segment to be embedded, and the target function to be processed relocatable target file is embedded, so that after the file corresponding to the target function code is executed, the artificial intelligence acceleration device can obtain a complete device-side executable file.

[0054] It can be understood that the above describes the method of obtaining various relocatable target files. In order to effectively start the device-side executable file, the kernel configuration parameter data of the artificial intelligence acceleration device can be obtained from the host-side code, which will be explained in conjunction with operations S303 to S305.

[0055] In operation S303, multiple host-side codes are parsed to obtain kernel configuration parameter data of each of the multiple artificial intelligence acceleration devices.

[0056] In some embodiments, when parsing multiple host-side codes, multiple syntax trees corresponding to multiple kernel configuration parameters can be constructed. There are differences between the kernel configuration parameters of different artificial intelligence acceleration devices. Based on the multiple syntax trees, the host-side compilation tool can identify the kernel configuration parameter data of each of the multiple artificial intelligence acceleration devices.

[0057] In some embodiments, the kernel configuration parameter data may indicate the hardware resources required to execute the device-side executable file. Taking the first artificial intelligence acceleration device including a general-purpose graphics processor as an example, the kernel configuration parameter data may indicate the size of a thread block and the size of a thread, and may also indicate the capacity of a shared memory unit occupied by the first device-side executable file.

[0058] In operation S304, a unified parameter data set is generated.

[0059] In some embodiments, the unified parameter data set may include kernel configuration parameter data of each of the plurality of artificial intelligence acceleration devices. For example, the data type of the unified parameter data set may be a structure (struct). The kernel configuration parameter data may include a plurality of fields. The order of the plurality of fields of the plurality of kernel configuration parameter data may be adjusted to generate the unified parameter data set.

[0060] In operation S305 , a unified parameter data set is provided to a plurality of artificial intelligence acceleration devices.

[0061] In some embodiments, the unified parameter data set can be provided to multiple startup interfaces for multiple artificial intelligence acceleration devices, respectively. For example, the unified parameter data set can be saved to the runtime application program interface of each of the multiple artificial intelligence acceleration devices to provide the unified parameter data set to multiple startup interfaces.

[0062] In other embodiments, the unified parameter data set can be combined as input parameters of each of the multiple target function codes. For example, the unified parameter data set can be combined as an input parameter of the first target function code. The unified parameter data set can be combined as an input parameter of the second target function code. The unified parameter data set can be combined as an input parameter of the third target function code. Through the embodiments of the present disclosure, the kernel configuration parameter data of different artificial intelligence processors can be passed to the corresponding artificial intelligence acceleration devices through the unified parameter data set, which can effectively realize that the device-side executable files of different types of artificial intelligence acceleration devices are started in a unified manner.

[0063] It can be understood that some methods of obtaining kernel configuration parameters are described above, and some methods of obtaining executable files of heterogeneous programs will be described below.

[0064] In operation S330, the host-side relocatable target file and the plurality of processed target function relocatable target files are linked to obtain a heterogeneous program executable file.

[0065] For example, the host-side relocatable target file file_30 and multiple target function relocatable target files file_31 may be linked to obtain a heterogeneous program executable file file_32.

[0066] It can be understood that the above describes some methods of obtaining the executable files of heterogeneous programs, and the following describes the methods of executing the executable files of heterogeneous programs.

[0067] In some embodiments, the heterogeneous program executable file may include a host-side executable file and a plurality of target function executable files corresponding to a plurality of artificial intelligence acceleration devices. The host-side executable file is executed to call the target function executable file. For example, after the host-side executable file is executed, the plurality of target function executable files may be called in an order corresponding to the plurality of host-side codes. The plurality of target function executable files may include a first target function executable file for a first artificial intelligence acceleration device, a second target function executable file for a second artificial intelligence acceleration device, and a third target function executable file for a third artificial intelligence acceleration device.

[0068] In some embodiments, the target function executable file is called to perform the following operations: controlling the artificial intelligence acceleration device to execute the device-side executable file.

[0069] In some embodiments, controlling the artificial intelligence acceleration device to execute the device-side executable file may include: calling the startup interface for the artificial intelligence acceleration device according to the kernel configuration parameter data of the artificial intelligence acceleration device to provide the device-side executable file to the artificial intelligence acceleration device. The kernel configuration parameter data of the artificial intelligence acceleration device may be obtained from the above-mentioned unified parameter data set. For example, taking the first artificial intelligence acceleration device as an example, the kernel configuration parameter data of the first artificial intelligence acceleration device may be obtained from the unified parameter data set. According to the kernel configuration parameter data of the first artificial intelligence acceleration device, the operating environment of the first artificial intelligence acceleration device is configured. Next, the startup interface for the first artificial intelligence acceleration device may be called to provide the first device-side executable file to the first artificial intelligence acceleration device. Via the startup interface for the first artificial intelligence acceleration device, a startup signal may be provided to the first artificial intelligence acceleration device so that the first artificial intelligence acceleration device starts to execute the first device-side executable file. When the first device-side executable file is executed or in the process of being executed. The first artificial intelligence acceleration device may provide the intermediate results and the execution results to the host side.

[0070] It can be understood that the method of the present disclosure is described above, and the device of the present disclosure will be described below.

[0071] Figure 4 is a block diagram of a code processing device according to an embodiment of the present disclosure.

[0072] like Figure 4 As shown, the apparatus 400 may include a first compiling module 410 , an obtaining module 420 , and a linking module 430 .

[0073] The first compiling module 410 is used to compile the target function codes of the respective artificial intelligence acceleration devices to obtain a plurality of target function relocatable target files to be processed. The plurality of artificial intelligence acceleration devices correspond to the host-side relocatable target files.

[0074] The acquisition module 420 is used to obtain multiple processed target function relocatable target files according to multiple device-side executable files for multiple artificial intelligence acceleration devices and multiple target function relocatable target files to be processed.

[0075] The linking module 430 is used to link the host-side relocatable target file and the multiple processed target function relocatable target files to obtain a heterogeneous program executable file.

[0076] In some embodiments, the first compilation module includes: a generation module, which is used to generate target function codes of multiple artificial intelligence acceleration devices according to the kernel information of the multiple device-side codes. The kernel information includes kernel type sub-information, and the kernel type sub-information is used to indicate the artificial intelligence acceleration device corresponding to the device-side code. The target function code includes the kernel type sub-information, and the device-side executable file is compiled from the device-side code. The first compilation sub-module is used to compile the target function codes of the multiple artificial intelligence acceleration devices using the host-side compilation tool to obtain multiple target function relocatable target files to be processed.

[0077] In some embodiments, the apparatus further comprises at least one of the following modules: a second compilation module, for compiling a plurality of host-side codes respectively used for a plurality of artificial intelligence acceleration devices, to obtain a host-side relocatable target file; a third compilation module, for compiling a plurality of device-side codes respectively used for a plurality of artificial intelligence acceleration devices, to obtain a plurality of device-side executable files.

[0078] In some embodiments, the device-side code for the artificial intelligence acceleration device is compiled using a device-side compilation tool for the artificial intelligence acceleration device.

[0079] In some embodiments, the obtaining module is further used to: embed the data segment to be embedded for the artificial intelligence acceleration device into the relocatable target file of the target function to be processed corresponding to the artificial intelligence acceleration device, and obtain the processed target function relocatable target file corresponding to the artificial intelligence acceleration device. The data segment to be embedded is obtained according to the device-side executable file for the artificial intelligence acceleration device.

[0080] In some embodiments, the heterogeneous program executable file includes a host-side executable file and a plurality of target function executable files corresponding to a plurality of artificial intelligence acceleration devices, and the host-side executable file is executed to call the target function executable file. The target function executable file is called to perform the following operations: control the artificial intelligence acceleration device to execute the device-side executable file.

[0081] In some embodiments, controlling the artificial intelligence acceleration device to execute the device-side executable file includes: calling a startup interface for the artificial intelligence acceleration device according to the kernel configuration parameter data of the artificial intelligence acceleration device to provide the device-side executable file to the artificial intelligence acceleration device. Controlling the artificial intelligence acceleration device to start executing the device-side executable file.

[0082] In some embodiments, the kernel configuration parameter data of the artificial intelligence acceleration device is obtained from a unified parameter data set, which includes the kernel configuration parameter data of each of the multiple artificial intelligence acceleration devices, and the kernel configuration parameter data of each of the multiple artificial intelligence acceleration devices is obtained by parsing multiple host-side codes.

[0083] In some embodiments, the apparatus further comprises at least one of the following modules: a first providing module, configured to provide the unified parameter data set to a plurality of startup interfaces respectively used for a plurality of artificial intelligence acceleration devices. A second providing module, configured to combine the unified parameter data set as input parameters of each of the plurality of objective function codes.

[0084] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved are in compliance with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0085] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium and a computer program product.

[0086] Figure 5 A schematic block diagram of an example electronic device 1400 that can be used to implement an embodiment of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.

[0087] like Figure 5As shown, the device 500 includes a computing unit 501, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 502 or a computer program loaded from a storage unit 508 to a random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of the device 500 can also be stored. The computing unit 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.

[0088] A number of components in the device 500 are connected to the I / O interface 505, including: an input unit 506, such as a keyboard, a mouse, etc.; an output unit 507, such as various types of displays, speakers, etc.; a storage unit 508, such as a disk, an optical disk, etc.; and a communication unit 509, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 509 allows the device 500 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0089] The computing unit 501 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSP), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 501 performs the various methods and processes described above, such as code processing methods. For example, in some embodiments, the code processing method may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 508. In some embodiments, part or all of the computer program may be loaded and / or installed on the device 500 via the ROM 502 and / or the communication unit 509. When the computer program is loaded into the RAM 503 and executed by the computing unit 501, one or more steps of the code processing method described above may be performed. Alternatively, in other embodiments, the computing unit 501 may be configured to execute the code processing method in any other appropriate manner (eg, by means of firmware).

[0090] Various embodiments of the systems and techniques described above herein may be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard parts (ASSPs), system on chip systems (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0091] The program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.

[0092] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or device, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include electrical connections based on one or more lines, portable computer disks, hard disks, random access memories, read-only memories, erasable programmable read-only memories (EPROM) or flash memories, optical fibers, portable compact disk read-only memories (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0093] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a cathode ray tube (CRT) display or a liquid crystal display (LCD)) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0094] The systems and techniques described herein may be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a Local Area Network (LAN), a Wide Area Network (WAN), and the Internet.

[0095] A computer system may include clients and servers. Clients and servers are generally remote from each other and usually interact through a communication network. The relationship of client and server is generated by computer programs running on respective computers and having a client-server relationship to each other.

[0096] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this document does not limit this.

[0097] The above specific implementations do not constitute a limitation on the protection scope of the present disclosure. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.

Claims

1. A code processing method, comprising: Compiling the target function codes of the respective artificial intelligence acceleration devices to obtain a plurality of target function relocatable target files to be processed, wherein the plurality of artificial intelligence acceleration devices correspond to the host-side relocatable target files; According to a plurality of device-side executable files for a plurality of the artificial intelligence acceleration devices and a plurality of the target function relocatable target files to be processed, a plurality of processed target function relocatable target files are obtained; The host-side relocatable target file and the plurality of processed target function relocatable target files are linked to obtain a heterogeneous program executable file.

2. The method according to claim 1, wherein: The target function codes of the respective multiple artificial intelligence acceleration devices are compiled to obtain multiple target function relocatable target files to be processed, including: Generate a plurality of target function codes for the artificial intelligence acceleration devices according to the kernel information of the plurality of device-side codes, wherein the kernel information includes kernel type sub-information, the kernel type sub-information is used to indicate the artificial intelligence acceleration device corresponding to the device-side code, the target function code includes the kernel type sub-information, and the device-side executable file is obtained by compiling the device-side code; The host-side compilation tool is used to compile the target function codes of each of the plurality of artificial intelligence acceleration devices to obtain a plurality of relocatable target files of the target functions to be processed.

3. The method according to claim 1, further comprising at least one of the following operations: Compiling a plurality of host-side codes respectively used for a plurality of the artificial intelligence acceleration devices to obtain the host-side relocatable target file; Multiple device-side codes respectively used for the multiple artificial intelligence acceleration devices are compiled respectively to obtain multiple device-side executable files.

4. The method according to claim 3, wherein: The device-side code for the artificial intelligence acceleration device is compiled using a device-side compilation tool for the artificial intelligence acceleration device.

5. The method according to claim 1, wherein: The step of obtaining a plurality of processed target function relocatable target files according to a plurality of device-side executable files for a plurality of the artificial intelligence acceleration devices and a plurality of the target function relocatable target files to be processed comprises: The data segment to be embedded for the artificial intelligence acceleration device is embedded in the relocatable target file of the target function to be processed corresponding to the artificial intelligence acceleration device, so as to obtain the processed target function relocatable target file corresponding to the artificial intelligence acceleration device, wherein the data segment to be embedded is obtained according to the device-side executable file for the artificial intelligence acceleration device.

6. The method according to claim 1, wherein: The heterogeneous program executable file includes a host-side executable file and a plurality of target function executable files corresponding one-to-one to the plurality of artificial intelligence acceleration devices, and the host-side executable file is executed to call the target function executable file. The target function executable file is called to perform the following operations: controlling the artificial intelligence acceleration device to execute the device-side executable file.

7. The method according to claim 6, wherein: The controlling the artificial intelligence acceleration device to execute the device-side executable file includes: Calling a startup interface for the artificial intelligence acceleration device according to the kernel configuration parameter data of the artificial intelligence acceleration device to provide the device-side executable file to the artificial intelligence acceleration device; Control the artificial intelligence acceleration device to start executing the device-side executable file.

8. The method according to claim 7, wherein: The kernel configuration parameter data of the artificial intelligence acceleration device is obtained from a unified parameter data set, which includes the kernel configuration parameter data of each of the multiple artificial intelligence acceleration devices, and the kernel configuration parameter data of each of the multiple artificial intelligence acceleration devices is obtained by parsing the multiple host-side codes.

9. The method according to claim 7, wherein: The method further comprises at least one of the following operations: Providing the unified parameter data set to a plurality of startup interfaces respectively used for a plurality of the artificial intelligence acceleration devices; The unified parameter data is combined as input parameters of each of the plurality of target function codes.

10. A code processing device, comprising: A first compiling module is used to compile the target function codes of the respective artificial intelligence acceleration devices to obtain a plurality of relocatable target files of the target functions to be processed, wherein the plurality of artificial intelligence acceleration devices correspond to the relocatable target files on the host side; An obtaining module, used for obtaining a plurality of processed target function relocatable target files according to a plurality of device-side executable files for a plurality of the artificial intelligence acceleration devices and a plurality of the target function relocatable target files to be processed; The linking module is used to link the host-side relocatable target file and the plurality of processed target function relocatable target files to obtain a heterogeneous program executable file.

11. An electronic device, comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 9.

12. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 9.

13. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 9.