Kernel function precompilation method, device, computer equipment and storage medium
By dividing the parameter list into sublists and precompiling on the slave nodes, executable files are generated, and the time consumption problem caused by MIOpen in real-time compilation in network model calculation is solved, and more efficient network model training is achieved.
Patent Information
- Application Number
- CN202111156980.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-30
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2041-09-30
AI Technical Summary
When performing network model calculations, MIOpen needs to be compiled in real time according to the actual kernel function size, resulting in a large amount of time consumption, increasing the startup preparation time of network model calculations, and affecting the entire training time and efficiency.
By obtaining the parameter list, dividing it into multiple sublists and issuing it to the corresponding slave nodes for precompilation, an executable file is generated, which reduces the compilation time of kernel function, and directly runs the executable file during network model training.
It effectively reduces the startup preparation time of network model calculation, shortens the training time, improves the efficiency of network model calculation, and greatly reduces the time consumption of kernel function compilation through distributed compilation.
Smart Images

Figure CN114035795B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a kernel function pre-compilation method, device, computer equipment and storage medium. Background Art
[0002] With the rapid development of artificial intelligence (AI) technology, deep learning, as an important branch of machine learning, has attracted widespread attention.
[0003] Deep learning is a machine learning technology used to build and simulate neural networks for analytical learning in the human brain, and to imitate the human brain's mechanism to interpret data. It requires high computing power to support it, so heterogeneous computing plays a pivotal role in it. In order to give full play to the computing power of heterogeneous chips and improve computing efficiency, chip manufacturers have launched efficient deep learning computing libraries. Heterogeneous chips can run various network models by calling the interface of the deep learning library, such as the open source high-performance machine intelligence library MIOpen under the ROCm (ROCplatforM, an open software ecosystem for accelerated computing) platform.
[0004] Currently, when performing network model calculations, MIOpen needs to perform real-time compilation based on the actual kernel function size, and then load the compiled kernel function into the running environment. However, real-time compilation of the kernel function consumes a lot of time, increases the startup preparation time for network model calculations, significantly affects the duration of the entire training, and reduces the efficiency of network model calculations. Summary of the invention
[0005] The embodiments of the present application provide a kernel function pre-compilation method, apparatus, computer equipment and storage medium, which can reduce the startup preparation time of network model calculation, shorten the training time, and improve the efficiency of network model calculation.
[0006] In a first aspect, an embodiment of the present application provides a kernel function pre-compilation method, comprising:
[0007] Obtaining a parameter list, wherein the parameter list is used to define parameter information of the kernel function;
[0008] Dividing the parameter list into a plurality of sub-lists according to a preset node list, wherein the preset node list includes a plurality of slave nodes;
[0009] Sending each of the sub-lists to a corresponding slave node respectively, each of the slave nodes including at least one kernel function, to instruct the slave node to precompile the corresponding kernel function according to the sub-list and generate an executable file;
[0010] Obtain the executable file sent by the slave node.
[0011] In the above embodiment, distributed preprocessing is achieved by dividing the parameter list into multiple sub-lists and sending them to the corresponding slave nodes for pre-compilation, which can effectively reduce the time of kernel function compilation. The kernel function is pre-compiled into an executable file, and the executable file can be directly run during network model training. Compared with real-time compilation, the preparation time for starting the network model is reduced.
[0012] In one embodiment, dividing the parameter list into a plurality of sub-lists according to the preset node list includes:
[0013] The parameter list is distributed and processed according to the preset node list to obtain a plurality of sub-lists, each of which includes at least one parameter.
[0014] In the above embodiment, the slave nodes are evenly divided into multiple sub-lists, and distributed pre-compilation is adopted, which can effectively reduce the kernel function compilation time.
[0015] In one embodiment, obtaining a parameter list includes:
[0016] A first parameter list is generated according to a preset configuration file and / or a second parameter list is extracted from a preset deep network.
[0017] In the above embodiment, the kernel function is compiled according to the parameter list to generate an executable file, which can increase the range of executable files that can be directly called, thereby reducing the occurrence of situations where the kernel function needs to be compiled in real time, and further reducing the preparation time for starting the network model.
[0018] In one embodiment, the method further comprises:
[0019] After all the executable files sent by the slave nodes are acquired, the executable files are merged into a preset database.
[0020] In the above embodiment, the precompiled executable files are integrated with the existing database, which can effectively reduce the proportion of real-time compilation required for common networks, thereby effectively reducing the time of the startup phase and achieving a performance improvement of more than 50%.
[0021] In one embodiment, merging the executable file into a preset database includes:
[0022] Performing a deduplication operation on all the executable files;
[0023] The executable files after the deduplication operation are merged into a preset database.
[0024] In the above embodiment, the executable files after deduplication are merged into the preset database, which not only increases the callable range, but also avoids the executable files from repeatedly occupying the database space. In addition, it avoids the existence of duplicate executable files in the database, which affects the efficiency of directly calling the executable files.
[0025] In one embodiment, before obtaining the parameter list, the method further includes:
[0026] Clear the local cache.
[0027] In the above embodiment, the local cache is cleared before the parameter list is obtained, so that the residual parameters of the previous compilation can be prevented from affecting the current compilation.
[0028] In one embodiment, the method further comprises:
[0029] Monitor the working status of each of the slave nodes.
[0030] In the above embodiment, the compilation status and progress of the slave nodes can be timely understood.
[0031] In a second aspect, an embodiment of the present application further provides a kernel function pre-compilation device, comprising:
[0032] A first acquisition module is used to acquire a parameter list, where the parameter list is used to define parameter information of a kernel function;
[0033] A division module, used for dividing the parameter list into a plurality of sub-lists according to a preset node list, wherein the preset node list includes a plurality of slave nodes;
[0034] A sending module, used for sending each of the sub-lists to the corresponding slave nodes respectively, so as to instruct the slave nodes to pre-compile the corresponding kernel function according to the sub-lists and generate an executable file;
[0035] The second acquisition module is used to acquire the executable file sent by the slave node.
[0036] In a third aspect, an embodiment of the present application further provides a computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above method when executing the computer program.
[0037] In a fourth aspect, an embodiment of the present application further provides a storage medium on which a computer program is stored, wherein the computer program implements the steps of the above method when executed by a processor.
[0038] The embodiment of the present application provides a kernel function pre-compilation method, device, computer equipment and storage medium, the method comprising: obtaining a parameter list, the parameter list is used to define the parameter information of the kernel function; dividing the parameter list into multiple sub-lists according to a preset node list, the preset node list includes multiple slave nodes; sending each sub-list to the corresponding slave node respectively to instruct the slave node to pre-compile the corresponding kernel function according to the sub-list and generate an executable file; obtaining the executable file sent by the slave node. The present application generates an executable file by performing distributed pre-compilation of the kernel function, so that the executable file can be directly run during network model training without the need to compile the kernel function in real time during the startup phase, thereby effectively reducing the startup time of the network model training, thereby improving the efficiency of the network model training, and distributed compilation can greatly reduce the time consumption of kernel function compilation. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.
[0040] Figure 1 It is a flowchart of the kernel function precompilation method provided in an embodiment of the present application.
[0041] Figure 2 is another flowchart of the kernel function precompilation method provided in an embodiment of the present application;
[0042] Figure 3 It is a schematic diagram of an application scenario of the kernel function precompilation method provided in an embodiment of the present application;
[0043] Figure 4 It is a structural diagram of a kernel function precompilation device provided in an embodiment of the present application;
[0044] Figure 5 It is a structural diagram of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0045] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of this application.
[0046] The present application provides a kernel function pre-compilation method, device, computer equipment and storage medium. Specifically, the present application provides a kernel function pre-compilation method applicable to a kernel function pre-compilation device, and the kernel function pre-compilation device can be integrated in a computer equipment.
[0047] See also Figure 1 , Figure 1 A flowchart of a kernel function precompilation method provided in an embodiment of the present application is provided. The method is applied to a computer device and may mainly include steps 101 to 103. The description of each step is as follows:
[0048] Step 101: Obtain a parameter list, where the parameter list is used to define parameter information of the kernel function.
[0049] Specifically, the computer device may include a parameter list generating component, and the parameter list generating component is used to generate a parameter list.
[0050] In this embodiment, step 101 may mainly include: generating a first parameter list according to a preset configuration file and / or extracting a second parameter list from a preset deep network.
[0051] The preset configuration file can be a YAML file, which is a language specifically used to write configuration files. It is very concise and powerful. It gives a template for generating parameters, including the starting conditions and step size of each parameter. The parameter list generation component can generate a parameter list based on the template and parameter range for generating parameters. The preset deep network can be an existing commonly used deep network, such as a convolutional neural network, and the parameter list generation component can directly extract the parameter list from the preset deep network.
[0052] For example, users can preset a YAML file containing commonly used parameters of convolutional neural networks, such as input width, input height, input dimension, output dimension, convolution kernel width, convolution kernel height, padding width, padding height, batch size, step width, step height, expansion width, expansion height, bias, number of groups, etc. The parameter list generation component can generate a parameter list using a given parameter template.
[0053] In some embodiments, before obtaining the parameter list, the method further includes: clearing the local cache.
[0054] Specifically, the computer device may further include a main control component for dividing the parameter list, and for issuing the parameter list in subsequent steps, monitoring the slave nodes, etc. Therefore, after generating the parameter list, the parameter list generation component sends the parameter list to the main control component for subsequent processing.
[0055] Specifically, in order to prevent the last cache data from affecting the current compilation, the main control component clears the local cache before obtaining the parameter list generated by the parameter list generation component.
[0056] Step 102: Divide the parameter list into a plurality of sub-lists according to a preset node list, wherein the preset node list includes a plurality of slave nodes.
[0057] Among them, the slave node can be understood as a slave device connected to the computer device. For example, the computer device can be a master server, and the slave node can be multiple slave servers. The slave server and the master server can be connected via wired or wireless means.
[0058] In this embodiment, step 102 may mainly include: performing allocation processing on the parameter list according to the preset node list to obtain a plurality of sub-lists, each of which includes at least one parameter.
[0059] For example, in some embodiments, step 102 may specifically include: obtaining the number of slave nodes; and equally dividing the parameter list according to the number of slave nodes to obtain multiple sub-lists.
[0060] It is easy to understand that when the performance of each slave server is not much different, the parameter list can be directly divided according to the number of slave nodes to ensure that the time for each slave node to compile the sub-list is roughly the same, which neither wastes slave node resources nor saves compilation time.
[0061] For example, assuming that the parameter template contains three parameters, namely length, width, and height, where the length can take values 1 and 3, the width can take values 2 and 4, and the height can take values 3 and 5. According to the parameter template, a parameter list as shown in Table 1 can be obtained. The parameter list includes eight entries. Assuming that the number of slave nodes is four, the parameter list is divided into four sublists, each of which contains two entries. It is easy to understand that dividing the parameter list into N sublists (N is a positive integer) and assigning them to N slave nodes for pre-compilation can shorten the compilation time by N times compared to using only one node for compilation.
[0062] Serial number long Width high 1 1 2 3 2 1 2 5 3 1 4 5 4 3 4 5 5 1 4 3 6 3 4 3 7 3 2 3 8 3 2 5
[0063] Table 1
[0064] For another example, in some embodiments, step 102 may mainly include: performing allocation processing on the parameter list according to the computing performance of each slave node in the preset node list to obtain a plurality of sub-lists.
[0065] It is easy to understand that if the performance of each slave server varies greatly, for example, the computing power varies greatly, and the parameter list is divided equally according to the number of slave nodes, it may result in that after a slave node is compiled, the compilation of other slave nodes has not been completed. In this case, the slave node can only wait for the compilation of other slave nodes, which will lead to a waste of slave node resources and reduce the compilation efficiency. Therefore, the parameter list can be allocated according to the computing power of each slave node in the preset node list.
[0066] In this embodiment, each slave node can also be configured with a parameter entry upper limit. When the parameters in the parameter list are allocated and processed, the upper limit of the parameter entry of the slave node must not be exceeded. In this way, a large number of compilation tasks can be avoided from being concentrated on a certain slave node, ensuring load balancing among multiple slave nodes, so as to improve the response speed and availability of the entire system.
[0067] Step 103: Send each sub-list to the corresponding slave node respectively, so as to instruct the slave node to pre-compile the corresponding kernel function according to the sub-list and generate an executable file.
[0068] Specifically, after the master control component divides the parameter list into multiple sub-lists, it sends the sub-lists to the slave nodes and sends a start instruction to instruct the slave nodes to precompile the corresponding kernel functions according to the sub-lists to generate executable files.
[0069] In some embodiments, if in the above step 102, the parameter list is directly divided equally according to the number of slave nodes, then step 103 may mainly include: sending each sub-list to the corresponding slave node in sequence; or, sending each sub-list to a slave node randomly.
[0070] As in the above example, assuming that there are slave nodes h1, h2, h3 and h4, the first and second items can be sent to h1 for compilation, and then the third and fourth items can be sent to h2 for compilation, and so on.
[0071] In some embodiments, if in the above step 102, the parameter list is distributed according to the computing performance of the slave node to obtain multiple sub-lists, then step 103 may mainly include: synchronously sending each sub-list to the corresponding slave node.
[0072] As in the above example, assuming that there are slave nodes h1, h2, h3 and h4, if the computing performance of slave node h1 is significantly higher than that of slave nodes h2, h3 and h4, the parameter list can be allocated into four sublists, wherein the first sublist includes entries one to five, and the other sublists each include one entry. Afterwards, the first sublist can be sent to slave node h1, and the other three sublists can be randomly sent to slave nodes h2, h3 and h4 to ensure that slave nodes with strong computing performance compile more parameter entries, and slave nodes with weak computing performance compile fewer parameter entries, so that the compilation time of parameters for each slave node is relatively uniform, which will not lead to waste of slave node resources and save compilation time.
[0073] In order to improve the distribution efficiency, the master control node can call multiple threads and use multiple threads to concurrently execute the distribution of multiple sub-lists.
[0074] Specifically, after receiving the compilation task from the node, the storage address of the executable file can be set first, and then the corresponding kernel function can be compiled according to the sub-list. After the executable file is generated, the executable file can be run to verify the availability and accuracy of the executable file.
[0075] Furthermore, when running executable files to verify their availability and accuracy, the running time of each executable file may be recorded, and when the executable files are subsequently called for running, the executable file with the shortest running time may be called for running.
[0076] In this embodiment, the method may further include: monitoring the working status of each slave node.
[0077] Specifically, after the master control component sends the sublist to the slave node and starts the slave node, it can monitor the slave node's working status regularly, so as to timely understand the slave node's compilation status and compilation progress. Specifically, the slave node will regularly feedback the master control component, feedback the current status information and progress, and when all items in the sublist are compiled, all generated executable files will be sent to the master control component.
[0078] For example, if the sublist contains a large number of entries, when each entry is compiled, the slave node will return status information to the master component, so that the master component can ensure that each entry in the sublist is compiled.
[0079] It is easy to understand that the method may also include: dynamically allocating parameter entries of each slave node.
[0080] As in the above example, assuming that there are slave nodes h1, h2, h3 and h4, if each slave node includes four parameter entries, when the master node determines that the four parameter entries assigned to slave node h1 have been compiled based on the status information returned by the slave node, and slave node h2 has only compiled one of the four parameter entries, the last parameter entry assigned to slave node h2 can be reallocated to slave node h1 for compilation. In this way, according to the real-time compilation status of each slave node, the parameter entries are dynamically allocated, which can save compilation time, improve compilation efficiency and the rationality of parameter entry allocation.
[0081] Step 104: Obtain the executable file sent from the node.
[0082] Specifically, the master control node can monitor the status information of the slave nodes regularly, so as to ensure that all entries of all slave nodes are compiled and that the executable files returned by each slave node are received normally.
[0083] In this embodiment, the method may further include: after acquiring all executable files sent from the nodes, merging the executable files into a preset database.
[0084] It is easy to understand that merging executable files into the preset database can increase the range of executable files that can be directly called, further reduce the occurrence of situations where real-time compilation is required, and thus effectively improve the startup time of deep network calculations.
[0085] Furthermore, the step of "merging the executable files into the preset database" may specifically include: performing a deduplication operation on all executable files; and merging the executable files after the deduplication operation into the preset database.
[0086] Specifically, the parameter list may contain multiple entries, and the executable files compiled by different entries may be exactly the same. When the master control component receives the executable files returned by all slave nodes, it can perform a deduplication operation on all executable files, and then merge the deduplication executable files into the preset database, thereby increasing the callable range and avoiding the executable files from repeatedly occupying the database space. In addition, it avoids the existence of duplicate executable files in the database, which affects the efficiency of directly calling the executable files.
[0087] The kernel function precompilation method provided in the embodiment of the present application obtains a parameter list, where the parameter list is used to define the parameter information of the kernel function, and then divides the parameter list into multiple sub-lists according to a preset node list, wherein the preset node list includes multiple slave nodes, and then sends each sub-list to the corresponding slave node respectively to instruct the slave node to pre-compile the corresponding kernel function according to the sub-list to generate an executable file, and then obtains the executable file sent by the slave node, so that the executable file can be directly run during network model training without the need to compile the kernel function in real time during the startup phase, thereby effectively reducing the startup time of network model training, thereby improving the efficiency of network model training, and distributed compilation can greatly reduce the time consumption of kernel function compilation.
[0088] See also Figure 2 , Figure 2 : is another flow chart of the kernel function precompilation method provided in an embodiment of the present application. The specific flow of the kernel function precompilation method provided in this embodiment can be as follows:
[0089] Step 201: Generate a parameter list according to a preset configuration file and a preset deep network, where the parameter list is used to define parameter information of a kernel function.
[0090] The preset configuration file can be a YAML file, which is a language specifically used to write configuration files. It is very concise and powerful. It gives a template for generating parameters, including the starting conditions and step size of each parameter. The parameter list generation component can generate a parameter list based on the template and parameter range for generating parameters. The preset deep network can be an existing commonly used deep network, such as a convolutional neural network, and the parameter list generation component can directly extract the parameter list from the preset deep network.
[0091] For example, see Figure 3 , Figure 3 A schematic diagram of an application scenario of the kernel function precompilation method provided in an embodiment of the present application, wherein the kernel function precompilation system includes a master control server and a slave server, wherein the master control server includes a parameter list generation component, and the parameter list generation component can generate a parameter list according to a preset configuration file (input.yaml) and a preset deep network.
[0092] Specifically, the computer device can also be a main control component for dividing the parameter list, and for issuing the parameter list in subsequent steps, monitoring the slave nodes, etc. Therefore, after generating the parameter list, the parameter list generation component sends the parameter list to the main control component for subsequent processing.
[0093] Specifically, in order to prevent the last cache data from affecting the current compilation, the main control component clears the local cache before obtaining the parameter list generated by the parameter list generation component.
[0094] Step 202: Divide the parameter list into multiple sub-lists according to the number of multiple slave nodes in the preset node list.
[0095] Please continue reading Figure 3 After obtaining the parameter list, the main control component divides the parameter list. As described in the above embodiment, assuming that the parameter list includes eight entries and the number of slave nodes is four, the parameter list is divided into four sub-lists, each sub-list contains two entries.
[0096] It is easy to understand that dividing the parameter list into N sublists (N is a positive integer) and allocating them to N slave nodes for pre-compilation can shorten the compilation time by N times compared to using only one node for compilation.
[0097] Step 203: Send each sub-list to the corresponding slave node respectively, so as to instruct the slave node to pre-compile the corresponding kernel function according to the sub-list to generate an executable file.
[0098] For details, please continue to see Figure 3 After the master control component divides the parameter list into multiple sub-lists, it sends the sub-lists to the slave node and sends a startup instruction. After the slave node receives the sub-list and the startup instruction, it can first set the storage address of the executable file, and then compile the corresponding kernel function according to the sub-list. After the executable file is generated, it can be run to verify the availability and accuracy of the executable file.
[0099] It is easy to understand that each slave node can include multiple kernel functions that implement the same function. At present, for a specific network model, the optimal kernel function is usually directly loaded into the execution program, and the real-time compilation of the kernel function and the real-time selection of the optimal kernel function are not performed. For example, most reasoning programs will directly burn the corresponding network parameters into a dedicated chip, and the kernel functions used are also fixed. In this embodiment, after the slave node obtains the sublist, the optimal kernel function can be selected for compilation according to the parameter type and hardware configuration, wherein the principle of selecting the optimal kernel function is to select the kernel function with the shortest compilation time for compilation.
[0100] In this embodiment, the method may further include: monitoring the working status of each slave node.
[0101] Specifically, Figure 3 As shown, after the master control component sends the sublist to the slave node and starts the slave node to work, the working status of the slave node can be monitored regularly, so as to timely understand the compilation status and compilation progress of the slave node.
[0102] Specifically, the slave node will periodically provide feedback to the master control component, feeding back the current status information and progress. When all items in the sublist are compiled, the generated executable file will be sent to the master control component.
[0103] For example, if the sublist contains a large number of entries, when each entry is compiled, the slave node will return status information to the master component, so that the master component can ensure that each entry in the sublist is compiled.
[0104] Step 204: Obtain the executable file sent from the node.
[0105] Step 205: After all executable files sent from the nodes are acquired, deduplication operation is performed on all executable files, and the executable files after deduplication operation are merged into a preset database.
[0106] It is easy to understand that merging executable files into the preset database can increase the range of executable files that can be directly called, further reduce the occurrence of situations where real-time compilation is required, and thus effectively improve the startup time of deep network calculations.
[0107] Specifically, after the main control component receives all executable files returned by the slave nodes, it can perform deduplication operations on all executable files, and then merge the deduplicated executable files into the preset database, thereby increasing the callable range and avoiding executable files from repeatedly occupying database space. In addition, it avoids the existence of duplicate executable files in the database, which affects the efficiency of directly calling executable files.
[0108] All of the above technical solutions can be arbitrarily combined to form optional embodiments of the present application, which will not be described in detail here.
[0109] The kernel function precompilation method provided in the embodiment of the present application generates a parameter list according to a preset configuration file and a preset deep network, the parameter list is used to define the parameter information of the kernel function, and then the parameter list is divided into multiple sub-lists according to the number of multiple slave nodes in the preset node list, and then each sub-list is sent to the corresponding slave node respectively to instruct the slave node to pre-compile the corresponding kernel function according to the sub-list to generate an executable file, and obtain the executable file sent by the slave node. After obtaining the executable files sent by all slave nodes, all executable files are deduplicated, and the executable files after the deduplication operation are merged into the preset database. The calling range of the executable files in the preset database can be increased, thereby reducing the need for real-time kernel function compilation, thereby effectively reducing the startup time of network model training, and distributed compilation can greatly reduce the time consumption of kernel function compilation.
[0110] In order to better implement the kernel function pre-compilation method of the embodiment of the present application, the embodiment of the present application also provides a kernel function pre-compilation device. Figure 4 , Figure 4 The structure diagram of the kernel function pre-compilation device provided in the embodiment of the present application is as follows: The kernel function pre-compilation device 10 may include a first acquisition module 11 , a division module 12 , a delivery module 13 and a second acquisition module 14 .
[0111] The first acquisition module 11 is used to acquire a parameter list, and the parameter list is used to define parameter information of the kernel function.
[0112] The dividing module 12 is used to divide the parameter list into multiple sub-lists according to the preset node list, and the preset node list includes multiple slave nodes.
[0113] The sending module 13 is used to send each sub-list to the corresponding slave node respectively, so as to instruct the slave node to pre-compile the corresponding kernel function according to the sub-list to generate an executable file.
[0114] The second acquisition module 14 is used to acquire the executable file sent from the node.
[0115] In some embodiments, the partitioning module 12 can be mainly used to: distribute the parameter list according to the preset node list to obtain a plurality of sub-lists, each of which includes at least one parameter.
[0116] In some embodiments, the first acquisition module 11 may be mainly used to: generate a first parameter list according to a preset configuration file and / or extract a second parameter list from a preset deep network.
[0117] In some embodiments, the kernel function pre-compilation device 10 may further include a merging module, which is used to: after acquiring all executable files sent from the nodes, merge the executable files into a preset database.
[0118] In some embodiments, the merging module may be specifically used to: perform deduplication operations on all executable files; and merge the executable files after the deduplication operations into a preset database.
[0119] In some embodiments, the kernel function pre-compilation device 10 may further include a cleaning module for cleaning up the local cache.
[0120] In some embodiments, the kernel function pre-compilation device 10 may further include a monitoring module for monitoring the working status of each slave node.
[0121] The kernel function precompilation device 10 provided in the embodiment of the present application obtains a parameter list through a first acquisition module 11, and the parameter list is used to define the parameter information of the kernel function. Then, the division module 12 divides the parameter list into multiple sub-lists according to a preset node list, and the preset node list includes multiple slave nodes. The sending module 13 sends each sub-list to the corresponding slave node respectively to instruct the slave node to precompile the corresponding kernel function according to the sub-list to generate an executable file. Then, the second acquisition module 14 obtains the executable file sent by the slave node, so that the executable file can be directly run during network model training without the need to compile the kernel function in real time during the startup phase, thereby effectively reducing the startup time of network model training, thereby improving the efficiency of network model training, and distributed compilation can greatly reduce the time consumption of kernel function compilation.
[0122] In addition, the embodiment of the present application also provides a computer device, which may be a terminal, and the terminal may be a notebook computer, a personal computer (PC, Personal Computer), a personal digital assistant (Personal Digital Assistant, PDA) and other terminal devices. Figure 5 As shown, Figure 5 A schematic diagram of the structure of a computer device provided in an embodiment of the present application. The computer device 2000 includes a processor 2001 having one or more processing cores, a memory 2002 having one or more computer-readable storage media, and a computer program stored in the memory 2002 and executable on the processor. The processor 2001 is electrically connected to the memory 2002. It will be understood by those skilled in the art that the computer device structure shown in the figure does not constitute a limitation on the computer device, and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.
[0123] Processor 2001 is the control center of computer device 2000. It uses various interfaces and lines to connect various parts of the entire computer device 2000, executes various functions of computer device 2000 and processes data by running or loading software programs and / or modules stored in memory 2002, and calling data stored in memory 2002, thereby monitoring computer device 2000 as a whole.
[0124] In the embodiment of the present application, the processor 2001 in the computer device 2000 will load instructions corresponding to the processes of one or more application programs into the memory 2002 according to the following steps, and the processor 2001 will run the application programs stored in the memory 2002, thereby realizing various functions:
[0125] Get the parameter list, which is used to define the parameter information of the kernel function;
[0126] Dividing the parameter list into a plurality of sub-lists according to a preset node list, wherein the preset node list includes a plurality of slave nodes;
[0127] Send each sublist to the corresponding slave node respectively to instruct the slave node to precompile the corresponding kernel function according to the sublist and generate an executable file;
[0128] Get the executable file sent from the node.
[0129] The specific implementation of the above operations can be found in the previous embodiments, which will not be described in detail here.
[0130] Optional, such as Figure 5 As shown, the computer device 2000 further includes: a touch screen 2003, a radio frequency circuit 2004, an audio circuit 2005, an input unit 2006, and a power supply 2007. The processor 2001 is electrically connected to the touch screen 2003, the radio frequency circuit 2004, the audio circuit 2005, the input unit 2006, and the power supply 2007. Those skilled in the art can understand that Figure 5 The illustrated computer device structure does not constitute a limitation on the computer device, and may include more or fewer components than illustrated, or combine certain components, or arrange the components differently.
[0131] The touch display screen 2003 can be used to display a graphical user interface and receive operation instructions generated by the user acting on the graphical user interface. The touch display screen 2003 may include a display panel and a touch panel. Among them, the display panel can be used to display information input by the user or information provided to the user and various graphical user interfaces of the computer device, which can be composed of graphics, text, icons, videos and any combination thereof. Optionally, the display panel can be configured in the form of a liquid crystal display (LCD, Liquid Crystal Display), an organic light emitting diode (OLED, Organic Light-Emitting Diode), etc. The touch panel can be used to collect the user's touch operations on or near it (such as the user's operation on the touch panel or near the touch panel using any suitable object or accessory such as a finger, stylus, etc.), and generate corresponding operation instructions, and the operation instructions execute the corresponding program. Optionally, the touch panel may include two parts: a touch detection device and a touch controller. Among them, the touch detection device detects the user's touch direction, detects the signal brought by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device, converts it into the touch point coordinates, and then sends it to the processor 2001, and can receive the command sent by the processor 2001 and execute it. The touch panel can cover the display panel. When the touch panel detects a touch operation on or near it, it is transmitted to the processor 2001 to determine the type of touch event, and then the processor 2001 provides corresponding visual output on the display panel according to the type of touch event. In an embodiment of the present application, the touch panel and the display panel can be integrated into the touch display screen 2003 to realize the input and output functions. However, in some embodiments, the touch panel and the touch panel can be used as two independent components to realize the input and output functions. That is, the touch display screen 2003 can also be used as a part of the input unit 2006 to realize the input function.
[0132] The radio frequency circuit 2004 may be used to send and receive radio frequency signals, so as to establish wireless communication with a network device or other computer device through wireless communication, and to send and receive signals between the network device or other computer device.
[0133] The audio circuit 2005 can be used to provide an audio interface between the user and the computer device through a speaker and a microphone. The audio circuit 2005 can transmit the electrical signal converted from the received audio data to the speaker, which is converted into a sound signal for output; on the other hand, the microphone converts the collected sound signal into an electrical signal, which is received by the audio circuit 2005 and converted into audio data, and then the audio data is output to the processor 2001 for processing, and then sent to another computer device through the radio frequency circuit 2004, or the audio data is output to the memory 2002 for further processing. The audio circuit 2005 may also include an earplug jack to provide communication between an external headset and the computer device.
[0134] The input unit 2006 may be used to receive input numbers, character information or user feature information (such as fingerprint, iris, facial information, etc.), and generate keyboard, mouse, joystick, optical or trackball signal input related to user settings and function control.
[0135] The power supply 2007 is used to supply power to various components of the computer device 2000. Optionally, the power supply 2007 can be logically connected to the processor 2001 through a power management system, so that the power management system can manage charging, discharging, and power consumption. The power supply 2007 can also include one or more DC or AC power supplies, recharging systems, power failure detection circuits, power converters or inverters, power status indicators, and other arbitrary components.
[0136] although Figure 5 Not shown, the computer device 2000 may also include a camera, a sensor, a wireless fidelity module, a Bluetooth module, etc., which will not be described in detail here.
[0137] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0138] From the above, it can be seen that the computer device provided in this embodiment obtains a parameter list, which is used to define the parameter information of the kernel function, and then divides the parameter list into multiple sub-lists according to a preset node list, and the preset node list includes multiple slave nodes. Each sub-list is then sent to the corresponding slave node to instruct the slave node to pre-compile the corresponding kernel function according to the sub-list to generate an executable file, and then obtain the executable file sent by the slave node, so that the executable file can be directly run during network model training without the need to compile the kernel function in real time during the startup phase, thereby effectively reducing the startup time of network model training, thereby improving the efficiency of network model training, and distributed compilation can greatly reduce the time consumption of kernel function compilation.
[0139] A person of ordinary skill in the art will appreciate that all or part of the steps in the various methods of the above embodiments may be completed by instructions, or by controlling related hardware through instructions. The instructions may be stored in a computer-readable storage medium and loaded and executed by a processor.
[0140] To this end, an embodiment of the present application provides a computer-readable storage medium, in which multiple computer programs are stored, and the computer program can be loaded by a processor to execute the steps in any one of the kernel function precompilation methods provided in the embodiment of the present application. For example, the computer program can execute the following steps: obtain a parameter list, the parameter list is used to define the parameter information of the kernel function; divide the parameter list into multiple sublists according to a preset node list, and the preset node list includes multiple slave nodes; send each sublist to the corresponding slave node respectively to instruct the slave node to generate an executable file after precompiling the corresponding kernel function according to the sublist; obtain the executable file sent by the slave node.
[0141] The specific implementation of the above operations can be found in the previous embodiments, which will not be described in detail here.
[0142] The storage medium may include: a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0143] Since the computer program stored in the storage medium can execute the steps in any kernel function precompilation method provided in the embodiments of the present application, the beneficial effects that can be achieved by any kernel function precompilation method provided in the embodiments of the present application can be achieved. Please see the previous embodiments for details and will not be repeated here.
[0144] The above is a detailed introduction to a kernel function precompilation method, device, storage medium and computer equipment provided in an embodiment of the present application. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea; at the same time, for technical personnel in this field, according to the idea of the present application, there will be changes in the specific implementation method and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.
Claims
1. A kernel function pre-compilation method, characterized in that: include: Obtaining a parameter list, wherein the parameter list is used to define parameter information of the kernel function; Dividing the parameter list into a plurality of sub-lists according to a preset node list, wherein the preset node list includes a plurality of slave nodes; Sending each of the sub-lists to the corresponding slave nodes respectively to instruct the slave nodes to pre-compile the corresponding kernel functions according to the sub-lists and generate executable files; Obtaining the executable file sent by the slave node; The step of dividing the parameter list into a plurality of sub-lists according to the preset node list comprises: Distributing the parameter list according to the preset node list to obtain a plurality of sub-lists, each of which includes at least one parameter; The acquisition parameter list includes: Generate a first parameter list according to a preset configuration file and / or extract a second parameter list from a preset deep network, wherein the first parameter list is generated according to a template and a parameter range for generating parameters; Directly divide the parameter list equally according to the number of slave nodes to obtain multiple sublists, and send each sublist to the corresponding slave node in sequence, or randomly send each sublist to a slave node; According to the computing performance of the slave nodes, the parameter list is distributed to obtain multiple sub-lists, and each sub-list is synchronously sent to the corresponding slave node.
2. The kernel function precompilation method according to claim 1, characterized in that: The method further comprises: After all the executable files sent by the slave nodes are acquired, the executable files are merged into a preset database.
3. The kernel function pre-compilation method according to claim 2, characterized in that: The step of merging the executable file into a preset database comprises: Performing a deduplication operation on all the executable files; The executable files after the deduplication operation are merged into a preset database.
4. The kernel function precompilation method according to claim 1, characterized in that: Before obtaining the parameter list, the method further includes: Clear the local cache.
5. The kernel function precompilation method according to claim 1, characterized in that: The method further comprises: Monitor the working status of each of the slave nodes.
6. A kernel function precompilation device, characterized in that: include: A first acquisition module is used to acquire a parameter list, where the parameter list is used to define parameter information of a kernel function; A division module, used for dividing the parameter list into a plurality of sub-lists according to a preset node list, wherein the preset node list includes a plurality of slave nodes; A sending module, used for sending each of the sub-lists to the corresponding slave nodes respectively, so as to instruct the slave nodes to pre-compile the corresponding kernel function according to the sub-lists and generate an executable file; A second acquisition module, used for acquiring the executable file sent by the slave node; The step of dividing the parameter list into a plurality of sub-lists according to the preset node list comprises: Distributing the parameter list according to the preset node list to obtain a plurality of sub-lists, each of which includes at least one parameter; The acquisition parameter list includes: Generate a first parameter list according to a preset configuration file and / or extract a second parameter list from a preset deep network, wherein the first parameter list is generated according to a template and a parameter range for generating parameters; Directly divide the parameter list equally according to the number of slave nodes to obtain multiple sublists, and send each sublist to the corresponding slave node in sequence, or randomly send each sublist to a slave node; According to the computing performance of the slave nodes, the parameter list is distributed to obtain multiple sub-lists, and each sub-list is synchronously sent to the corresponding slave node.
7. A computer device, characterized in that: The method comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the method according to any one of claims 1 to 5 when executing the computer program.
8. A storage medium, characterized in that: A computer program is stored, wherein when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Kernel function construction method and device, data prediction method and device, equipment, and storage medium
CN108776717A
Code compiling method and device, storage medium and electronic equipment
CN109710263A
Pseudo instruction compiling method and device, computer equipment and storage medium
CN113050952A