Neural network model calculation task processing method, compiling method and related device

By obtaining the bandwidth aware information of electronic devices and selecting matching candidate execution files, the problem that neural network accelerators cannot achieve optimal performance under different bandwidth resource allocation is solved, and efficient task execution is achieved.

CN120029781APending Publication Date: 2025-05-23SMARTER SILICON (SHANGHAI) TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510164589.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

During the operation phase of the neural network task, the available bandwidth resources allocated by the electronic device to the neural network accelerator are inconsistent with the actual allocated bandwidth resources in the compilation phase, resulting in the neural network accelerator being unable to achieve the expected optimal execution performance.

Method used

By obtaining the bandwidth-aware information of the electronic device, select the candidate execution file that matches the current bandwidth-aware information as the target execution file and load it into the neural network accelerator to ensure the best match of bandwidth resources.

Benefits of technology

It is realized that under the allocation of resources of different bandwidths, neural network accelerators can achieve optimal performance, avoiding the impact of performance degradation and task execution efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120029781A_ABST
    Figure CN120029781A_ABST
Patent Text Reader

Abstract

The invention discloses a neural network model calculation task processing method, a compiling method and a related device, and relates to the technical field of neural networks, bandwidth perception information of electronic equipment is acquired, and the bandwidth perception information represents the bandwidth resource allocation condition of at least one hardware device of the electronic equipment; the at least one hardware device comprises a neural network accelerator used for realizing a calculation task of a neural network model so as to select a candidate execution file matched with the bandwidth perception information from a plurality of candidate execution files corresponding to a to-be-executed target calculation task as a current to-be-executed target execution file according to the bandwidth perception information, and loading the target execution file into the neural network accelerator to enable the neural network accelerator to execute the target execution file so as to complete the target calculation task. Wherein the plurality of candidate execution files respectively correspond to different bandwidth ranges and are used for completing the same calculation task.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of neural network technology, and in particular to a neural network model computing task processing method, a compilation method and related devices. Background Art

[0002] In recent years, neural networks have been widely used in many fields. In order to improve the reasoning speed and resource utilization of neural networks, a large number of neural network accelerators have been proposed to perform neural network tasks such as real-time translation, speech recognition or image segmentation to meet practical application needs.

[0003] Among them, in order to improve the operating efficiency of neural network tasks, during the compilation stage of the original neural network model, the compiler usually optimizes the running speed of the original neural network model based on given bandwidth resources through operator fusion or pipeline arrangement, so that the neural network accelerator can load and execute the optimized neural network model's executable file to obtain the best execution performance.

[0004] However, during the running phase of the neural network task, the available bandwidth resources allocated by the electronic device to the neural network accelerator are often inconsistent with the bandwidth resources actually allocated during the compilation phase, which can easily lead to the neural network accelerator failing to achieve the expected optimal execution performance, or even causing the performance of the neural network accelerator to decline, thereby affecting the task execution efficiency. Summary of the invention

[0005] In view of the above problems, this application provides the following solutions:

[0006] The first aspect of the present application provides a method for processing a neural network model calculation task, the method comprising:

[0007] Acquiring bandwidth perception information of the electronic device; the bandwidth perception information represents bandwidth resource allocation of at least one hardware device of the electronic device, wherein the at least one hardware device includes a neural network accelerator for implementing a computing task of the neural network model;

[0008] According to the bandwidth perception information, from a plurality of candidate execution files corresponding to the target computing task to be executed, a candidate execution file matching the bandwidth perception information is selected as the target execution file to be executed currently; the plurality of candidate execution files respectively correspond to different bandwidth ranges and are used to complete the same computing task;

[0009] The target execution file is loaded into the neural network accelerator so that the neural network accelerator executes the target execution file to complete the target computing task.

[0010] In a possible implementation, the acquiring bandwidth perception information of the electronic device includes:

[0011] Sending a bandwidth test instruction to a neural network accelerator to be executed with respect to the target computing task, so that the neural network accelerator executes the bandwidth test instruction at least once to obtain bandwidth test data; the bandwidth test instruction includes at least part of the computing instructions required by the target computing task;

[0012] Receiving the bandwidth test data fed back by the neural network accelerator;

[0013] Based on the bandwidth test data, bandwidth perception information for the neural network accelerator is obtained.

[0014] In a possible implementation, sending a bandwidth test instruction to a neural network accelerator to execute the target computing task includes any one of the following:

[0015] If the computing tasks of the neural network model include different computing tasks corresponding to each neural network layer, before the target computing task corresponding to the neural network layer to be executed is executed, a bandwidth test instruction is sent to the neural network accelerator to execute the target computing task;

[0016] According to the preset bandwidth test frequency, a bandwidth test instruction is sent to the neural network accelerator to execute the target computing task.

[0017] In a possible implementation, the acquiring bandwidth awareness information of the electronic device further includes:

[0018] Determine a change in available bandwidth resources allocated to the neural network accelerator based on a plurality of bandwidth test data continuously fed back by the neural network accelerator;

[0019] According to the change in the available bandwidth resources, the bandwidth test frequency of the neural network accelerator is adjusted, so that according to the adjusted bandwidth test frequency, the step of sending a bandwidth test instruction to the neural network accelerator to execute the target computing task is executed.

[0020] In a possible implementation, the multiple candidate execution files are compiled based on the original execution file using compilation parameters that match different bandwidth ranges.

[0021] In a possible implementation, selecting, based on the bandwidth awareness information, from multiple candidate execution files corresponding to the target computing task to be executed, a candidate execution file matching the bandwidth awareness information as the target execution file to be executed currently includes any of the following:

[0022] According to the number of hardware devices participating in bandwidth allocation contained in the bandwidth perception information, or the available bandwidth resources allocated to the neural network accelerator, it is determined that the available bandwidth resources of the neural network accelerator are greater than the bandwidth threshold, and a candidate execution file corresponding to the maximum bandwidth range is selected from multiple candidate execution files corresponding to the target computing task to be executed as the target execution file to be executed currently;

[0023] According to the number of hardware devices participating in bandwidth allocation contained in the bandwidth perception information, or the available bandwidth resources allocated to the neural network accelerator, it is determined that the available bandwidth resources of the neural network accelerator are less than the bandwidth threshold, and a candidate execution file matching the available bandwidth resources of the neural network accelerator is selected from multiple candidate execution files corresponding to the target computing task to be executed as the target execution file to be executed currently;

[0024] The bandwidth threshold is determined based on available bandwidth resources of the electronic device.

[0025] A second aspect of the present application provides a neural network model compilation method, the neural network model compilation method comprising:

[0026] Acquire multiple bandwidth ranges that can be obtained by a neural network accelerator in a working state for executing a computing task of the neural network model;

[0027] According to the multiple bandwidth ranges, respectively configuring compilation parameters matching the multiple bandwidth ranges for each of the computing tasks;

[0028] According to the different compilation parameters, the original execution files corresponding to the computing tasks are compiled respectively to obtain a plurality of candidate execution files corresponding to the computing tasks.

[0029] In a possible implementation, the neural network model compilation method further includes:

[0030] Obtaining weight files required to implement each of the computing tasks in the neural network model;

[0031] Multiple candidate execution files corresponding to the same computing task are associated with the weight file and stored, so that when the neural network accelerator executes any of the candidate execution files of the computing task, the weight parameters required for the computing task are retrieved from the associated stored weight file and loaded into the neural network accelerator to complete the computing task.

[0032] In a possible implementation, the obtaining of multiple bandwidth ranges that can be obtained by the neural network accelerator in a working state for executing the computing task of the neural network model includes any one of the following:

[0033] Obtaining a bandwidth variation range during the historical operation of the neural network accelerator, segmenting the bandwidth variation range according to the variation range and variation trend of the historical bandwidth resources, and determining a plurality of bandwidth ranges that the neural network accelerator can obtain to implement the computing task of the neural network model;

[0034] According to the bandwidth configuration information of the electronic device, the original bandwidth range allocated to the neural network accelerator used to execute the computing task of the neural network model is determined, and the original bandwidth range is segmented to obtain multiple bandwidth ranges that the neural network accelerator can obtain when it is in a working state.

[0035] In a third aspect, the present application provides a neural network model calculation task processing device, the neural network model calculation task processing device comprising:

[0036] A bandwidth perception information acquisition module, used to acquire bandwidth perception information of the electronic device; the bandwidth perception information represents bandwidth resource allocation of at least one hardware device of the electronic device, and the at least one hardware device includes a neural network accelerator to implement the computing task of the neural network model;

[0037] a target execution file selection module, configured to select, from a plurality of candidate execution files corresponding to the target computing task to be executed, a candidate execution file matching the bandwidth perception information as the target execution file to be executed currently; the plurality of candidate execution files respectively correspond to different bandwidth ranges and are used to complete the same computing task;

[0038] A target execution file loading module is used to load the target execution file into the neural network accelerator so that the neural network accelerator executes the target execution file to complete the target computing task.

[0039] In a fourth aspect, the present application provides a neural network model compilation device, the neural network model compilation device comprising:

[0040] A bandwidth range acquisition module, used to acquire multiple bandwidth ranges that can be acquired by a neural network accelerator in a working state for executing a computing task of the neural network model;

[0041] A compilation parameter configuration module, configured to configure compilation parameters matching each of the multiple bandwidth ranges for each of the computing tasks according to the multiple bandwidth ranges;

[0042] The compiling module is used to compile the original execution files of the corresponding computing tasks according to different compiling parameters to obtain multiple candidate execution files of the corresponding computing tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the accompanying drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and the originals and elements are not necessarily drawn to scale.

[0044] Figure 1 A schematic diagram of the hardware structure of an electronic device proposed in this application;

[0045] Figure 2 A schematic diagram of a scenario of a neural network model computing task processing method and a compiling method applicable to an electronic device;

[0046] Figure 3 A flowchart of a neural network model computing task processing method proposed in Example 1 of the present application;

[0047] Figure 4 A flowchart of a neural network model computing task processing method proposed in Example 2 of the present application;

[0048] Figure 5 A flowchart of a neural network model computing task processing method proposed in Example 3 of the present application;

[0049] Figure 6 A flowchart of a neural network model computing task processing method proposed in Example 4 of the present application;

[0050] Figure 7 A flowchart of a neural network model compilation method proposed in Example 1 of the present application;

[0051] Figure 8 A flowchart of a neural network model compilation method proposed in Example 2 of the present application;

[0052] Fig. 9 A flowchart of a neural network model compilation and computing task processing method proposed in an embodiment of the present application;

[0053] Fig.10 A schematic diagram of the structure of a neural network model computing task processing device provided in an embodiment of the present application;

[0054] Fig.11 A schematic diagram of the structure of a neural network model compilation device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0055] The embodiments of the present application are described below in conjunction with the drawings in the embodiments of the present application. The terms used in the implementation mode of the present application are only used to explain the specific embodiments of the present application and are not intended to limit the present application. The embodiments of the present application are described below in conjunction with the drawings. It is known to those of ordinary skill in the art that with the development of technology and the emergence of new scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.

[0056] The terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and need not be used to describe a specific order or sequential order. It should be understood that the terms used in this way can be interchangeable under appropriate circumstances, which is only to describe the distinction mode adopted by the objects of the same attributes when describing in the embodiments of the present application. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, so that the process, method, system, product or equipment comprising a series of units need not be limited to those units, but may include other units that are not clearly listed or inherent to these processes, methods, products or equipment.

[0057] In order to solve the above problems, the embodiments of the present application provide a neural network model computing task processing method and a neural network model compilation method. The present application can be applied to application fields of many industries based on artificial intelligence technology, which may include computer vision fields such as target detection, scene recognition or image segmentation; natural language processing fields such as automatic search engines, dialogue service robots, text classification or intelligent translation, such as speech processing fields such as speech recognition and voiceprint recognition, as well as physics (high-energy particle collision classification or cosmic celestial body map data analysis, etc.), biology / medicine (such as protein folding prediction, etc.) and one or more application fields, so as to promote the implementation of artificial intelligence technology in various industries through hardware devices specially used to accelerate neural network computing tasks in electronic devices, such as face recognition, security monitoring, medical image analysis, speech recognition, machine translation, sentiment analysis, autonomous driving, smart home or smart city, etc., so as to meet the actual application needs reliably, efficiently and with high performance.

[0058] It should be noted that the present application can be, but is not limited to, applied to terminal devices (such as smart phones, tablet computers, wearable devices, vehicle-mounted devices, augmented reality (AR) / virtual reality (VR) devices, laptops, edge devices, robots, smart medical / transportation equipment or other Internet of Things devices, etc.) or servers (such as cloud servers or physical servers, etc.) with data processing capabilities in the corresponding fields listed above, or through communication between terminal devices and servers, to implement the neural network model computing task processing method and neural network model compilation method proposed in the present application. The present application does not impose any restrictions on the operating environment of the application scenario.

[0059] Generally speaking, for computing-intensive tasks such as model training, large-scale data processing, and complex reasoning tasks, the method proposed in this application can be implemented by the server; for real-time processing tasks based on optimized reasoning models with low latency and low power consumption, the method proposed in this application can be implemented by the terminal device; for computing tasks that need to be implemented across platforms, the method proposed in this application can be implemented by the terminal device and the server in cooperation, such as the neural network model compilation method implemented by the terminal device, the neural network model computing task processing method implemented by the server, etc., which can be determined in combination with the task requirements, hardware resources, and performance requirements of the application scenario. The following first describes in detail some operating environments suitable for the method proposed in this application in conjunction with the accompanying drawings.

[0060] Reference Figure 1 , is a schematic diagram of the hardware structure of an electronic device proposed in this application. As analyzed above, the electronic device can be at least one of a terminal device or a server, such as Figure 1 As shown, the electronic device may include but is not limited to: at least one communication element 11, at least one memory 12 and at least one processor 13, wherein:

[0061] At least one communication element 11, at least one memory 12 and at least one processor 13 can communicate with each other via a bus. The bus can be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 1 Only one bidirectional line is used in the diagram, but this does not mean that there is only one bus or only one type of bus.

[0062] The communication element 11 can be used to obtain bandwidth perception information of the electronic device, input data of the computing task implemented based on the neural network model (which can be determined according to the actual application scenario, including but not limited to one or more of audio data, text data, image data and video data), etc., as well as various intermediate data generated during the implementation of the neural network model computing task processing method proposed in the embodiment of the present application, and / or various intermediate data generated in the neural network model compilation method proposed in the embodiment of the present application, as well as to realize data or instruction transmission between internal components of the electronic device, etc., depending on the situation.

[0063] Based on this, in the embodiment of the present application, the communication element 11 may include a communication element corresponding to one or more wireless communication methods such as WIFI, Bluetooth, 5G / 6G, Global System of Mobile communication (GSM), General Packet Radio Service (GPRS), etc., so that the electronic device can realize data transmission with other devices through the communication element. Of course, in order to realize data transmission between the components inside the electronic device, the communication element 11 may also include one or more interfaces supporting wired communication methods, such as a general-purpose input / output (GPIO) interface, a USB interface, a universal asynchronous receiver / transmitter (UART) interface, and one or more combinations thereof. The present application does not limit the composition structure of the communication element 11 to realize this function and its corresponding communication transmission mechanism.

[0064] The memory 12 can be used to store multiple first computer instructions for implementing the neural network model computing task processing method proposed in the embodiment of the present application, and / or multiple second computer instructions for implementing the neural network model compilation method proposed in the embodiment of the present application. The processor 13 can load and execute the first computer instruction to implement the neural network model computing task processing method proposed in the embodiment of the present application, or load and execute the second computer instruction to implement the neural network model compilation method proposed in the present application. The implementation process can refer to the description of the corresponding embodiment below, and this embodiment is not described here.

[0065] In the embodiment of the present application, the memory 12 may include storage media such as a floppy disk, a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk. The processor 13 may include any one or more combinations of a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), a digital signal processor (DSP), a tensor processing unit (TPU), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), and other dedicated neural network accelerators.

[0066] In one possible implementation, refer to Figure 2 As shown in the flowchart, a CPU processor can be used to run a compilation tool, and according to the neural network model compilation method proposed in the embodiment of the present application, the neural network model can be compiled into an execution file suitable for a specific hardware device (such as various types of neural network accelerators or other hardware devices, etc.), that is, multiple candidate execution files for implementing the same computing task of the neural network model. During operation, the CPU can select a candidate execution task that matches the bandwidth perception information from multiple candidate execution files corresponding to the target computing task to be executed according to the neural network model computing task processing method proposed in the embodiment of the present application, based on the bandwidth perception information of the electronic device, as the current target execution file to be executed, and load it into at least one neural network accelerator for execution to implement the target computing task. The implementation process can refer to the description process of the corresponding method embodiment below.

[0067] Among them, the compilation tools may include but are not limited to: official tools and libraries such as OpenVINO (Open Visual Inference & Neural Network Optimization Toolkit, an open source tool suite for accelerating the development and deployment of computer vision and deep learning applications), cuDNN (Compute Unified Device Architecture Deep Neural Network library, a set of GPU acceleration libraries designed specifically for deep learning, including many optimized deep learning operators such as convolution, pooling, normalization, etc.), SDKs (Software Development Kits) provided by hardware manufacturers, such as QNNPACK (Quantized Neural Network PACKage, a high-performance kernel library) and other open source tools. One or more combinations thereof can be selected based on the hardware device that executes the target execution files of each computing task to obtain the candidate execution files of the corresponding computing task. This application does not limit the type of compilation tools for obtaining multiple candidate execution files corresponding to each computing task of the neural network model, which can be determined according to the circumstances.

[0068] It should be understood that Figure 1 The structure of the electronic device shown does not constitute a limitation on the electronic device in the embodiment of the present application. In practical applications, the electronic device may include Figure 1 More or fewer components as shown, or a combination of certain components, such as when the electronic device is a terminal device, may also include microphones, speakers, displays, various sensors, antennas, power modules, RF components, external ports and other input / output components, etc., which can be determined based on processing function requirements, and this application does not give detailed examples one by one.

[0069] In combination with the above description of the electronic device and its operating scenario, the neural network model computing task processing method of the embodiment of the present application will be introduced in detail with reference to the accompanying drawings.

[0070] Reference Figure 3 , is a flowchart of a neural network model computing task processing method proposed in the first embodiment of the present application. The method can be applied to electronic devices as described in the above embodiments to implement neural network model computing tasks in application scenarios such as face recognition, security monitoring, medical image analysis, speech recognition, machine translation, sentiment analysis, autonomous driving, smart home or smart city. Based on this, when an electronic device runs a neural network model in any application scenario to implement the corresponding task, Figure 3As shown, the neural network model calculation task processing method performed by the electronic device may include but is not limited to the following steps:

[0071] Step S31, obtaining bandwidth perception information of the electronic device; the bandwidth perception information represents bandwidth resource allocation of at least one hardware device of the electronic device, and the at least one hardware device includes a neural network accelerator to implement a computing task of a neural network model;

[0072] In the embodiment of the present application, in the process of realizing data processing tasks (such as image segmentation, scene recognition, intelligent translation, text classification or speech recognition, etc.) in the corresponding application scenarios based on the neural network model, the data processing tasks realized based on the neural network model can be divided into different computing tasks according to the various processing stages or processing steps of the data processing tasks realized by the neural network model, or the multiple computing steps contained in each processing stage / step, the computing steps performed by each component contained in the neural network model, etc. For example, preprocessing such as normalization, resizing or enhancement of input data; forward propagation such as linear operations, activation operations and pooling operations; various computing tasks such as post-processing such as classification or threshold processing, model optimization and parallel computing, etc., and even computing tasks can be divided according to the neural network layer to determine the computing tasks based on each neural network layer itself or the computing tasks realized by multiple neural network layers. This application does not limit the computing task determination method, task content and task quantity of the neural network model, which can be determined according to the situation.

[0073] In practical applications, combined with the above analysis, it is known that the bandwidth resources required or allocated to different computing tasks of the same or different neural network accelerators to implement the neural network model may be different. Since the bandwidth resources of the electronic device or the hardware devices that need to use the bandwidth resources may change dynamically, the bandwidth resources that can be allocated to the same neural network accelerator to implement the same type of computing tasks at different times may be different. In other words, during the actual operation of the neural network accelerator in the electronic device, the actual allocated available bandwidth resources are dynamically changing, which can change with the changes in the bandwidth resource allocation of each hardware device in the electronic device that requires bandwidth resources to run.

[0074] Among them, the neural network accelerator in this application may refer to a hardware device specifically used to accelerate the computing tasks of the neural network model, such as one or more of the NPC, CPU or dedicated processor listed above, which can significantly improve the training and training speed of the neural network model and reduce power consumption through parallel computing, dedicated instructions, optimized storage architecture, etc. This application does not limit the working principle of each neural network accelerator, which can be determined based on the optimized network structure of the neural network model used in the application scenario and the computing tasks in the application scenario. It should be understood that if there are multiple neural network accelerators, the embodiment of the present application can obtain the available bandwidth resources allocated to each neural network accelerator, and the acquisition method is not described in detail in this embodiment.

[0075] According to the above analysis, in order to ensure the optimal performance of the neural network accelerator, avoid the situation that the available bandwidth resources actually allocated to it are insufficient to support its computing tasks and cause task failure, or the remaining bandwidth resources for computing tasks are too much, resulting in resource waste, and reducing the accuracy of computing tasks, etc., the present application proposes that during the compilation of the neural network model part for each computing task, different bandwidth ranges are used to compile multiple candidate execution files for the same computing task, that is, based on the compilation parameters that match different bandwidth ranges for the original execution file of each computing task, multiple candidate execution files for implementing the computing task are compiled, so that the most suitable candidate execution file can be selected from them as the target execution file for implementing the computing task during operation. Among them, the compilation process of the candidate execution file can refer to the neural network model compilation method described in the following embodiment, and this embodiment will not be described in detail here.

[0076] In order to accurately select the appropriate target execution file for each computing task, the embodiment of the present application proposes to obtain the actual bandwidth perception information of the electronic device during the operation of the neural network model and before executing the computing task, so as to determine the actual bandwidth resource allocation of at least one hardware device of the electronic device (such as a neural network accelerator, and may also include other processors or communication elements, etc.), accurately understand whether the hardware device currently running the device uses the bandwidth resources configured by the electronic device, and / or the usage of the bandwidth resources by the hardware device (that is, the available bandwidth resources that can be allocated), etc. The present application does not limit the content of the bandwidth perception information and the method of obtaining it.

[0077] It should be noted that the present application does not limit the neural network accelerator to obtain the bandwidth perception information of the current electronic device before executing each computing task. According to the actual situation, step S31 can also be repeated periodically, according to preset rules, or after certain conditions are met. However, before implementing the first computing task of the neural network model, step S31 is usually required to determine the actual bandwidth resource allocation of the electronic device running the neural network model.

[0078] Step S32, selecting, from among a plurality of candidate execution files corresponding to the target computing task to be executed, a candidate execution file matching the bandwidth perception information as the target execution file to be executed currently, according to the bandwidth perception information; the plurality of candidate execution files respectively correspond to different bandwidth ranges and are used to complete the same computing task;

[0079] Following the above analysis, in order to be applicable to operating scenarios such as different operating times, bandwidth resource allocation situations or different neural network accelerators, each neural network accelerator can implement computing tasks with high performance, this application no longer uses a fixed execution file (which is compiled based on fixed bandwidth resources) to execute the computing task to implement the computing task, but pre-obtains multiple candidate execution files for implementing the same computing task, and these multiple candidate execution files each correspond to a different bandwidth range, so that the neural network accelerator executes different candidate execution files of the same computing task in the same operating scenario, and the task performance that can be achieved is different.

[0080] Based on this, when the neural network accelerator implements its target computing task to be executed (which can be the computing task currently required to be implemented by the neural network accelerator among the various computing tasks of the neural network model), in different operation scenarios of the electronic device, in order to make the performance of the neural network accelerator to achieve the expected optimal performance, it is necessary to dynamically and flexibly select the most suitable candidate execution file from the various candidate execution files for implementing the target computing task, that is, execute the candidate execution file to implement the target computing task, which will neither waste the available bandwidth resources allocated to the neural network accelerator nor use it as the current target execution file to be executed for implementing the target computing task because the actual available bandwidth resources allocated are insufficient to support the neural network accelerator to implement the target computing task. It can be seen that the target execution file selected for the same target computing task in different operation scenarios of the electronic device can be different.

[0081] For the operation scenario of the electronic device as described above, it can be determined by the currently acquired bandwidth perception information of the electronic device to match it with the bandwidth range corresponding to each candidate execution file corresponding to the target computing task, and according to the corresponding bandwidth matching result, the candidate execution file corresponding to the bandwidth range matching the current bandwidth perception information is selected as the target execution file, that is, the execution file that needs to be executed by the neural network accelerator to implement the current target computing task. This application does not elaborate on the selection process of the target execution file. As analyzed above, it should be understood that as the bandwidth perception information of the electronic device changes, the target execution file selected for the same target computing task to be executed may be different, and under the same bandwidth perception information, the bandwidth range corresponding to the target execution file selected for different target computing tasks to be executed may be different.

[0082] In an embodiment of the present application, each candidate execution file corresponding to a computing task can represent the processing logic of the computing task that implements the neural network model, and can include the operation code that implements the processing logic. If needed, it can also include the weight parameters of the neural network model required to implement the processing logic, etc. Of course, the neural network accelerator can also call the corresponding weight parameters of the neural network model according to the operation requirements during the execution of the operation code, that is, the candidate execution file may not contain the weight parameters required during the execution of the operation code, and the present application does not limit the content of each candidate execution file.

[0083] Among them, in actual applications, in the process of determining multiple different bandwidth ranges for the same computing task and compiling multiple candidate execution files to implement the computing task, it is mainly achieved by adjusting the original processing logic represented by the original execution file of the computing task, so that the adjusted processing logic represented by each of the multiple candidate execution files is different, but the weight parameters required for each of the multiple candidate execution files to implement the computing task are usually the same. This application does not limit how to adjust the implementation process of the original processing logic for implementing the computing task according to different bandwidth ranges, as well as the adjusted processing logic content and the general weight parameter content required to implement the computing task.

[0084] It should be noted that for multiple candidate execution files that implement the same computing task of the neural network model, the electronic device can run at least one of the compilation tools listed above to obtain them, and the candidate execution files corresponding to different computing tasks can be obtained by running corresponding different compilation tools, especially in different scenarios of hardware devices that implement different computing tasks. In order to obtain candidate execution files suitable for the hardware device, matching compilation tools can be used to obtain them. The implementation process of this application embodiment is not described in detail here. Among them, in the process of obtaining multiple candidate execution files corresponding to any computing task, the original processing logic of the neural network model to implement the computing task can be converted into operation codes suitable for execution by hardware devices (such as neural network accelerators) for their respective corresponding bandwidth ranges, and the hardware device can be guaranteed to use bandwidth resources within the corresponding bandwidth range, so that the computing task can be efficiently and reliably implemented.

[0085] It can be seen that the multiple candidate execution files corresponding to the same computing task represent different processing logics for implementing the computing task. The more complex the processing logic is, the more bandwidth resources the hardware device requires, that is, the more bandwidth resources the corresponding bandwidth range contains, and usually the higher the accuracy of implementing the same computing task. This application does not limit the content of the multiple candidate execution files corresponding to each computing task and the method of obtaining the files.

[0086] Step S33, loading the target execution file into the neural network accelerator so that the neural network accelerator executes the target execution file to complete the target computing task.

[0087] In the process of running a neural network model on an electronic device to realize any application scenario task such as face recognition, security monitoring, medical image analysis, speech recognition, machine translation, sentiment analysis or autonomous driving, for each target computing task to be executed (which may be a computing task currently to be executed during the sequential execution of multiple computing tasks of the neural network model to realize the corresponding application scenario task), according to the method described above, Figure 2 As shown, after dynamically selecting a target execution file that matches the bandwidth perception information of the current electronic device, the target execution file is loaded into the neural network accelerator, and the neural network accelerator executes the target execution file to complete the target computing task, thereby ensuring that the neural network accelerator can achieve optimal task performance under the current bandwidth allocation of at least one hardware device of the electronic device.

[0088] It can be seen that, compared with obtaining fixed execution files corresponding to each computing task of the neural network model, the present application obtains multiple candidate execution files corresponding to each computing task before running the neural network model, and these multiple candidate execution files correspond to different bandwidth ranges. In this way, before executing the current target computing task to be executed, first, based on the bandwidth perception information of the current electronic device, from the multiple candidate execution files corresponding to the target computing task, quickly and reliably select the target execution file that matches the bandwidth perception information, that is, accurately select the target execution file that enables the current neural network accelerator to achieve the expected optimal performance of the target computing task, and then load the selected target execution file into the neural network accelerator for execution to complete the target computing task, which reliably avoids the mismatch between the bandwidth range based on the target execution file acquisition process and the available bandwidth resources allocated to the neural network accelerator during actual operation, resulting in the neural network accelerator being unable to achieve the expected optimal performance, such as the target computing task performance being reduced or even the target computing task being unable to be achieved due to the actual small available bandwidth resources allocated, thereby improving the processing efficiency and reliability of the neural network model computing tasks.

[0089] Reference Figure 4 , is a flowchart of the neural network model calculation task processing method proposed in the second embodiment of the present application. This embodiment can describe an optional implementation method of how to obtain the bandwidth perception information of the electronic device in the neural network model calculation task processing method proposed below. The implementation process of other steps of the neural network model calculation task processing method can refer to the description of the corresponding part of the context embodiment, and this embodiment will not be repeated. Based on this, Figure 4 As shown, the implementation method of obtaining the bandwidth perception information of the electronic device may include but is not limited to:

[0090] Step S41, sending a bandwidth test instruction to the neural network accelerator to execute the target computing task, so that the neural network accelerator executes the bandwidth test instruction at least once to obtain bandwidth test data; the bandwidth test instruction includes at least part of the computing instructions required for the target computing task;

[0091] Step S42, receiving the bandwidth test data fed back by the neural network accelerator;

[0092] Step S43: Obtain bandwidth perception information for the neural network accelerator based on the bandwidth test data.

[0093] According to the above analysis of the technical solution of the present application, during the operation of a neural network model used to implement any of the application scenario tasks listed above, for multiple computing tasks of the neural network model, before executing the computing task that currently needs to be executed (i.e., the target computing task to be executed), in order to be able to select a target execution file that can achieve the optimal task performance from its corresponding multiple candidate execution files, an embodiment of the present application proposes to perform a bandwidth test on a neural network accelerator used to implement the target computing task to determine the available bandwidth resources currently allocated to the neural network accelerator, and / or whether the electronic device has other hardware devices that compete with the neural network accelerator for bandwidth resources, and determine bandwidth perception information such as the number of other hardware devices or the bandwidth resources competed for.

[0094] Therefore, before loading the execution file to the neural network accelerator to execute the target computing task, a bandwidth test instruction is sent to the neural network accelerator so that the neural network accelerator obtains bandwidth test data by executing the bandwidth test instruction at least once. The bandwidth test data may include data required for calculating bandwidth perception information, such as bandwidth test data such as the amount of access data and the running time during the execution of the bandwidth test instruction in order to calculate the available bandwidth resources currently allocated to the neural network accelerator, so as to obtain the available bandwidth resources of the neural network accelerator by calculating the amount of access data per unit time, and even to determine the bandwidth resource competition situation between the neural network accelerator and other hardware devices in the electronic device by calculating the ratio of the available bandwidth resources to the total bandwidth resources currently configured by the electronic device. The present application does not limit the implementation method of how the neural network accelerator obtains the bandwidth test data, and the content of the bandwidth test data.

[0095] Among them, the bandwidth test instructions may include at least part of the computing instructions required for the target computing task, such as at least part of the computing instructions in the operation code used to implement the target computing task described above, so as to determine the bandwidth perception information of the neural network accelerator by recording the start time and end time of the neural network accelerator executing the bandwidth test instructions, as well as the bandwidth test data such as the IO access data volume during the execution (such as input and output access to the port for external communication), so as to evaluate the performance of the current neural network accelerator and determine which bandwidth range the execution file is suitable for.

[0096] In some embodiments, in the process of determining the bandwidth test instruction, it is possible to determine which computing instructions the bandwidth test instruction needs to include based on one or more of the bandwidth test requirements of the data processing task to be implemented by the neural network model, the bandwidth test requirements of the target computing task to be executed, or the bandwidth test requirements of the corresponding neural network accelerator, and may include the corresponding contents of each computing instruction. Among them, the bandwidth test requirements (which can also be said to be the bandwidth test effect) may include but are not limited to one or more test indicators such as bandwidth test time, bandwidth test accuracy, and resource consumption of bandwidth test, so that the bandwidth test instruction obtained thereby is sent to the corresponding neural network accelerator for bandwidth testing, which can achieve the expected bandwidth test effect, that is, each actual test indicator meets the bandwidth test requirements.

[0097] According to the above analysis, in a possible implementation, the bandwidth test instruction may include at least some of the computing instructions included in the computing task actually possessed by the neural network model, that is, some of the computing instructions actually to be executed to implement the computing task. In the process of selecting this part of the computing instructions, the present application may analyze the expected demand for bandwidth resources when executing each computing instruction in the computing task, and select some of the computing instructions with higher expected demand, such as selecting some of the computing instructions corresponding to the expected demand greater than the bandwidth threshold, or selecting a preset number of computing instructions with a higher ranking after sorting from large to small according to the expected demand, etc. That is to say, the present embodiment may select some of the computing instructions in the computing task that most require bandwidth resources to support execution, to form the bandwidth test instruction.

[0098] In another possible implementation of selecting at least some of the computing instructions included in the bandwidth test instruction, the present application may also select some of the computing instructions from all the computing instructions according to a selection rule such as a preset ratio or selection probability when all the computing instructions included in the computing task are determined. Alternatively, after determining the various instruction types included in all the computing instructions, select at least one computing instruction corresponding to each instruction type, and determine it as the partial computing instructions that need to be included in the corresponding bandwidth test instruction, so that the bandwidth test instruction covers all instruction types in all the computing instructions included in the computing task. Optionally, the present application may also perform an importance analysis on all the computing instructions included in the computing task, and select some computing instructions that are more important for implementing the computing task. The present application does not limit the implementation method of how to select at least some of the computing instructions in the computing task included in the bandwidth test instruction. In practical applications, one or more combination methods listed above may be used for implementation.

[0099] In some other embodiments, the present application may obtain the bandwidth test effect (as represented by one or more test indicators described above) of the neural network accelerator executing the bandwidth test instruction after preliminarily determining the bandwidth test instruction according to but not limited to the method described above, and then adjust the bandwidth test instruction based on the bandwidth test effect, such as adjusting one or more of the instruction content (such as the above-selected partial calculation instructions or other instructions) and the number of instructions (such as the above-selected partial calculation instructions or other instructions) contained in the bandwidth test instruction, so as to obtain a bandwidth test instruction (i.e., the adjusted bandwidth test instruction) that meets the bandwidth test requirements (such as meeting the bandwidth test accuracy requirements).

[0100] Optionally, after each adjustment of the bandwidth test instruction, the present application can re-acquire the bandwidth test effect of the neural network accelerator executing the adjusted bandwidth test instruction. If the bandwidth test effect does not meet the bandwidth test requirements, continue to adjust the bandwidth test instruction; if the bandwidth test effect meets the bandwidth test requirements, store the final adjusted bandwidth test instruction to implement actual bandwidth testing of the neural network accelerator.

[0101] Optionally, in the process of adjusting the bandwidth test instructions, the present application can also adjust the bandwidth test instructions by measuring multiple test indicators such as bandwidth test time and bandwidth test accuracy, so that the bandwidth test of the neural network accelerator is performed based on the adjusted bandwidth test instructions, and the actual bandwidth test time generated is shorter and has higher bandwidth test accuracy, thereby ensuring the accuracy and reliability of the bandwidth perception information obtained thereby, and further improving the accuracy of the target execution file selected thereby for implementing the computing task. It should be noted that the implementation method for determining the content of the bandwidth test instruction includes but is not limited to the several methods listed above.

[0102] Optionally, the bandwidth test instruction can also be a general bandwidth test code file applicable to the corresponding neural network accelerator and communication port, which can be constructed during the compilation phase of the neural network model. The theoretical IO access volume (i.e., the above-mentioned data access volume) required for the operation of the bandwidth test code file can also be recorded as needed. In order to reduce the waste of bandwidth resources, the theoretical IO access volume is minimized as much as possible without affecting the bandwidth test function. In this way, the bandwidth test code file is loaded into the neural network accelerator for execution. After obtaining the execution time of the bandwidth test code file, the execution time and IO access volume can be used to obtain the bandwidth perception information of the neural network accelerator.

[0103] It should be noted that in the actual application of this application, for each computing task corresponding to the neural network model, it may be completed by the same neural network accelerator according to the method proposed in this application, or it may be completed by the cooperation of different neural network accelerators. That is to say, in the case where the electronic device includes multiple neural network accelerators, different computing tasks may be distributed to different neural network accelerators. The distribution relationship between each computing task and different neural network accelerators can be determined based on the respective configuration information and hardware resources of different neural network accelerators, as well as the task requirements of each computing task. This application does not limit the content of this distribution relationship and its acquisition method.

[0104] Among them, for each neural network accelerator to which at least one computing task is distributed, when the computing task distributed to it is the target computing task to be currently executed, the bandwidth test of the neural network accelerator can be realized according to the method described in this embodiment, and the bandwidth perception information of the neural network accelerator can be obtained. The implementation process is similar, and this application will not elaborate one by one. In addition, during the bandwidth test process of each neural network accelerator, a bandwidth test instruction can be executed according to the method described above to obtain the corresponding bandwidth test data, or the bandwidth test instruction can be executed multiple times, and the mean operation can be performed on the bandwidth test data obtained from each test, or the bandwidth test data obtained from each test can be summed to obtain the corresponding bandwidth test data, so as to improve the reliability and accuracy of the bandwidth perception information.

[0105] In the actual application of this application, the neural network accelerator of the electronic device can be one or more of the above-mentioned CPU, GPU, FPGA, TPU, or other dedicated neural network acceleration chips, etc. The configuration information of multiple neural network accelerators of the same type may be the same or different, so as to meet the needs of different computing tasks. Exemplarily, for large-scale matrix computing tasks in the neural network model, it can be distributed to a GPU with excellent parallel computing capabilities to improve the task execution efficiency; for matrix computing and tensor operation computing tasks, it can be distributed to a TPU to reduce the task execution power consumption; for each subtask included in complex control processes and lightweight neural network tasks, it can be distributed to the CPU to utilize the multi-core architecture and multi-thread model of the CPU to improve the task processing efficiency, etc. This application does not limit the number and type of neural network accelerators for implementing each computing task of the neural network model, which can be determined according to the situation.

[0106] In addition, in a possible implementation, according to the neural network model computing task processing method described in the above embodiment, during the operation of the neural network model, before executing each target computing task to be executed, the neural network accelerator used to implement the target computing task can be tested according to the bandwidth perception information acquisition method described above, that is, a bandwidth test instruction is sent to the neural network accelerator of the target computing task to be executed, so as to Figure 4 The method shown obtains bandwidth perception information for the neural network accelerator, which is used to select a target execution file matching the bandwidth perception information from multiple candidate execution files corresponding to the target computing task and load it into the neural network accelerator to implement the target computing task, thereby ensuring that the neural network accelerator can achieve the optimal task performance under the current bandwidth resource allocation conditions.

[0107] In combination with the above analysis, in some embodiments, referring to Figure 5 The flowchart of the neural network model computing task processing method proposed in the third embodiment of the present application is shown in FIG. 1 , where the computing tasks of the neural network model include different computing tasks corresponding to each neural network layer, such as Figure 5 As shown, the neural network model computing task processing method proposed in this embodiment may include the following steps:

[0108] Step S51, in the process of sequentially executing the computing tasks corresponding to each neural network layer of the neural network model, before each execution of the target computing task corresponding to the current neural network layer to be executed, sending a bandwidth test instruction to the neural network accelerator of the target computing task to be executed, so that the neural network accelerator executes the bandwidth test instruction at least once to obtain bandwidth test data;

[0109] Step S52, receiving bandwidth test data fed back by the neural network accelerator;

[0110] Step S53, obtaining bandwidth perception information for the neural network accelerator according to the bandwidth test data;

[0111] Step S54, selecting, from a plurality of candidate execution files corresponding to the target computing task, a candidate execution file matching the bandwidth perception information as the target execution file to be executed currently, according to the bandwidth perception information;

[0112] Step S55, loading the target execution file into the neural network accelerator so that the neural network accelerator executes the target execution file to complete the target computing task;

[0113] Following the above analysis, for each computing task of the neural network model, if it is divided according to the various neural network layers contained in the neural network model, that is, each neural network layer corresponds to a computing task, such as the convolution operation task corresponding to the convolution layer, the pooling operation task corresponding to the pooling layer, etc., before running each neural network layer of the neural network model, that is, before executing the target computing task corresponding to the current neural network layer to be executed each time, a bandwidth test instruction is first sent to the neural network accelerator used to implement the target computing task to obtain the current bandwidth perception information for the neural network accelerator. The process of obtaining the bandwidth perception information is not described in detail in this application.

[0114] Afterwards, a candidate execution file that matches the bandwidth perception information obtained from this bandwidth test can be selected from the multiple candidate execution files corresponding to the currently stored target computing task to be executed as the currently executed target execution file, and the target execution file is loaded into the corresponding neural network accelerator, and the neural network accelerator executes the target execution file to complete the corresponding target computing task. In this way, before executing the computing task corresponding to the next neural network layer to be executed (i.e., the new target computing task to be executed), a bandwidth test instruction is resent to the neural network accelerator used to implement the computing task (which may be the same or different from the neural network accelerator that implements the previous target computing task) to obtain the current bandwidth perception information for the neural network accelerator. The implementation process is similar to this application and will not be described in detail. In this way, according to the running order of each neural network layer contained in the neural network model, the target computing task corresponding to the corresponding neural network layer is implemented until the target computing task corresponding to the last neural network layer of the neural network model is completed.

[0115] It can be seen that when the neural network accelerator completes the computing task corresponding to each neural network layer, the executed target execution file matches the bandwidth perception information for the neural network accelerator obtained from the current test, ensuring that the neural network accelerator executes the target execution file and can achieve the expected optimal performance. In particular, in the scenario where the bandwidth resource allocation of at least one hardware device of the electronic device changes dynamically, the processing method proposed in this embodiment reliably avoids the bandwidth resources required to be consumed in the actual execution process exceeding the available bandwidth resources actually allocated to the neural network accelerator, resulting in the current target computing task to be executed. The execution time is too long (that is, the task execution efficiency is low) or the execution fails, or the bandwidth resources actually consumed in the execution process are less than the available bandwidth resources actually allocated to the neural network accelerator, and the bandwidth difference between the two is greater than the threshold, indicating that the bandwidth resources actually consumed in the execution process are far less than the available bandwidth resources actually allocated to the neural network accelerator, that is, the neural network accelerator does not fully utilize its available bandwidth resources in the process of implementing its distributed computing tasks, which not only causes a waste of bandwidth resources, but also reduces the accuracy of the completed computing tasks, thereby affecting the task accuracy of the entire neural network model, such as the accuracy of any application scenario task such as image recognition, text classification or speech recognition, resulting in the inability to reliably meet the task requirements.

[0116] In some embodiments, when an electronic device is configured with multiple neural network accelerators to respectively implement different computing tasks of a neural network model, in order to improve processing efficiency, during the process of executing the target execution file corresponding to the target computing task to be executed currently on the first neural network accelerator, if the next target computing task to be executed is distributed to the second neural network accelerator, it is not necessary to wait for the completion of the target computing task to be executed currently. The method described above can be used to directly perform a bandwidth test on the second neural network accelerator, select the target execution file corresponding to the next target computing task to be executed, and load it directly to the second neural network accelerator. In this way, when the current computing task is completed and the next target computing task to be executed needs to be executed, the second neural network accelerator can be controlled to directly execute the already loaded target execution file corresponding to the next target computing task to be executed, without waiting for the bandwidth test on the second neural network accelerator to select the target execution file corresponding to the next target computing task to be executed, that is, the target execution file corresponding to the new target computing task to be executed, thereby shortening the running time of the entire neural network model and improving task processing efficiency.

[0117] Thus, in the embodiment described in the previous paragraph, for two adjacent computing tasks to be executed of the neural network model, when one of the computing tasks (recorded as the first computing task) is distributed to the first neural network accelerator and the other computing task (recorded as the second computing task) is distributed to the second neural network accelerator, the target execution file execution process of the first neural network accelerator and the bandwidth test and target execution file selection process of the second neural network accelerator can be executed in parallel, so that after the first neural network accelerator completes the first computing task, the second neural network accelerator can directly execute the target execution file corresponding to the second computing task to complete the second computing task.

[0118] It should be understood that if two adjacent computing tasks to be executed are distributed to a neural network accelerator, the method described in the third embodiment above can be used. Before executing each computing task, the neural network accelerator performs a bandwidth test, selects the target execution file corresponding to the target computing task to be executed, and loads it into the neural network accelerator for execution. After completing this target computing task, the bandwidth resource allocation of at least one hardware device of the electronic device may change during this period, resulting in changes in the available bandwidth resources allocated to the neural network accelerator. The bandwidth perception information obtained by the last bandwidth test may no longer be accurate. If the multiple candidate execution files corresponding to the next computing task are selected according to the bandwidth perception information, it is easy to cause the selected target execution file to not be the most suitable candidate execution file at present, that is, it is not a candidate execution file that can obtain the optimal performance of the neural network accelerator. Therefore, the present application can still re-test the bandwidth of the neural network accelerator before executing the next computing task, repeat the above operation, and accurately select the most suitable target execution file corresponding to the next computing task to complete the next computing task.

[0119] In other embodiments, in the process of implementing the neural network model computing task processing method described in the above embodiments, that is, in the process of sequentially executing the computing tasks corresponding to each neural network layer of the neural network model according to the above method, if the computing tasks corresponding to several neural network layers are completed in sequence, the bandwidth resource allocation of at least one hardware device of the electronic device may be stabilized, and even the available bandwidth resources allocated to the neural network accelerator may be stabilized. There is no need to frequently perform bandwidth testing on the neural network accelerator in the manner described in the above embodiments. By reducing the bandwidth testing frequency of the neural network accelerator, the bandwidth testing time during the entire neural network model operation can be shortened, thereby improving the processing efficiency of the neural network model computing tasks. Based on this, the bandwidth testing frequency can be adjusted according to but not limited to the method described in the following steps.

[0120] Step S56, in the process of sequentially executing the computing tasks corresponding to each neural network layer of the neural network model, if multiple bandwidth test data continuously fed back by the neural network accelerator are obtained, the change amount of the available bandwidth resources allocated to the neural network accelerator is determined according to the multiple bandwidth test data;

[0121] Step S57, according to the change in the available bandwidth resources, adjust the bandwidth test frequency of the neural network accelerator, so that before the target computing task corresponding to the corresponding neural network layer to be executed is executed, a bandwidth test instruction is sent to the neural network accelerator to execute the target computing task according to the adjusted bandwidth test frequency, and return to step S52 to continue to implement the target computing task until all computing tasks of the neural network model are completed.

[0122] Following the above analysis, during the operation of the neural network model, the bandwidth test data obtained from each bandwidth test of the neural network accelerator can be recorded, and the difference between the bandwidth test data obtained from two adjacent tests can be calculated. For example, the difference between the available bandwidth resources obtained from two adjacent bandwidth tests of the same neural network accelerator can be calculated to obtain the change in the available bandwidth resources of the neural network accelerator, which can be used to evaluate the bandwidth fluctuation of the neural network accelerator. If the change in the available bandwidth resources is large, it can be considered that the bandwidth fluctuation of the neural network accelerator is large and not stable enough. In order to ensure the performance of the neural network accelerator, the bandwidth test of the corresponding neural network accelerator can still be performed according to the method described above before executing the computing task corresponding to each neural network layer, so as to accurately select the most suitable target execution file at the moment.

[0123] On the contrary, if the change in available bandwidth resources allocated to the neural network accelerator is small, or the change in available bandwidth resources is within a smaller bandwidth fluctuation range, it can be considered that the bandwidth fluctuation of the neural network accelerator is small, that is, the available bandwidth resources allocated to the neural network accelerator during this period are stable, and the neural network accelerator does not need to be tested for bandwidth before executing the computing tasks corresponding to each neural network layer. The bandwidth test frequency of the neural network accelerator can be adjusted (reduced here) according to the currently obtained change in available bandwidth resources. For example, after completing the computing tasks corresponding to several neural network layers (such as 3 or 4 consecutive neural network layers, which can be determined according to the adjusted bandwidth test frequency), the neural network accelerator is tested for bandwidth once. In this way, the number of bandwidth tests on the neural network accelerator during the operation of the entire neural network model can be reduced, the total bandwidth test time can be shortened, and the task processing efficiency can be improved.

[0124] In a possible implementation, during the adjustment of the bandwidth test frequency of the neural network accelerator, if the currently obtained change in available bandwidth resources is less than or equal to a first threshold (which can be determined based on the actual application scenario task requirements or experiments or experience, and the present application does not limit the value of the first threshold), it means that the available bandwidth resources allocated to the neural network accelerator during this period are basically stable, and the bandwidth test frequency of the neural network accelerator can be reduced, such as increasing the number of computing tasks corresponding to the neural network layer completed between two adjacent bandwidth tests, that is, the number of target execution files executed by the neural network accelerator, from 1 in the above step (which corresponds to the first bandwidth test frequency) to 2 or 3 or 4, etc., and can be adjusted according to a pre-configured default value, that is, the currently executed first bandwidth test frequency is adjusted to a pre-configured second bandwidth test frequency, and the first bandwidth test frequency is greater than the second bandwidth test frequency. The present application does not limit the value of the second bandwidth test frequency.

[0125] Optionally, the present application may also select a matching candidate bandwidth test frequency as the second bandwidth test frequency from the pre-configured candidate bandwidth test frequencies based on the currently obtained change in available bandwidth resources, or the difference between the change in available bandwidth resources and the first threshold, etc. The current bandwidth test frequency may also be adjusted according to the pre-configured bandwidth test frequency adjustment rule based on the currently obtained change in available bandwidth resources, or the difference between the change in available bandwidth resources and the first threshold, to obtain an adjusted bandwidth test frequency, such as the above-mentioned second bandwidth test frequency, etc.

[0126] Similarly, if the currently obtained change in available bandwidth resources is greater than the first threshold, it means that the available bandwidth resources allocated to the neural network accelerator during this period are unstable. If the current bandwidth test frequency is the first bandwidth test frequency (that is, before executing the computing task corresponding to each neural network layer to be executed, a bandwidth test is performed on the neural network accelerator), the neural network accelerator can continue to be tested according to the first bandwidth test frequency; if the current bandwidth test frequency is the second bandwidth test frequency or other bandwidth test frequencies less than the first bandwidth test frequency, the bandwidth test frequency of the neural network accelerator can also be increased according to the change trend of the available bandwidth resource change. Subsequently, the increased bandwidth test frequency can be recorded as the third bandwidth test frequency (which is less than or equal to the first bandwidth test frequency) to continue to perform bandwidth testing on the neural network accelerator. Among them, the bandwidth test process of each neural network accelerator is similar, and the description of the corresponding part of the above embodiment can be referred to, and this embodiment will not be repeated here.

[0127] From the above analysis, it can be seen that during the operation of the entire neural network model, by dynamically adjusting the bandwidth test frequency of the neural network accelerator according to the method described above, it can be ensured that for each target computing task to be executed, a suitable candidate execution file for the current bandwidth allocation of the electronic device is accurately selected as the target execution file to be executed, so as to obtain the optimal performance of the neural network accelerator to achieve the target computing task. At the same time, when the available bandwidth resources of the neural network accelerator are stable, the bandwidth test frequency can be reduced in time, thereby shortening the bandwidth test time, reducing the resource consumption of the bandwidth test process, and improving the task processing efficiency.

[0128] It should be noted that after completing the adjustment of the bandwidth test frequency once, you can still continue to monitor the change in the available bandwidth resources actually allocated to the neural network accelerator according to the method described above, so that after the comparison result between it and the first threshold changes, you can adjust the current bandwidth test frequency in time according to but not limited to the adjustment method described above, so as to ensure the performance of the neural network accelerator and improve the operating efficiency of the entire neural network model and the quality of the tasks achieved.

[0129] In some other embodiments proposed in the present application, starting from the neural network model, a bandwidth test instruction can be sent to the neural network accelerator of the target computing task to be executed according to the pre-configured preset bandwidth test frequency before the corresponding target computing task to be executed is executed, so as to obtain bandwidth perception information for the neural network accelerator according to the bandwidth test method described above, so as to select the most suitable target execution file for currently implementing the target computing task, load it into the neural network accelerator for execution, complete the target computing task, and ensure the performance of the neural network accelerator. Among them, the preset bandwidth test frequency can be obtained based on empirical values, experimental values, or based on the historical data analysis of the neural network model during the execution of the same type of application scenario tasks, and the present application does not limit the value of the preset bandwidth test frequency.

[0130] In one possible implementation, in the embodiment described in the previous paragraph, each computing task of the neural network model can be completed according to the preset bandwidth test frequency. In another possible implementation, after implementing several bandwidth tests on the neural network accelerator according to the preset bandwidth test frequency and obtaining the corresponding multiple bandwidth test data, the preset bandwidth test frequency can still be dynamically adjusted according to the method described in the corresponding parts of step S56 and step S57, so as to continue to complete the computing tasks corresponding to the subsequent neural network layers of the neural network model according to the adjusted bandwidth test frequency. In this subsequent execution process, the bandwidth test frequency of the neural network accelerator can still be adjusted. The implementation process is not described in detail in this application.

[0131] In addition, when multiple neural network accelerators of an electronic device complete different computing tasks of a neural network model, the bandwidth test frequency of each neural network accelerator can be determined or dynamically adjusted according to the method described above. Optionally, the present application can also adjust the bandwidth test frequency of the multiple neural network accelerators according to the change in available bandwidth resources allocated to each of the multiple neural network accelerators, so as to implement bandwidth testing of each neural network accelerator according to the adjusted bandwidth test frequency. It should be noted that if the difference between the change in available bandwidth resources of each of the multiple neural network accelerators is large, such as greater than the second threshold (which can be determined based on the actual application scenario task requirements or experiments or experience, and the present application does not limit the value of the first threshold), the method described in the above embodiment can be used to implement dynamic adjustment of the bandwidth test frequency of each neural network accelerator to ensure the performance of each neural network accelerator.

[0132] Reference Figure 6 , is a flowchart of the neural network model computing task processing method proposed in the fourth embodiment of the present application. Combining the various detailed implementation methods of the neural network model computing task processing method described in the above embodiments, the selection and implementation process of the target execution file corresponding to each computing task to be executed can be described in detail, such as Figure 6 As shown, the neural network model calculation task processing method proposed in this embodiment may include but is not limited to the following steps:

[0133] Step S61, obtaining bandwidth perception information of the electronic device; the bandwidth perception information represents bandwidth resource allocation of at least one hardware device of the electronic device, and the at least one hardware device includes a neural network accelerator to implement a computing task of a neural network model;

[0134] Regarding the implementation process of step S61, reference may be made to the description of the corresponding parts of the above embodiments, such as the description of the corresponding parts of Embodiment 2 and Embodiment 3, and the embodiments of the present application will not be elaborated here.

[0135] Step S62, determining whether the available bandwidth resources of the neural network accelerator are greater than the bandwidth threshold according to the number of hardware devices participating in bandwidth allocation contained in the bandwidth perception information, or the available bandwidth resources allocated to the neural network accelerator, and if so, executing step S63; if not, proceeding to step S64;

[0136] In an embodiment of the present application, the bandwidth perception information of the electronic device is obtained according to the method described above, which may include the available bandwidth resources allocated to the neural network accelerator and one or a combination of the two information of the number of hardware devices participating in the bandwidth allocation, thereby determining whether the neural network accelerator in the current electronic device monopolizes the total bandwidth resources configured for the electronic device, or whether there are other hardware devices (which can be recorded as second hardware devices) participating in the bandwidth allocation, that is, there are second hardware devices competing with the neural network accelerator for the total bandwidth resources configured for the electronic device. If there are second hardware devices, the number of second hardware devices can be further determined, thereby evaluating the bandwidth competition situation between the second hardware device and the neural network accelerator to determine the available bandwidth resources that can be allocated to the neural network accelerator. The implementation process is not described in detail in this application.

[0137] Among them, if there is no second hardware device in the electronic device that competes with the neural network accelerator for bandwidth resources, that is, the neural network accelerator monopolizes the bandwidth resources, or the number of second hardware devices is small (such as less than the first number threshold, such as 2 or 3, etc., and the present application does not limit its value), it can be determined that the available bandwidth resources of the neural network accelerator are greater than the bandwidth threshold (which can be determined based on the available bandwidth resources of the current electronic device, that is, the above-mentioned total bandwidth resources, such as 80% or 85% of the total bandwidth resources, etc., or can be determined based on experience, and the present application does not limit the value of the bandwidth threshold), that is, it is considered that the neural network accelerator can be allocated with more available bandwidth resources, which is sufficient to execute the candidate execution files corresponding to the maximum bandwidth range, so as to maximize the performance of the neural network accelerator.

[0138] On the contrary, if the number of second hardware devices running in the electronic device that compete with the neural network accelerator for bandwidth resources is large (such as greater than the first number threshold, or even greater than the second number threshold, such as 5 or 8, which are greater than the first number threshold, and this application does not limit the value of the second number threshold), it can be determined that the available bandwidth resources of the neural network accelerator are less than the bandwidth threshold, that is, it is considered that the available bandwidth resources that can be allocated to the neural network accelerator are small and insufficient to execute candidate execution files corresponding to a larger bandwidth range.

[0139] Therefore, the present application can timely and accurately determine whether the available bandwidth resources of the neural network accelerator are greater than the bandwidth threshold through the content of the bandwidth perception information. For different judgment results, different selection rules can be used to implement the selection operation of multiple candidate execution files corresponding to the target computing task to be executed, and determine the most suitable target execution file under the bandwidth resource allocation of each hardware device of the current electronic device to ensure the performance of the neural network accelerator.

[0140] Step S63, selecting a candidate execution file corresponding to the maximum bandwidth range from a plurality of candidate execution files corresponding to the target computing task to be executed as the target execution file to be executed currently;

[0141] Step S64, selecting a candidate execution file that matches the available bandwidth resources of the neural network accelerator from a plurality of candidate execution files corresponding to the target computing task to be executed as the target execution file to be executed currently;

[0142] Following the above analysis, when the neural network accelerator of an electronic device monopolizes the available bandwidth resources of the electronic device, or the number of second hardware devices participating in bandwidth allocation is small, there is no need to perform bandwidth testing on the neural network accelerator subsequently, and the candidate execution file corresponding to the largest bandwidth range can be directly selected from multiple candidate execution files corresponding to the target computing task to be executed as the current target execution file to be executed.

[0143] When there are a large number of second hardware devices participating in bandwidth allocation in the electronic device, the available bandwidth resources of the neural network accelerator can be compared with the bandwidth ranges of multiple candidate execution files corresponding to the target computing tasks to be executed, and the candidate execution files that match the available bandwidth resources of the neural network accelerator are determined, such as the candidate execution files corresponding to the bandwidth range where the available bandwidth resources of the neural network accelerator are located, and determined as the target execution files to be executed currently, but it is not limited to this method of selecting target execution files.

[0144] Step S65, loading the target execution file into the neural network accelerator so that the neural network accelerator executes the target execution file to complete the target computing task.

[0145] According to the method described above, after determining the most suitable target execution file corresponding to the current target computing task to be executed, it is loaded into the neural network accelerator for execution to complete the target computing task, which can ensure that the performance of the neural network accelerator is optimal under the bandwidth resource allocation of each hardware device of the electronic device. This application does not limit the implementation method of the neural network accelerator executing the loaded target execution file to complete the corresponding target computing task.

[0146] In one possible implementation, if the target execution file includes, in addition to the operation code for implementing the target computing task, the weight parameters used in the neural network model to implement the target computing task, the neural network accelerator can read the required one or more weight parameters from the target execution file during the execution of the operation code, process the input data for the target computing task, and complete the target computing task. In another possible implementation, if the target execution file does not include the weight parameters, during the execution of the target execution file by the neural network accelerator, it can read at least one corresponding weight parameter from the weight file used to store the weight parameters of the neural network model according to processing needs, process the input data, and complete the target computing task, etc.

[0147] In some embodiments, in combination with any of the selection rules described above, during the operation of the neural network model, the neural network accelerator can still be tested for bandwidth according to the method described above to determine that the available bandwidth resources of the neural network accelerator change from greater than the bandwidth threshold to less than the bandwidth threshold, and the target execution file of the subsequent computing task can be selected according to the corresponding selection rules in the above steps. Optionally, the present application can monitor the change in the available bandwidth resources allocated to the neural network accelerator, and dynamically adjust the available bandwidth resources of the neural network accelerator accordingly, so as to continue to perform bandwidth testing on the neural network accelerator according to the adjusted available bandwidth resources before the subsequent computing tasks to be executed are executed, and the target execution file corresponding to the subsequent computing tasks to be executed is selected according to the available bandwidth resources allocated to the neural network accelerator, so as to load it into the neural network accelerator for execution and complete the subsequent computing tasks. The implementation process can refer to the description of the corresponding part of the above embodiment, and this embodiment will not be repeated.

[0148] To sum up, in the neural network model computing task processing method proposed in the present application, the bandwidth resource allocation situation of each hardware device in the current electronic device can be dynamically perceived, and the competition situation of multiple hardware devices (i.e., at least one neural network accelerator and each second hardware device) for the available bandwidth resources of the electronic device can be determined. Based on this, the selection rules of the target execution files corresponding to the computing tasks to be executed are adjusted in time to ensure that the actually selected target execution files are loaded into the neural network accelerator, so as to obtain better performance, and even achieve the optimal performance under the current circumstances, effectively alleviating the adverse effects on the performance of the neural network accelerator when a large number of hardware devices compete for bandwidth resources, thereby reliably meeting the processing requirements of tasks in various application scenarios.

[0149] In combination with the neural network model computing task processing method described in the above embodiment, the relevant description of the multiple candidate execution files corresponding to each computing task itself, the neural network model compilation method proposed in the embodiment of the present application will be introduced in detail in combination with the service.

[0150] Reference Figure 7 , is a flowchart of the neural network model compilation method proposed in the first embodiment of the present application. The method can be applied to electronic devices, such as electronic devices used to implement the above-mentioned neural network model calculation task processing method, or another electronic device different from the electronic device used to implement the above-mentioned neural network model calculation task processing method, etc., depending on the situation. Figure 7 As shown, the neural network model compilation method proposed in this embodiment may include but is not limited to the following steps:

[0151] Step S71, obtaining multiple bandwidth ranges that can be obtained by a neural network accelerator in a working state for executing a computing task of a neural network model;

[0152] As can be seen from the above analysis, during the process of an electronic device running a neural network model, the available bandwidth resources allocated to the neural network accelerator may change, and other hardware devices may also participate in bandwidth allocation, such as increasing or decreasing a second hardware device that competes with the neural network accelerator for bandwidth resources, causing the bandwidth resource allocation of at least one hardware device in the electronic device to change. If fixed bandwidth resources are directly used to compile the original execution files of each computing task of the neural network model, and obtain an execution file corresponding to each computing task itself, so that when the computing task is to be executed, the corresponding execution file is directly loaded into the neural network accelerator for execution, it is easy to reduce the performance of the neural network accelerator due to the inconsistency (i.e., mismatch) between the fixed bandwidth resources used during translation and the available bandwidth resources when the computing task is actually completed.

[0153] In order to improve the above-mentioned problems and ensure the processing efficiency of computing tasks of the neural network model, it is proposed that in the compilation stage, for each computing task of the neural network model, multiple candidate execution files suitable for different bandwidth ranges are compiled and generated. For these multiple different bandwidth ranges, they can be multiple bandwidth ranges that can be obtained by the neural network accelerator in the working state, that is, within the available bandwidth resource range configured by the electronic device, multiple ranges of available bandwidth resources that can be allocated to the neural network accelerator for use. The present application does not limit the method for obtaining these multiple bandwidth ranges, including but not limited to the several acquisition methods listed below.

[0154] In one possible implementation, the present application may be pre-configured for the above-mentioned multiple different bandwidth ranges, and the values ​​of each bandwidth range may be determined based on the design data of each hardware device involved in bandwidth resource allocation in the electronic device, or based on the empirical value determined by the historical operation data of the neural network model. In another possible implementation, the present application may also compile the execution file of each computing task according to the preset bandwidth range and load it into the neural network accelerator for execution, obtain the actual bandwidth resources used to implement the corresponding computing task, use it as a reference bandwidth resource, and re-determine the multiple bandwidth ranges corresponding to each computing task itself.

[0155] Optionally, the present application may also determine the bandwidth ranges that may be allocated to each neural network accelerator in the electronic device under different load states by considering the bandwidth ranges that may be allocated to each hardware device participating in bandwidth resource allocation in the electronic device when the electronic device is running under different load states (i.e., low load state, balanced / intermediate load state, high load state, and ultra-high load state, etc.). In this way, before the electronic device performs the neural network model calculation task, the target load state in which the electronic device is currently located can be obtained, and the bandwidth ranges that may be allocated to each neural network accelerator corresponding to the target load state can be directly determined as multiple bandwidth ranges that can be obtained by the neural network accelerator that implements the neural network model calculation task. Generally, if the electronic device is currently running in a high load state, the bandwidth resources that can be allocated to the neural network accelerator are relatively small, and the values ​​of the multiple bandwidth ranges configured are relatively low; conversely, if the electronic device is currently running in a low load state, the bandwidth resources that can be allocated to the neural network accelerator are relatively large, and the values ​​of the multiple bandwidth ranges configured are relatively high.

[0156] In addition, the present application can also obtain the maximum bandwidth resource and the minimum bandwidth resource of the electronic device, and then directly select a number of bandwidth resources from the numerical bandwidth resources (i.e., the intermediate bandwidth resources) between the maximum bandwidth resource and the minimum bandwidth resource to form multiple bandwidth ranges, so as to compile the original execution file of each computing task and obtain multiple candidate execution files for implementing the computing task. The present application does not restrict the value-taking method and quantity of each intermediate bandwidth resource, such as random or fixed bandwidth interval or other value-taking rules.

[0157] In some embodiments, if the computing tasks of the neural network model include computing tasks corresponding to each neural network layer of the neural network model, in the implementation process of step S71, multiple bandwidth ranges that may be obtained by the corresponding neural network accelerator under the corresponding computing tasks can be obtained based on the computing tasks corresponding to each neural network layer, so that the multiple bandwidth ranges corresponding to different computing tasks can be different, so as to avoid configuring too small a bandwidth range for computing tasks with large computing amount, or configuring too large a bandwidth range for computing tasks with small computing amount, resulting in the actual operation process. The available bandwidth resources within the bandwidth range will not be used, resulting in resource waste. Optionally, in the process of obtaining multiple bandwidth ranges, it is also possible to combine the task requirements of the current application scenario implemented by the neural network model, such as low latency, high throughput, high precision or balance requirements, etc. Since the various computing tasks of the neural network model have different degrees of influence on the realization of the task requirements, for computing tasks with higher influence, these requirements need to be considered more in the process of obtaining their bandwidth ranges, so as to ensure that the corresponding candidate execution files compiled according to the multiple bandwidth ranges can meet the corresponding requirements.

[0158] Therefore, the present application can determine the relevance of each computing task of the neural network model to the task requirements of the neural network model in the current application scenario. For example, the greater the impact of the computing task on the task requirements, the higher the corresponding relevance; conversely, the lower the corresponding relevance. The present application does not limit the method for determining the relevance. For example, it can be determined based on experience or experiments, or the correspondence between various computing tasks and task requirements of different dimensions can be obtained based on historical data analysis. After determining the various computing tasks and task requirements of the neural network model, the correspondence can be queried to determine the relevance of each computing task to the task requirements in the current application scenario.

[0159] In some other embodiments, for the computing tasks corresponding to each neural network layer of the neural network model, the multiple bandwidth ranges based on which the multiple candidate execution files for different computing tasks are compiled and implemented can be the same, and there is no need to spend a long time to determine the different multiple bandwidth ranges corresponding to each computing task. At this time, the method for obtaining the multiple bandwidth ranges based on which the multiple candidate execution files for compiling and implementing the computing tasks are compiled can refer to but is not limited to the method described in the above embodiment. In the process of obtaining the above multiple bandwidth ranges, it is necessary not only to consider the demand for bandwidth resources for the implementation of the computing tasks, but also to consider the influence of other aspects of the electronic device, such as because other hardware devices running on the electronic device will compete with the neural network accelerator that implements the computing task for the total bandwidth resources of the electronic device. If the number or type of other hardware devices competing for bandwidth resources changes, it will affect the bandwidth resources that the neural network accelerator can obtain. At this time, in the process of compiling the candidate execution files of the computing tasks, it is necessary to adaptively reduce the values ​​of the various bandwidth ranges based on, so that the bandwidth range based on the compilation process is basically consistent with the bandwidth range actually used in the process of the neural network accelerator implementing the computing task to execute the corresponding candidate execution file, while ensuring the efficiency and reliability of computing task processing, avoiding bandwidth resource waste.

[0160] In addition, when the electronic device is configured with multiple neural network accelerators and different computing tasks of the neural network model are distributed to different neural network accelerators for implementation, in the above-mentioned process of obtaining multiple bandwidth ranges corresponding to the different candidate execution files for compiling each computing task, it can also be implemented in combination with the configuration information of the neural network accelerator distributed to the computing task (such as the hardware resources it has, etc.). The implementation process is not described in detail in this application.

[0161] Step S72, according to the multiple bandwidth ranges, respectively configuring compilation parameters matching the multiple bandwidth ranges for each computing task;

[0162] Step S73, compiling the original execution files of the corresponding computing tasks according to different compilation parameters to obtain multiple candidate execution files of the corresponding computing tasks.

[0163] In order to ensure that the compiled candidate execution file is loaded into the neural network accelerator for execution, the bandwidth resources required by the neural network accelerator to execute the candidate execution file are within the expected bandwidth range, so as to avoid wasting bandwidth resources or causing failure of computing tasks due to insufficient available bandwidth resources. The present application proposes to compile multiple different candidate execution files for implementing the computing task according to multiple different bandwidth ranges of each computing task, so that these multiple candidate execution files correspond to a bandwidth range respectively, so that when the computing task is subsequently implemented, a matching candidate execution file can be directly selected as the target execution file based on the bandwidth resource allocation situation, and loaded into the neural network accelerator for execution, instead of executing a fixed execution file to complete the computing task, thereby ensuring that the electronic device can obtain better performance of the neural network accelerator under different bandwidth resource allocation situations.

[0164] Based on this, for each computing task of the neural network model, in the process of converting the original execution file corresponding to the computing task into multiple candidate execution files corresponding to each bandwidth range according to multiple different bandwidth ranges, the compilation parameters matching each bandwidth range can be configured for the computing task according to the multiple different bandwidth ranges. The compilation parameters can be used to optimize the performance of the neural network model part (such as at least one neural network layer or its local combination, etc.) corresponding to the computing task to adapt to the neural network accelerator for distribution, and can include the hardware characteristics of the neural network accelerator to which the computing task is distributed (optimized code for characteristics such as memory bandwidth, number of computing units, supported precision, etc.), optimization parameters for the corresponding neural network model part (such as one or more of quantization precision, operator fusion, parallel computing, and memory optimization), calculation graph optimization parameters (such as simplified calculation graph, automatic tuning, etc.), optimization parameters for the operation code that implements the computing task, performance optimization parameters (such as selecting an appropriate batch size to reuse the parallel computing capability of the neural network accelerator, setting multi-threaded / multi-stream execution to improve operating efficiency, and optimizing cache usage to reduce memory access latency), and parameters of a specific compiler (such as ‌ (Open Neural Network Compiler, an open source, modular, reusable compiler algorithm and toolchain library) model, TVM_T compiler that supports neural network training, nncase neural network compiler that supports dividing neural network models into subgraphs, etc.) One or more of these. This application does not limit the content of the compilation parameters.

[0165] It can be seen that for a smaller bandwidth range corresponding to a computing task, the corresponding compilation parameters assigned to the computing task can optimize the compilation of the original execution file of the computing task to a greater extent, so as to reduce the complexity of the original processing logic of the computing task to a greater extent, so that the complexity of the processing logic represented by the compiled corresponding candidate execution file is smaller. Conversely, for a larger bandwidth range corresponding to a computing task, the corresponding compilation parameters assigned to the computing task can optimize the compilation of the original execution file of the computing task to a greater extent, so that the complexity of the processing logic represented by the compiled corresponding candidate execution file is greater. Among them, the more complex the processing logic of the candidate execution file using the same computing task, the larger the bandwidth resources required, which may affect the accuracy of the computing task, but reasonable compilation optimization processing can also ensure that the computing task accuracy that can be achieved by the execution files before and after compilation is basically the same. This application does not elaborate on the compilation implementation method of the original execution file.

[0166] In the embodiment of the present application, in combination with the above analysis, the electronic device can run one or more compilation tools to compile the original execution file of each computing task according to different compilation parameters for the computing task, and obtain different candidate execution files corresponding to the computing task, so that the neural network accelerator can improve the execution efficiency and adaptability of the computing task within the corresponding bandwidth. The implementation process of step S73 can be determined in combination with the working principle of the compilation tool used, and this application will not elaborate on it.

[0167] It can be seen that, relative to configuring respective compilation parameters for each computing task of the neural network model according to a fixed bandwidth range, for each computing task, compiling its original execution file according to the fixed compilation parameters to obtain a compilation method for an execution file for implementing the computing task, the embodiment of the present application proposes first obtaining multiple bandwidth ranges that can be obtained by the neural network accelerator used to execute the computing task in a working state, and configuring multiple different compilation parameters respectively configured for each computing task according to the multiple bandwidth ranges, thereby compiling the original execution file of the computing task respectively according to the multiple different compilation parameters, and obtaining multiple candidate execution files corresponding to each of the multiple different bandwidth ranges for implementing the one computing task, so that the bandwidth resource allocation of each hardware device of the electronic device matched by different candidate execution files is different.

[0168] In this way, when the bandwidth resource allocation situation of each hardware device of the electronic device changes dynamically, the present application can directly select a candidate execution file that matches the changed bandwidth resource allocation situation (such as the bandwidth perception information of the electronic device currently obtained) from multiple pre-compiled candidate execution files that are used to implement this computing task before implementing the current computing task to be executed, as the target execution file, and then load it into the neural network accelerator for execution to complete the computing task. Compared with the processing method in which the neural network accelerator can only execute fixed execution files to implement the computing task under various bandwidth resource allocation situations, the present application can obtain better performance of the neural network accelerator, and even achieve optimal performance.

[0169] Therefore, the neural network model compilation method proposed in this application lays a foundation for coping with the dynamic changes in the bandwidth resource allocation of various hardware devices during the operation of the neural network model and obtaining better performance of the neural network accelerator, improves the compilation flexibility and adaptability of the execution files corresponding to the computing tasks of the neural network model, and helps to improve the operating efficiency and performance of the neural network model.

[0170] In combination with the above description of the contents of the candidate execution file, in some embodiments, if the candidate execution file contains the operation code for implementing the corresponding computing task, but does not contain the weight parameters of the neural network model required to implement the computing task, during the compilation of the neural network model, the weight files required to implement each computing task in the neural network model can also be obtained. The weight file may include all the weight parameters required to implement the corresponding computing task. This application does not limit the content of the weight file and the method of obtaining it.

[0171] Afterwards, multiple candidate execution files and weight files corresponding to the same computing task can be associated and stored, so that when the neural network accelerator executes any candidate execution file of the computing task, the weight parameters required for the computing task can be retrieved from the weight file stored in association with the candidate execution file and loaded into the neural network accelerator to complete the computing task. Regarding the implementation process of the neural network accelerator completing the computing task, the corresponding part of the description in the above embodiment of the neural network model computing task processing method can be referred to, and this embodiment will not be repeated here.

[0172] In a possible implementation, in the process of obtaining the above-mentioned weight file, since multiple candidate execution files used to complete the same computing task are compiled from the original execution file of the computing task, although the processing logics represented by each are different, the weight parameters required to implement each computing task often have some repeated weight parameters. In order to save storage resources, it is not necessary to repeatedly store the same weight parameters. Therefore, the present application can remove the weight parameters required during the execution of multiple candidate execution files corresponding to the same computing task, and store all the weight parameters of the computing task implemented by the corresponding different processing logics as a weight file. In this way, during the operation of the neural network model, according to the method described above, after selecting the target execution file of the target computing task to be executed, the target execution file can be loaded into the neural grid accelerator first. When the neural network accelerator executes the operation code in the target execution file, one or more required weight parameters can be retrieved from the weight file stored in the memory of the electronic device, and the input data for implementing the computing task (which can be the output data of the computing task executed last time, or other external data, etc., the present application does not limit the content of the input data) is read and processed to complete the computing task.

[0173] In one possible implementation, the repeated weight parameters required during the execution of each candidate execution file of the same computing task can be determined and stored as a weight file, and the non-repeated weight parameters required between the candidate execution files are stored in the corresponding candidate execution files, so that one of the candidate execution files is selected as the target execution file and loaded into the neural network accelerator for execution. The neural network accelerator can retrieve the weight parameters required during the execution process from the weight file, and process them in combination with the input data for implementing the computing task to complete the computing task.

[0174] Of course, for each candidate execution file of the same computing task, it can also contain the operation code and weight parameters for implementing the computing task at the same time, so that the selected target execution file can be loaded into the neural network accelerator according to the method described above, and the neural network accelerator can directly run the operation code contained in the target execution file, and process the corresponding weight parameters and input data to complete the computing task.

[0175] Reference Figure 8 , is a flowchart of the neural network model compilation method proposed in Example 2 of the present application, such as Figure 8 As shown, the method may include:

[0176] Step S81, determining the original bandwidth range allocated to the neural network accelerator for executing the computing task of the neural network model according to the bandwidth configuration information of the electronic device;

[0177] Step S82, segmenting the original bandwidth range to obtain multiple bandwidth ranges that the neural network accelerator can obtain when in a working state;

[0178] In the actual application of the present application, multiple bandwidth ranges that can be obtained by the neural network accelerator can be obtained according to the bandwidth resource allocation that may occur during the operation of the neural network model by the electronic device, such as directly segmenting the original bandwidth range that can be allocated during the operation of the neural network accelerator of the electronic device, and taking each bandwidth as a bandwidth range to perform subsequent steps. Exemplarily, since the bandwidth is the amount of data passing through a communication port (such as a port for external communication of an electronic device, etc.) per unit time, if the original bandwidth range is 0~64GB / s (gigabytes per second), that is, the maximum available bandwidth resource is 64GB / s, it can be divided into 32 segments, and the multiple bandwidth ranges obtained can be: 0~1GB / s, 2GB / s~3 GB / s, ..., 63GB / s~64GB / s, etc. 32 bandwidth ranges, but are not limited to 32 segments.

[0179] It should be noted that in the implementation process of step S82, the uniform segmentation method described in the above example can be used to make the bandwidth difference between the minimum bandwidth resource and the maximum bandwidth resource of each bandwidth range the same, or the non-uniform segmentation method can be used to make the bandwidth difference between the minimum bandwidth resource and the maximum bandwidth resource in different bandwidth ranges different. This application does not limit the segmentation method of the original bandwidth range.

[0180] In some other embodiments proposed in the present application, the bandwidth variation range during the historical operation of the neural network accelerator can also be obtained, so as to segment the bandwidth variation range according to the variation range and variation trend of the historical bandwidth resources, and determine the multiple bandwidth ranges that the neural network accelerator can obtain to implement the computing task of the neural network model. The segmentation method used in the implementation process can refer to the description method of the corresponding part of the above embodiment.

[0181] Step S83, according to the multiple bandwidth ranges, respectively configuring compilation parameters matching the multiple bandwidth ranges for each computing task;

[0182] Step S84, compiling the original execution files of the corresponding computing tasks according to different compilation parameters to obtain multiple candidate execution files of the corresponding computing tasks.

[0183] Regarding the implementation method of step S83 and step S84, reference may be made to the description of the corresponding parts of the above embodiment, which will not be elaborated in this embodiment.

[0184] In combination with the above analysis, in some embodiments, Fig. 9As shown, n candidate execution files corresponding to the same computing task are compiled according to the above method, which can be represented as: bw_kernel_1 file, bw_kernel_2 file, ..., bw_kernel_n file, n is the number of bandwidth ranges corresponding to the computing task, such as 32, etc., and this application does not limit its value. These n candidate execution files can be packaged into one file, such as a fat.bin file, which is a mirror file in the file system format of the neural network model. At this time, a bandwidth test instruction for implementing the bandwidth test of the neural network accelerator can also be constructed, such as a bandwidth test code file, which can be represented as a select_kernel file, and it is packaged together with the above n bw_kernel files into a unified fat.bin file.

[0185] In this way, the fat.bin files corresponding to each computing task of the neural network model are obtained, that is, after the compilation and optimization of the neural network model is completed, such as Fig. 9 As shown, during the operation of the neural network model, as analyzed above, the fat.bin file of the target computing task can be run before the target computing task to be executed is executed, and the select_kernel file can be loaded into the neural network accelerator to run, and the bandwidth test of the neural network accelerator can be completed, so as to obtain the corresponding bandwidth test data and obtain the current bandwidth perception information for the neural network accelerator. Based on this, the target bw_kernel file that matches the bandwidth perception information is selected from the n bw_kernel files corresponding to the target computing task and loaded into the neural network accelerator for execution to complete the target computing task. Regarding the implementation process of each computing task of the neural network model, the corresponding part of the description of the embodiment of the neural network model computing task processing method can be referred to, and this embodiment will not be repeated here.

[0186] In combination with the above-introduced neural network model calculation task processing method provided by the embodiment of the present application, the device for executing the above-mentioned neural network model calculation task processing method will be introduced below.

[0187] Reference Fig.10 , is a schematic diagram of the structure of a neural network model computing task processing device provided in an embodiment of the present application. Fig.10 As shown, the neural network model calculation task processing device may include:

[0188] The bandwidth perception information acquisition module 101 is used to acquire bandwidth perception information of the electronic device; the bandwidth perception information represents the bandwidth resource allocation of at least one hardware device of the electronic device, and the at least one hardware device includes a neural network accelerator to implement the computing task of the neural network model;

[0189] The target execution file selection module 102 is used to select, from a plurality of candidate execution files corresponding to the target computing task to be executed, a candidate execution file matching the bandwidth perception information as the target execution file to be executed currently; the plurality of candidate execution files respectively correspond to different bandwidth ranges and are used to complete the same computing task;

[0190] The target execution file loading module 103 is used to load the target execution file into the neural network accelerator so that the neural network accelerator executes the target execution file to complete the target computing task.

[0191] Optionally, the multiple candidate execution files corresponding to the above-mentioned one computing task are compiled based on the original execution file using compilation parameters that match different bandwidth ranges.

[0192] In some embodiments, the bandwidth awareness information acquisition module 101 may include:

[0193] A bandwidth test instruction sending unit is used to send a bandwidth test instruction to the neural network accelerator to be executed by the target computing task, so that the neural network accelerator executes the bandwidth test instruction at least once to obtain bandwidth test data; the bandwidth test instruction includes at least part of the computing instructions required by the target computing task;

[0194] A bandwidth test data receiving unit, used to receive the bandwidth test data fed back by the neural network accelerator;

[0195] The bandwidth perception information obtaining unit is used to obtain bandwidth perception information for the neural network accelerator based on the bandwidth test data.

[0196] In a possible implementation, the bandwidth test instruction sending unit may include any of the following sending units:

[0197] A first sending unit is used to send a bandwidth test instruction to a neural network accelerator to execute the target computing task before the target computing task corresponding to the neural network layer to be executed is executed if the computing task of the neural network model includes different computing tasks corresponding to each neural network layer;

[0198] The second sending unit is used to send a bandwidth test instruction to the neural network accelerator to execute the target computing task according to a preset bandwidth test frequency.

[0199] Optionally, the bandwidth awareness information acquisition module 101 may further include:

[0200] An available bandwidth resource change amount determining unit, configured to determine the available bandwidth resource change amount allocated to the neural network accelerator based on a plurality of bandwidth test data continuously fed back by the neural network accelerator;

[0201] A bandwidth test frequency adjustment unit is used to adjust the bandwidth test frequency of the neural network accelerator according to the change in the available bandwidth resources, so as to execute the step of sending a bandwidth test instruction to the neural network accelerator to execute the target computing task according to the adjusted bandwidth test frequency.

[0202] In some embodiments, the target executable file selection module 102 may include any of the following selection units:

[0203] A first selection unit is used to determine, based on the number of hardware devices participating in bandwidth allocation contained in the bandwidth perception information or the available bandwidth resources allocated to the neural network accelerator, that the available bandwidth resources of the neural network accelerator are greater than a bandwidth threshold, and select a candidate execution file corresponding to a maximum bandwidth range from a plurality of candidate execution files corresponding to the target computing task to be executed as the target execution file to be executed currently;

[0204] a second selection unit, configured to determine, based on the number of hardware devices participating in bandwidth allocation contained in the bandwidth perception information or the available bandwidth resources allocated to the neural network accelerator, that the available bandwidth resources of the neural network accelerator are less than a bandwidth threshold, and select, from a plurality of candidate execution files corresponding to the target computing task to be executed, a candidate execution file that matches the available bandwidth resources of the neural network accelerator as the target execution file to be executed currently;

[0205] The bandwidth threshold is determined based on available bandwidth resources of the electronic device.

[0206] In combination with the neural network model compilation method provided by the embodiment of the present application introduced above, the device for executing the above-mentioned neural network model compilation method will be introduced below.

[0207] Reference Fig.11 , is a schematic diagram of the structure of a neural network model compilation device provided in an embodiment of the present application. Fig.11 As shown, the neural network model compilation device may include:

[0208] A bandwidth range acquisition module 111 is used to acquire multiple bandwidth ranges that can be acquired by a neural network accelerator that is used to perform a computing task of the neural network model when it is in a working state;

[0209] A compilation parameter configuration module 112, configured to configure compilation parameters matching each of the multiple bandwidth ranges for each of the computing tasks according to the multiple bandwidth ranges;

[0210] The compiling module 113 is used to compile the original execution files corresponding to the computing tasks according to different compiling parameters to obtain multiple candidate execution files corresponding to the computing tasks.

[0211] Optionally, the neural network model compiling device may further include:

[0212] A weight file acquisition module is used to acquire the weight files required to implement each of the computing tasks in the neural network model;

[0213] A storage module is used to associate and store multiple candidate execution files corresponding to the same computing task with the weight file, so that when the neural network accelerator executes any of the candidate execution files of the computing task, the weight parameters required for the computing task are retrieved from the associated stored weight file and loaded into the neural network accelerator to complete the computing task.

[0214] Optionally, the bandwidth range acquisition module 111 may include any of the following units:

[0215] An acquisition unit, used to acquire a bandwidth variation range during the historical operation of the neural network accelerator, segment the bandwidth variation range according to the variation range and variation trend of the historical bandwidth resources, and determine a plurality of bandwidth ranges that can be acquired by the neural network accelerator to implement the computing task of the neural network model;

[0216] A segmentation unit is used to determine the original bandwidth range allocated to the neural network accelerator used to execute the computing task of the neural network model based on the bandwidth configuration information of the electronic device, and to segment the original bandwidth range to obtain multiple bandwidth ranges that the neural network accelerator can obtain when it is in a working state.

[0217] Also provided in an embodiment of the present application is a computer program product including computer-readable instructions. When the computer-readable instructions are executed on an electronic device, the electronic device implements any one of the neural network model computing task processing and / or neural network model compilation methods provided in the embodiments of the present application.

[0218] A computer-readable storage medium is also provided in an embodiment of the present application. The storage medium carries one or more computer programs. When one or more computer programs are executed by an electronic device, the electronic device can implement any one of the neural network model computing task processing and / or neural network model compilation methods provided in the embodiment of the present application.

[0219] It should also be noted that the device embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed over multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. In addition, in the drawings of the device embodiments provided by the present application, the connection relationship between the modules indicates that there is a communication connection between them, which may be specifically implemented as one or more communication buses or signal lines.

[0220] Through the description of the above implementation mode, in the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website site, a computer, a training device or a data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website site, computer, training device or data center. The computer-readable storage medium can be any available medium that a computer can store or a data storage device such as a training device, a data center, etc. that contains one or more available media integrated. The available medium may be a magnetic medium (eg, a floppy disk, a hard disk, a magnetic tape), an optical medium (eg, a DVD), or a semiconductor medium (eg, a solid state disk (SSD)).

[0221] In addition, the various embodiments in this specification are described in a progressive or parallel manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other. For the devices, electronic devices, products and media disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the method part description.

Claims

1. A method for processing a neural network model calculation task, the method comprising: Obtaining bandwidth perception information of electronic devices; The bandwidth perception information represents bandwidth resource allocation of at least one hardware device of the electronic device, wherein the at least one hardware device includes a neural network accelerator for implementing a computing task of the neural network model; According to the bandwidth perception information, from a plurality of candidate execution files corresponding to the target computing task to be executed, a candidate execution file matching the bandwidth perception information is selected as the target execution file to be executed currently; the plurality of candidate execution files respectively correspond to different bandwidth ranges and are used to complete the same computing task; The target execution file is loaded into the neural network accelerator so that the neural network accelerator executes the target execution file to complete the target computing task.

2. According to the method of claim 1, the acquiring of bandwidth perception information of the electronic device comprises: Sending a bandwidth test instruction to a neural network accelerator to be executed with respect to the target computing task, so that the neural network accelerator executes the bandwidth test instruction at least once to obtain bandwidth test data; the bandwidth test instruction includes at least part of the computing instructions required by the target computing task; Receiving the bandwidth test data fed back by the neural network accelerator; Based on the bandwidth test data, bandwidth perception information for the neural network accelerator is obtained.

3. According to the method of claim 2, the sending of a bandwidth test instruction to the neural network accelerator to execute the target computing task comprises any one of the following: If the computing tasks of the neural network model include different computing tasks corresponding to each neural network layer, before the target computing task corresponding to the neural network layer to be executed is executed, a bandwidth test instruction is sent to the neural network accelerator to execute the target computing task; According to the preset bandwidth test frequency, a bandwidth test instruction is sent to the neural network accelerator to execute the target computing task.

4. The method according to claim 2 or 3, wherein the acquiring bandwidth awareness information of the electronic device further comprises: Determine a change in available bandwidth resources allocated to the neural network accelerator based on a plurality of bandwidth test data continuously fed back by the neural network accelerator; According to the change in the available bandwidth resources, the bandwidth test frequency of the neural network accelerator is adjusted, so that according to the adjusted bandwidth test frequency, the step of sending a bandwidth test instruction to the neural network accelerator to execute the target computing task is executed.

5. The method according to claim 1, wherein: The multiple candidate execution files are obtained by compiling the original execution file using compilation parameters that match different bandwidth ranges.

6. The method according to claim 1, wherein, based on the bandwidth awareness information, selecting a candidate execution file that matches the bandwidth awareness information from multiple candidate execution files corresponding to the target computing task to be executed as the target execution file to be executed currently comprises any of the following: According to the number of hardware devices participating in bandwidth allocation contained in the bandwidth perception information, or the available bandwidth resources allocated to the neural network accelerator, it is determined that the available bandwidth resources of the neural network accelerator are greater than the bandwidth threshold, and a candidate execution file corresponding to the maximum bandwidth range is selected from multiple candidate execution files corresponding to the target computing task to be executed as the target execution file to be executed currently; According to the number of hardware devices participating in bandwidth allocation contained in the bandwidth perception information, or the available bandwidth resources allocated to the neural network accelerator, it is determined that the available bandwidth resources of the neural network accelerator are less than the bandwidth threshold, and a candidate execution file matching the available bandwidth resources of the neural network accelerator is selected from multiple candidate execution files corresponding to the target computing task to be executed as the target execution file to be executed currently; in, The bandwidth threshold is determined based on available bandwidth resources of the electronic device.

7. A neural network model compilation method, the neural network model compilation method comprising: Acquire multiple bandwidth ranges that can be obtained by a neural network accelerator in a working state for executing a computing task of the neural network model; According to the multiple bandwidth ranges, respectively configuring compilation parameters matching the multiple bandwidth ranges for each of the computing tasks; According to the different compilation parameters, the original execution files corresponding to the computing tasks are compiled respectively to obtain a plurality of candidate execution files corresponding to the computing tasks.

8. The method according to claim 7, wherein the neural network model compilation method further comprises: Obtaining weight files required to implement each of the computing tasks in the neural network model; Multiple candidate execution files corresponding to the same computing task are associated with the weight file and stored, so that when the neural network accelerator executes any of the candidate execution files of the computing task, the weight parameters required for the computing task are retrieved from the associated stored weight file and loaded into the neural network accelerator to complete the computing task.

9. According to the method of claim 7, the step of acquiring multiple bandwidth ranges that can be acquired by the neural network accelerator in a working state for executing the computing task of the neural network model comprises any one of the following: Obtaining a bandwidth variation range during the historical operation of the neural network accelerator, segmenting the bandwidth variation range according to the variation range and variation trend of the historical bandwidth resources, and determining a plurality of bandwidth ranges that the neural network accelerator can obtain to implement the computing task of the neural network model; According to the bandwidth configuration information of the electronic device, the original bandwidth range allocated to the neural network accelerator used to execute the computing task of the neural network model is determined, and the original bandwidth range is segmented to obtain multiple bandwidth ranges that the neural network accelerator can obtain when it is in a working state.

10. A neural network model calculation task processing device, the neural network model calculation task processing device comprising: A bandwidth perception information acquisition module, used to acquire bandwidth perception information of an electronic device; The bandwidth perception information represents bandwidth resource allocation of at least one hardware device of the electronic device, wherein the at least one hardware device includes a neural network accelerator for implementing a computing task of the neural network model; a target execution file selection module, configured to select, from a plurality of candidate execution files corresponding to the target computing task to be executed, a candidate execution file matching the bandwidth perception information as the target execution file to be executed currently; the plurality of candidate execution files respectively correspond to different bandwidth ranges and are used to complete the same computing task; A target execution file loading module is used to load the target execution file into the neural network accelerator so that the neural network accelerator executes the target execution file to complete the target computing task.

11. A neural network model compilation device, the neural network model compilation device comprising: A bandwidth range acquisition module, used to acquire multiple bandwidth ranges that can be acquired by a neural network accelerator in a working state for executing a computing task of the neural network model; A compilation parameter configuration module, configured to configure compilation parameters matching each of the multiple bandwidth ranges for each of the computing tasks according to the multiple bandwidth ranges; The compiling module is used to compile the original execution files of the corresponding computing tasks according to different compiling parameters to obtain multiple candidate execution files of the corresponding computing tasks.