Multimodal neural network timing method, timing processor and storage medium
By using a multi-mode neural network timing processor to partition processing units into clusters and design a shared cache, the problems of high cost and insufficient flexibility in timing processing under different application scenarios are solved, achieving low-cost multi-neural network timing processing and improving the system's versatility and efficiency.
Patent Information
- Application Number
- CN202210041259.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-14
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2042-01-14
AI Technical Summary
In the fields of radiation detection and medical imaging, existing technologies struggle to achieve timed processing of multiple neural networks across different application scenarios, resulting in high costs and insufficient system flexibility and versatility.
A multi-mode neural network timing processor is provided, which realizes timing processing in different application scenarios by clustering the processing unit array and sharing the feature map cache, combined with the global cache unit and input and output logic units.
Under low cost conditions, multi-neural network timing processing in different application scenarios is realized, which improves the flexibility and versatility of the system structure and adapts to the needs of complex scenarios such as high count rate, single event effect and uncertainty estimation.
Smart Images

Figure CN114386581B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of signal processing, in particular to a multi-mode neural network timing method, a timing processor and a storage medium. BACKGROUND
[0002] In radiation detection and new medical imaging equipment, there is a wide demand for time resolution of the detector. For example, in positron emission tomography (PET) based on time-of-flight (TOF) technology, the time difference of paired photons generated by positron-electron pair annihilation reaching the detector module needs to be measured to determine the specific location of annihilation. For another example, in high-energy physics experiments, the time resolution capability of the energy meter helps to identify particles and identify primary / secondary vertex decay processes.
[0003] As a new feature extraction algorithm, the target neural network takes all sampling point signals in a period of time as input and time feature data as output, which can more fully utilize the beneficial information of the sampling points for the target task and has the potential to improve time resolution. Further, the target neural network accelerator designed for specific operations in the target neural network improves the efficiency of forward inference of the target neural network from the hardware level and can be implemented on an application specific integrated circuit (ASIC) or a field programmable gate array (FPGA).
[0004] In the field of radiation detection and medical imaging, different applications often have different technical requirements, and different target neural network timing processors need to be designed for different application scenarios, which is relatively high in cost. SUMMARY
[0005] Therefore, the present application provides a multi-mode neural network timing method, a timing processor and a storage medium, which can realize timing of multiple neural networks in different application scenarios at a low cost.
[0006] The delivery route determination method and device can quickly and accurately determine an optimized delivery route set.
[0007] To solve the above technical problems, the technical solution of the present application is as follows:
[0008] In one embodiment, a multi-mode neural network timing processor is provided, comprising: a processing unit, a global cache unit, an input logic unit and an output logic unit, data transmission is performed through an on-chip interconnection; and the processing unit array is divided into processing unit clusters according to application scenarios, and the processing units in each of the divided processing unit clusters share a feature map cache; wherein the processing unit array is a set of all processing units in the timing processor.
[0009] The processing unit is configured to cache feature map and convolution kernel data and perform related operations.
[0010] The global cache unit is configured to cache convolution kernel data required for timing.
[0011] The input logic unit is configured to input a sampling pulse point signal.
[0012] The output logic unit is configured to output a timing result.
[0013] In another embodiment, a multi-mode neural network timing method is provided, which is applied to a timing processor in a high count rate optimization operation scenario, and the method comprises:
[0014] The global cache unit loads convolution kernel data of each layer of a target neural network load to the convolution kernel cache of the processing unit cluster corresponding to the layer.
[0015] A sampling point pulse signal arrives at the input logic unit.
[0016] If the processing unit cluster corresponding to the first layer of the target neural network load is idle, the sampling point pulse signal is loaded into the feature map cache of the processing unit cluster corresponding to the first layer through the input logic unit, and forward inference operation of the target neural network is performed.
[0017] When the processing unit cluster corresponding to the current layer completes operation of the corresponding layer, if the processing unit cluster corresponding to the next layer is idle, the operation result of the current layer is loaded into the feature map cache of the processing unit cluster corresponding to the next layer.
[0018] When the processing unit cluster corresponding to the last layer completes operation of the corresponding layer, the operation result is transmitted to the output logic unit.
[0019] The output logic unit outputs a timing result corresponding to the sampling point pulse signal.
[0020] In another embodiment, a multi-mode neural network timing method is provided, which is applied to a timing processor in a single event effect optimization operation scenario, and the method comprises:
[0021] The global cache unit simultaneously transmits the same convolution kernel data to the convolution kernel cache of each processing unit combination;
[0022] The sampling point pulse signal reaches the input logic unit;
[0023] If it is determined that all the processing unit combinations are idle, the sampling point pulse signal is simultaneously loaded from the input logic unit to the feature map cache of the first preset number of processing unit combinations, and the forward inference operation of the target neural network is started;
[0024] When all the processing units in the processing unit combination complete the operation, the operation result is transmitted to the output logic unit;
[0025] Based on the operation result of the first preset number of processing unit combinations, the output logic unit determines the timing result corresponding to the sampling point pulse signal according to a preset decision principle.
[0026] In another embodiment, a multi-mode neural network timing method is provided, which is applied to a timing processor in an uncertainty estimation optimization operation scenario, and the method comprises:
[0027] The global cache unit simultaneously transmits the same convolution kernel data to the convolution kernel cache of each processing unit combination;
[0028] The sampling point pulse signal reaches the input logic unit;
[0029] If it is determined that all the processing unit combinations are idle, the sampling point pulse signal is simultaneously loaded from the input logic unit to the feature map cache of the first preset number of processing unit combinations, and the forward inference operation of the target neural network is started;
[0030] When all the processing units in the processing unit combination complete the operation, the operation result is transmitted to the output logic unit;
[0031] Based on the operation result, the output logic unit determines the uncertainty corresponding to the sampling point pulse signal.
[0032] In another embodiment, a computer readable storage medium is provided, which stores a computer program, and the program is executed by a processor to implement the timing processing method of the timing processor in different application scenarios.
[0033] In another embodiment, a computer program product is provided, which comprises a computer program, and the computer program is executed by a processor to implement the timing processing method of the timing processor in different application scenarios.
[0034] From the above technical solutions, it can be seen that in the above embodiments, based on the general timer, the processing unit array is divided into processing unit clusters in different application scenarios, and the processing units in each processing unit cluster share the feature map cache; the timer processing can be performed for various application scenarios under the condition of multiplexing hardware resources. The timer processor can realize the timing of multiple neural networks in different application scenarios under the condition of low cost. The technical solution can greatly improve the flexibility and universality of the design system structure. BRIEF DESCRIPTION OF DRAWINGS
[0035] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0036] Figure 1 The structure schematic diagram of the multi-mode target neural network timing processor in the embodiments of the present application is shown in the figure.
[0037] Figure 2 The structure schematic diagram of the timing processor in the high counting rate optimization operation scenario in the embodiments of the present application is shown in the figure.
[0038] Figure 3 The timing flow schematic diagram in the high counting rate optimization operation scenario in the embodiments of the present application is shown in the figure.
[0039] Figure 4 The structure schematic diagram of the timing processor in the single event effect optimization operation scenario in the embodiments of the present application is shown in the figure.
[0040] Figure 5 The timing flow schematic diagram in the single event effect optimization operation scenario in the embodiments of the present application is shown in the figure.
[0041] Figure 6 The structure schematic diagram of the timing processor in the uncertainty estimation optimization operation scenario in the embodiments of the present application is shown in the figure.
[0042] Figure 7 The timing flow schematic diagram in the uncertainty estimation optimization operation scenario in the embodiments of the present application is shown in the figure. DETAILED DESCRIPTION
[0043] The technical solutions in the embodiments of the present application will be described clearly and completely in the following with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0044] The terms "first", "second", "third", "fourth" and the like in the description and in the claims of the present application, and above-mentioned drawings, if any, are used as identifiers to distinguish between similar objects, and are not necessarily intended to describe a particular sequential or chronological order. It will be understood that the use of such terms is interchangeable under appropriate circumstances such that the embodiment of the application described herein are capable of operating in other sequences than those illustrated or otherwise described herein. Furthermore, the terms "comprise", "comprising", "include", "including" and "has", "having" and any variations thereof are intended to cover a non-exclusive inclusion, such that a process, method, system, product or apparatus that comprises a list of steps or units is not necessarily limited to those steps or units that are expressly listed, but can include other steps or units not expressly listed or inherent to such process, method, product or apparatus.
[0045] The technical solutions of the present application will be described in detail below with specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments.
[0046] In the embodiments of the present application, a multi-mode neural network timing processor is provided, which has a general target neural network timing processor architecture, and can perform timing processing for various application scenarios through division of processing units under the condition of multiplexing hardware resources. The timing processor can realize timing processing in various application scenarios under the condition of low cost. The technical solution can greatly improve the flexibility and universality of the design system structure.
[0047] The structure of the general multi-mode target neural network timer provided in the embodiments of the present application will be described below in conjunction with the drawings.
[0048] Referring to Figure 1 , Figure 1 The structure of the multi-mode target neural network timing processor in the embodiments of the present application is shown in the figure. Figure 1 The timing processor in the figure includes a processing unit, a global cache unit, an input logic unit and an output logic unit, and data transmission is performed through on-chip interconnection;
[0049] The processing unit is used to cache feature map and convolution kernel data, and perform related operations;
[0050] The processing unit includes a multiply-add operation core, a feature map cache and a convolution kernel cache;
[0051] The multiply-add operation core is used to perform multiply-add related operations;
[0052] The feature map cache is used to cache feature maps;
[0053] The convolution kernel cache is used to cache convolution kernel data.
[0054] In order to work normally, the processing unit also includes necessary control logic, the implementation of the control logic in the embodiments of the application is not changed, therefore, it is not described here.
[0055] A global cache unit is configured to cache convolution kernel data required for timing.
[0056] An input logic unit is configured to input a sampling pulse point signal; the sampling pulse signal herein can be an ADC sampling pulse signal.
[0057] An output logic unit is configured to output a timing result.
[0058] Figure 1 For example, data transmission is performed by on-chip interconnection, data transmission between processing units can also be performed by a private interconnection bus.
[0059] Based on the above-mentioned generally-structured timing processor, when applied in different application scenarios, the processing unit array can be divided into processing unit clusters according to the application scenarios, and the processing units in each processing unit cluster after division share a feature map cache; wherein the processing unit array is a collection of all processing units in the timing processor. Figure 1 The left column in FIG. 1 is an example of a processing unit cluster, and the processing units in the processing unit cluster share a feature map cache.
[0060] The following describes the timing process in different application scenarios using the above-mentioned timing processor.
[0061] If the application scenario is a high count rate optimization operation scenario, the processing units in the processing unit array are divided into N processing unit clusters based on the number of layers of the target neural network load, and the data storage requirement and operation amount of each layer; wherein N is an integer greater than 1.
[0062] When the processing unit clusters are divided, the processing units in each processing unit cluster are adjacent in space or topology, and the processing units included in the processing unit cluster after division do not coincide in space or topology, that is, a processing unit is not divided into multiple processing unit clusters.
[0063] The processing unit cluster corresponding to the first layer of the target neural network is connected to the input logic unit;
[0064] The feature map caches of the processing unit clusters of different layers constitute a cascade structure of a pipeline;
[0065] The processing unit cluster corresponding to the last layer of the target neural network is connected to the output logic unit.
[0066] In a specific implementation, taking N equal to 3 as an example, the processing unit array is divided into 3 processing unit clusters, and the processing units in each processing unit cluster share the feature map cache. Referring to Figure 2 , Figure 2 FIG. 1 is a schematic diagram of a timing processor structure in a high count rate optimization operation scenario according to an embodiment of the present application.
[0067] Figure 2 In the embodiment, the processing unit array is divided into 3 processing unit clusters, i.e., processing unit cluster 21, processing unit cluster 22 and processing unit cluster 23.
[0068] The feature map caches of the processing units in each processing unit cluster are shared, i.e., the feature map cache units are shared.
[0069] The processing unit cluster 21 corresponding to the first layer of the target neural network is connected to the input logic unit;
[0070] The feature map caches of the processing unit clusters (processing unit cluster 21, processing unit cluster 22 and processing unit cluster 23) of different layers constitute a cascade structure of the pipeline;
[0071] The processing unit cluster 23 corresponding to the last layer of the target neural network is connected to the output logic unit.
[0072] The process of timing processing in a high count rate optimization operation scenario will be described in detail below with reference to the accompanying drawings.
[0073] Referring to Figure 3 , Figure 3 FIG. 2 is a schematic diagram of a timing process in a high count rate optimization operation scenario according to an embodiment of the present application. The specific steps are as follows:
[0074] In step 301, the global cache unit loads the convolution kernel data of each layer of the target neural network through the convolution kernel cache of the processing unit cluster corresponding to the layer.
[0075] The target neural network here is the neural network used for timing in this application scenario. For example, Figure 3 In the embodiment, the global cache unit loads the convolution kernel data of each layer through the convolution kernel cache of the three processing unit clusters respectively.
[0076] This step is a preparation work before the operation timing. If the convolution kernel data does not need to be updated, this preparation work does not need to be performed every time the sampling point pulse signal is received, but can be performed once.
[0077] In step 302, the sampling point pulse signal reaches the input logic unit.
[0078] At this time, the sampling point pulse signal that needs to be processed is received.
[0079] Step 303, if the processing unit cluster corresponding to the first layer of the target neural network load is idle, load the sample point pulse signal into the feature map buffer of the processing unit cluster corresponding to the first layer through the input logic unit, and perform forward inference operation of the target neural network.
[0080] This step needs to judge whether the processing unit cluster corresponding to the first layer of the target neural network load is idle, that is, whether the processing unit cluster 21 in the formula (1) is idle, if yes, load the sample point pulse signal for processing; otherwise, wait until the processing unit cluster corresponding to the first layer of the target neural network load is idle to load the sample point pulse signal. Figure 2
[0081] Here, only the processing unit cluster corresponding to the first layer needs to be idle, and all processing unit clusters do not need to be idle, so that in the case that the previous sample point pulse signal has not completed operation, the new sample point pulse signal is allowed to enter the pipeline, and the parallel operation of the processing unit cluster in space and time can be realized, and the operation efficiency is improved.
[0082] Step 304, when the processing unit cluster corresponding to the current layer completes the operation of the corresponding layer, if the processing unit cluster corresponding to the next layer is idle, load the operation result of the current layer into the feature map buffer of the processing unit cluster corresponding to the next layer.
[0083] As shown in the formula (2), when the processing unit cluster 21 completes the operation, if the processing unit cluster 22 is idle, load the operation result of the processing unit cluster 21 into the feature map buffer of the processing unit cluster 22; if the processing unit cluster 22 is not idle, first cache the operation result in the feature map buffer of the processing unit cluster 21 until the processing unit cluster 22 is idle, and so on, to realize the pipeline operation. Figure 2 Step 305, when the processing unit cluster corresponding to the last layer completes the operation of the corresponding layer, transmit the operation result to the output logic unit.
[0084] As shown in the formula (3), when the processing unit cluster 23 completes the operation, transmit the operation result to the output logic unit through the on-chip interconnection.
[0085] Figure 2 The operation result here is the processing timing result.
[0086] Step 306, output the timing result corresponding to the sample point pulse signal through the output logic unit.
[0087] Step 306, output the timing result corresponding to the sample point pulse signal through the output logic unit.
[0088] The processing unit array is divided into a plurality of processing unit clusters in the embodiment, and the feature map cache is shared within the processing unit cluster; the input logic, the processing unit cluster and the output logic form a hierarchical pipeline structure through on-chip interconnection, and the data of the convolution kernel cache is preloaded; the input sampling point pulse signal is subjected to the pipeline neural network inference operation in an event-driven manner, and the timing result is transmitted from the last stage of the pipeline; the pipeline enables the system to perform parallel processing in space and time on the input sampling point pulse signals of different events, thereby improving the throughput of the system and enabling the system to adapt to the demand of high counting rate.
[0089] If the application scenario is a single event effect optimization operation scenario, the processing units in the processing unit array are divided into a first preset number of processing unit combinations according to the adjacent relationship in space or topology;
[0090] The processing unit combinations are divided into a plurality of processing unit clusters.
[0091] Or, each processing unit combination is taken as a processing unit cluster.
[0092] In the specific implementation, the processing units can be first divided into a first preset number of processing unit combinations, that is, divided into a plurality of parts, if a certain processing unit combination can be divided again, the combination is divided into a plurality of processing unit clusters, if a certain processing unit combination is not divided into a plurality of processing unit clusters, the combination is taken as a processing unit cluster, that is, the first preset number of processing units can all be divided into processing unit clusters, or none of the processing units can be divided into processing unit clusters, or some of the processing unit combinations can be divided into processing unit clusters.
[0093] Taking the first preset value as 3 and not dividing all the processing unit combinations, each processing unit combination is taken as a processing unit cluster as an example.
[0094] The feature map cache is shared in the same processing unit cluster.
[0095] The number of processing units in each processing unit cluster can be evenly divided, if the total number of processing units cannot be divided by N and the remainder is M, M processing unit clusters are selected from N processing unit clusters to respectively assign one processing unit.
[0096] In the division of the processing unit cluster, the processing units in each processing unit cluster are adjacent in space or topology, and the processing units included in the divided processing unit cluster do not overlap in space or topology, that is, a processing unit is not divided into a plurality of processing unit clusters.
[0097] Each processing unit cluster is connected with an input logic unit and an output logic unit.
[0098] There is no connection between different processing unit clusters.
[0099] In the specific implementation, the processing unit array is divided into three processing unit combinations, without further division into processing unit clusters. Each combination functions as a single processing unit cluster, allowing the processing units within each combination to share the feature map cache. See [link to implementation details]. Figure 4 , Figure 4 This is a schematic diagram of the timer processor structure in the single-event effect optimization operation scenario in the embodiments of this application.
[0100] Figure 4 The processing unit array is divided into three combinations: processing unit combination 41, processing unit combination 42, and processing unit combination 43.
[0101] The feature map cache is shared among the processing units in each processing unit combination, that is, the feature map cache unit is shared.
[0102] Processing unit combination 41, processing unit combination 42 and processing unit combination 43 are respectively connected to the input logic unit and the output logic unit.
[0103] The following section, with reference to the accompanying diagram, details the process of timing processing in a single-event effect optimization scenario.
[0104] See Figure 5 , Figure 5 This is a schematic diagram of the timing process in the single-event effect optimization operation scenario in this application embodiment. The specific steps are as follows:
[0105] Step 501: Simultaneously transmit the same convolution kernel data to the convolution kernel cache of each processing unit through the global cache unit.
[0106] This step is a preparatory step before the operation begins. If the convolution kernel data does not need to be updated, this preparatory step does not need to be performed before each sampling point pulse signal is received; it only needs to be performed once.
[0107] like Figure 4 In this process, the global cache unit loads the same convolution kernel data into the convolution kernel caches of the three processing unit clusters through on-chip interconnects.
[0108] Step 502: The sampling point pulse signal arrives at the input logic unit.
[0109] Step 503: If it is determined that all the processing unit combinations are idle, the sampling point pulse signals are simultaneously loaded from the input logic unit into the feature map cache of the first preset number of processing unit combinations, and the forward inference operation of the target neural network begins.
[0110] If the processing unit combination is not divided into processing unit clusters, all processing unit combinations need to be idle;
[0111] If the processing unit combination is divided into processing unit clusters, and the first processing unit cluster of each processing unit combination is idle, it is determined that all processing unit combinations are idle.
[0112] If one combination is divided into multiple processing unit clusters, that is, there are multiple processing unit clusters in one processing unit combination, it can be processed in a way that optimizes the pipeline in the high-count-rate optimization operation scenario.
[0113] If all the convolution kernel data cannot be cached at one time, and the operation of the current convolution kernel data has been completed, the same convolution kernel data is loaded again through the global cache unit.
[0114] Step 504, when all processing units in the processing unit combination complete the operation, the operation result is transmitted to the output logic unit.
[0115] When each processing unit combination completes the operation, the operation result is transmitted to the output logic unit through the on-chip interconnection.
[0116] Step 505, based on the operation result of the first preset number of processing unit combinations, the output logic unit determines the timing result corresponding to the sampling point pulse signal according to a preset decision principle and outputs it.
[0117] The preset decision principle here is related to the first preset value. For example, if the first preset value is 3, the decision principle can be that the result is valid, that is, the timing processing result, when the operation results of two processing unit combinations are the same.
[0118] If the first preset value is 5, the decision principle can be that the result is valid, that is, the timing processing result, when the operation results of three processing unit combinations are the same.
[0119] In this embodiment, if single event effect contaminates the operation result of one of the multiple processing unit combinations, or several operation results, but the remaining operation results are all correct, the neural network inference process is still valid, so this embodiment can reduce the influence of single event effect to a certain extent. And multiple processing unit combinations are in different positions in space or topology, and the possibility of single event effect occurring at multiple positions and affecting the result in one neural network inference process is extremely small.
[0120] If the application scenario is the uncertainty estimation optimization operation scenario, based on the number of layers and the operation scale of the target neural network load, the processing unit array is divided into a second preset number of processing unit clusters.
[0121] The second preset value can be an integer greater than 4.
[0122] The processing units in each processing unit cluster share a feature map cache.
[0123] In a specific implementation, the processing units can be divided into second preset value processing unit sets.
[0124] The number of processing units in each processing unit cluster can be evenly divided. If the total number of processing units cannot be divided by N, and the remainder is M, M processing unit clusters in N processing unit clusters are selected to respectively assign one processing unit.
[0125] In the division of the processing unit cluster, the processing units in each processing unit cluster are adjacent in space or topology, and the processing units included in the divided processing unit cluster do not overlap in space or topology, that is, a processing unit is not divided into multiple processing unit clusters.
[0126] Each processing unit cluster is connected to an input logic unit and an output logic unit.
[0127] Different processing unit clusters are not associated.
[0128] In actual application, in order to complete the uncertainty estimation, preferably, the implementation scheme of dividing the processing unit array into more than 5 processing unit clusters is preferred, but the number of divided processing unit clusters is not limited.
[0129] In the embodiment of the application, in order to make the image clearer, the processing unit array is divided into two processing unit clusters, and the processing units in each processing unit cluster share a feature map cache. Figure 6 , Figure 6 The figure is a schematic diagram of the timing processor structure in the uncertainty estimation optimization operation scenario in the embodiment of the application.
[0130] Figure 6 In the figure, two processing unit clusters are divided, which are processing unit cluster 61 and processing unit cluster 62. The processing units in each processing unit cluster share a feature map cache.
[0131] The processing unit cluster 61 and the processing unit cluster 62 are respectively connected to an input logic unit and an output logic unit. The processing units in each processing unit cluster have an adjacent relationship in space or topology, and there is no same processing unit in the processing unit cluster 61 and the processing unit cluster 62.
[0132] The process of timing processing in the uncertainty estimation optimization operation scenario is described in detail below with reference to the accompanying drawings.
[0133] Referring to Figure 7 ,Figure 7 A timing flow diagram in an uncertainty estimation optimization operation scenario in an embodiment of the present application is shown. The specific steps are as follows:
[0134] In step 701, the global cache unit transmits different versions of target neural network convolution kernel data to the convolution kernel cache of each processing unit cluster.
[0135] This step is a preparation work before the operation. If the convolution kernel data does not need to be updated, the preparation work before receiving the sampling point pulse signal each time is not needed, and it can be performed once.
[0136] As Figure 6 In the embodiment, the global cache unit loads different versions of target neural network convolution kernel data to the convolution kernel cache of the two processing unit clusters through on-chip interconnection.
[0137] In step 702, the sampling point pulse signal reaches the input logic unit.
[0138] In step 703, if it is determined that the processing unit clusters are all idle, the sampling point pulse signal is loaded into the feature map cache of each processing unit cluster through the input logic unit, and the forward inference operation of the target neural network is started.
[0139] If the processing unit clusters complete the operation, the operation result is transmitted to the output logic unit.
[0140] If the processing unit clusters do not cache all the convolution kernel data, and the operation of the currently cached convolution kernel data is completed, the global cache unit loads the corresponding version of the convolution kernel data to the convolution kernel cache of each processing unit cluster again, refreshes the convolution kernel data, and continues the forward inference of the target neural network until the operation is completed.
[0141] If the processing unit clusters do not complete the inference process of all versions of the neural network, and the operation of the currently cached version of the convolution kernel data is completed, the global cache unit loads the remaining version of the convolution kernel data to the convolution kernel cache of each processing unit cluster again, refreshes the convolution kernel data, and continues the forward inference of the target neural network until the operation is completed.
[0142] In step 704, when the processing unit clusters complete the operation, the operation result is transmitted to the output logic unit.
[0143] In step 705, the uncertainty corresponding to the sampling point pulse signal is determined based on the operation result through the output logic unit, and output.
[0144] The operation result in the embodiment of the present application includes:
[0145] a predicted value and a prediction uncertainty;
[0146] or, a predicted value.
[0147] In a specific implementation, if the operation result includes a predicted value and a prediction uncertainty, and N versions of the neural network convolution kernel are used for inference, the uncertainty corresponding to the sampling point pulse signal is determined by the following formula:
[0148]
[0149]
[0150] where i is an integer from 1 to N, μ i is the predicted value corresponding to the i-th version of the neural network convolution kernel, and σ i is the uncertainty corresponding to the i-th version of the neural network convolution kernel.
[0151] If the operation result only includes a predicted value, and N versions of the neural network convolution kernel are used for inference, the uncertainty corresponding to the sampling point pulse signal is determined by the following formula:
[0152]
[0153]
[0154] where i is an integer from 1 to N, μ i is the predicted value corresponding to the i-th version of the neural network convolution kernel.
[0155] For the uncertainty estimation scenario, multiple versions of the neural network convolution kernel are used for multiple times of inference.
[0156] In the embodiments of the present application, multiple versions of the neural network convolution kernel can be pre-trained by randomly initializing parameters and stored in a global cache; according to the needs, the processing unit array is divided into a plurality of processing unit clusters; once a new sampling point pulse signal arrives at the input logic, it is loaded into the feature map cache of each processing unit cluster at the same time; each processing unit cluster uses one version of the convolution kernel for inference, and the timing result of the inference is transmitted to the output logic; if the number of processing unit clusters is less than the number of versions of the neural network convolution kernel, the convolution kernel cache of the processing unit cluster is refreshed through the global cache, and the new convolution kernel is used for inference until all versions of the neural network convolution kernel are inferred.
[0157] In another embodiment, a computer readable storage medium is also provided, which stores a computer program that is executed by a processor to implement the method of timing the timing processor in different application scenarios.
[0158] In another embodiment, a computer program product is also provided, comprising a computer program which, when executed by a processor, implements the method for timing the processor timing in different application scenarios.
[0159] Those skilled in the art can clearly understand the implementation of the embodiments by means of software and the necessary general hardware platform, and of course, the embodiments can also be implemented by hardware. Based on such understanding, the above technical solutions, essentially or in the sense of contribution to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer readable storage medium, such as a ROM / RAM, a magnetic disk, or an optical disc, and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0160] The above merely describes the preferred embodiments of the present application and is not intended to limit the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the scope of protection of the present application.
Claims
1. A multi-mode neural network timing method, characterized by, The timing processor applied to the uncertainty estimation optimization operation scene includes a processing unit, a global cache unit, an input logic unit and an output logic unit; data transmission is performed through an on-chip interconnection; and based on the number of layers and operation scale of a target neural network load, the processing unit array is divided into a preset number of processing unit clusters, and the processing units in each of the divided processing unit clusters share a feature map cache; wherein the processing unit is used to cache feature map and convolution kernel data; the processing unit array is a collection of all processing units in the timing processor; the method comprises: The global cache unit transmits different versions of target neural network convolution kernel data to the convolution kernel cache of each processing unit cluster; A sampling point pulse signal arrives at the input logic unit; If it is determined that the processing unit clusters are all idle, the sampling point pulse signal is loaded into the feature map cache of each processing unit cluster through the logic input unit; and the processing unit starts forward inference operation of the target neural network; When the processing unit clusters complete the operation, the operation result is transmitted to the output logic unit; Based on the operation result, the output logic unit determines the uncertainty corresponding to the sampling point pulse signal.
2. The method of claim 1, wherein, The operation result includes: a predicted value and a predicted uncertainty; Or, a predicted value.
3. The method of claim 1 or 2, wherein, The method further comprises: If the processing unit clusters do not cache all the convolution kernel data, and the operation of the currently cached convolution kernel data is completed, the global cache unit loads the corresponding version of the convolution kernel data into the convolution kernel cache of each processing unit cluster again, refreshes the convolution kernel data, and continues the forward inference of the target neural network until the operation is completed; If the processing unit clusters do not complete the inference process of all versions of the neural network, and the operation of the currently cached version of the convolution kernel data is completed, the global cache unit loads the remaining version of the convolution kernel data into the convolution kernel cache of each processing unit cluster again, refreshes the convolution kernel data, and continues the forward inference of the target neural network until the operation is completed.
4. A multi-mode neural network timing processor, characterized by, The timing processor applied to the uncertainty estimation optimization operation scene includes a processing unit, a global cache unit, an input logic unit and an output logic unit; data transmission is performed through an on-chip interconnection; and based on the number of layers and operation scale of a target neural network load, the processing unit array is divided into a preset number of processing unit clusters, and the processing units in each of the divided processing unit clusters share a feature map cache; wherein the processing unit array is a collection of all processing units in the timing processor; The global cache unit is configured to transmit different versions of target neural network convolution kernel data to the convolution kernel cache of each processing unit cluster; The input logic unit is configured to input a sampling pulse point signal; if it is determined that the processing unit clusters are all idle, the sampling point pulse signal is loaded into the feature map cache of each processing unit cluster; The processing unit is configured to cache feature map data and convolution kernel data, and perform forward inference operation of the target neural network; and transmit an operation result to the output logic unit when the cluster completes the operation. The output logic unit is configured to determine an uncertainty corresponding to the sampling point pulse signal based on the operation result.
5. The timing processor of claim 4, wherein, The operation result includes: a predicted value and a predicted uncertainty; or a predicted value.
6. The timing processor of claim 4 or 5, wherein: The global cache unit is further configured to, if all the convolution kernel data is not cached by the processing unit cluster, and operation of the currently cached convolution kernel data is completed, load corresponding versions of the convolution kernel data to the convolution kernel cache of each processing unit cluster again to refresh the convolution kernel data, and continue to perform the forward inference of the target neural network until the operation is completed; and if the processing unit cluster does not complete the inference process of all versions of the neural network, and operation of the currently cached version of the convolution kernel data is completed, load remaining versions of the convolution kernel data to the convolution kernel cache of each processing unit cluster again to refresh the convolution kernel data, and continue to perform the forward inference of the target neural network until the operation is completed. The program is executed by a processor to implement the method of any one of claims 1-3.
7. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by a processor to implement the method of any one of claims 1-3.
8. A computer program product comprising a computer program, characterized in that,
Citation Information
Patent Citations
Configurable convolution accelerator applied to convolutional neural network
CN110751280A
Neural network scheduling method and apparatus
WO2021237755A1