Low-power neural network system based on predictive exit and its implementation method
By predicting the exit point and calculating the remaining amount in the neural network, adjusting the processor voltage frequency in real time, and activating the corresponding exit layer, the problem of excessive calculation and energy overhead in the early exit design in the prior art is solved, and neural network reasoning with lower computing power and power consumption is achieved.
Patent Information
- Application Number
- CN202210630781.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-06
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2042-06-06
AI Technical Summary
In the early exit design, existing neural network hardware accelerators are difficult to effectively solve the dilemma of fine-grained and coarse-grained exit layers, resulting in excessive computing and energy overhead, and the inability to adjust the hardware circuit voltage and frequency in real time to achieve a design with lower computing power and power consumption.
A low-power neural network system based on predicting exit is proposed. By predicting the possible exit points and remaining calculation amount of the neural network, the processor voltage frequency is adjusted in real time, the corresponding exit layer is activated, and the appropriate calculation frequency and voltage are selected to run neural network inference.
A neural network inference process with lower computing power and lower power consumption is realized, which reduces the additional computing overhead introduced by early exit, makes full use of the opportunity of early exit, and saves computing power and energy consumption.
Smart Images

Figure CN114997370B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of neural network processing, and in particular to a low-power neural network system based on predictive exit and an implementation method thereof. Background Art
[0002] Although early exit is used in the design of neural network hardware accelerators to reduce computational and energy costs, it still cannot solve the following challenges. First, there is a dilemma between the fine-grained and coarse-grained placement of existing exit layers. Fine-grained exit layer placement results in significant computational and energy overhead due to the frequent execution of exit layers. On the contrary, coarse-grained placement of exit layers may miss the opportunity to exit earlier. Second, exit layers have different topologies at different locations, which will add an additional burden to map computations to computing hardware resources. Third, existing neural network accelerator methods based on early exit can only reduce the hardware circuit voltage and frequency after exiting. During the inference process, it is impossible to adjust the runtime hardware accelerator voltage and frequency in real time to achieve neural network design and implementation with lower computing power and lower power consumption.
[0003] Patent document EP3997621A1 discloses a system, method and device for early exit convolution, which compares the dot product value of a subset of an operand set with a threshold value through at least one PE circuit; at least one PE circuit determines whether to activate a node of a neural network based on at least the result of the comparison. This is the general implementation method of early neural network, which determines whether a neural network node is activated or not based on the limited setting of the exit layer and the setting of the exit layer. This ignores the possibility that neural network reasoning can achieve early exit at a certain network layer between the exit layers, and cannot adjust the processor voltage frequency in real time to achieve a neural network reasoning process with lower computing power and lower power consumption. Summary of the invention
[0004] In view of the defects in the prior art, the present invention proposes a predictive exit algorithm, which is simple and easy to understand and has a wide range of applications. The biggest feature of predictive exit is that by predicting the possible exit points of the neural network in advance and calculating the remaining computing amount, the processor voltage frequency is adjusted in real time, and finally a neural network reasoning process with lower computing power and lower power consumption is achieved.
[0005] A low-power neural network system based on predicted exit provided according to the present invention includes a low-power and low-computing power exit predictor;
[0006] The low-power and low-computing power exit predictor is capable of predicting the early exit point of neural network inference and activating the exit layer in response to the predicted early exit point, while selecting the required computing frequency and voltage to run the neural network inference based on the remaining computing workload before the predicted early exit point.
[0007] Furthermore, the low-power and low-computing power exit predictor comprises a zero-filling module, a one-dimensional convolutional layer module, a prediction function module, an activation start exit module, and a processor voltage and frequency adjustment module;
[0008] The zero filling module expands the dimension according to the result vector obtained by the low-power and low-computing power exit predictor from the predicted exit starting layer set in the neural network and the exit layer execution result vector corresponding to the result vector to obtain an expanded dimension output vector;
[0009] The one-dimensional convolution layer module performs a one-dimensional convolution operation on the dimension-expanded output vector of the zero-filling module and obtains a new feature weight vector;
[0010] The prediction function module determines whether the feature weight vector of the one-dimensional convolutional layer module satisfies the prediction function condition. When the feature weight vector satisfies the prediction function condition, the neural network level corresponding to the feature weight vector is the predicted early exit point.
[0011] The activation start exit module activates the positioning exit layer according to the early exit point of the prediction function module, and the positioning exit layer is the exit layer where the early exit point is located;
[0012] The processor voltage and frequency adjustment module selects a computing frequency and voltage to run the neural network reasoning before the positioning exit layer activated by the activation start exit module is executed according to the remaining computing workload in the neural network of the positioning exit layer.
[0013] Furthermore, the calculation process of the dimension-expanded output vector of the zero-filling module is as follows:
[0014]
[0015] The predicted exit starting layer L0 in the neural network execution result vector y0 and the corresponding exit layer execution result vector To expand the dimension, by Add 0 at both ends to convert the vector Depend on Dimensions expanded to dimension, and recorded as Where N c For vector The length is also the number of inference result categories, and K is the number of feature vectors in the feature pooling group.
[0016] Furthermore, the calculation process of the feature weight vector of the one-dimensional convolutional layer module is:
[0017]
[0018] in is the feature weight vector, is the expanded output vector, the convolution weight is a one-dimensional vector h of length K, K is the number of feature vectors in the feature pooling group, i = [1, 2, 3...], k = [1, 2, 3...], and L0 is the prediction exit starting layer.
[0019] Furthermore, the prediction function condition is:
[0020]
[0021] When satisfy When the condition The layer is the network layer level where the predicted early exit point is located, where Specify level thresholds for users; is the feature weight vector, and L0 is the prediction exit starting layer.
[0022] Furthermore, the exit layer is arranged at each level in the neural network, and the exit layer and the neural network share the same topological structure.
[0023] Further, the topological structure of the exit layer includes a feature pooling group and a fully connected layer;
[0024] The feature pooling group compresses the reasoning information to reduce redundancy;
[0025] The fully connected layer classifies the information processed by the feature pooling group and outputs a classification evaluation value;
[0026] When the evaluation value is greater than the threshold value entered by the user, the inference result is obtained in advance and the neural network is exited; otherwise, the neural network is returned to continue reasoning.
[0027] Furthermore, the placement of the exit layer in the neural network adopts a fine-grained exit layer virtual placement method;
[0028] The fine-grained exit layer virtual placement method is that the exit layer is virtually placed after each convolutional layer, and the reasoning information of the neural network has the opportunity to execute the exit layer after passing through the neural network calculation of each layer.
[0029] The present invention also provides a method for implementing a low-power neural network system based on predictive exit, which adopts the above-mentioned low-power neural network system based on predictive exit and also includes the following implementation steps:
[0030] Step 1 The neural network inputs the predicted exit starting position and related parameters specified by the user, and enters the first layer of the neural network to start calculation;
[0031] Step 2: When the neural network reaches the predicted exit starting position specified by the user, it starts to execute the predicted exit algorithm in the low-power and low-computing power exit predictor and obtains the predicted early exit point;
[0032] Step 3: The low-power and low-computing power exit predictor activates the early exit layer located at the early exit point; after executing the network layer located at the early exit point, the neural network inference will execute the early exit layer located at the network layer and exit early;
[0033] In Step 4, at the same time as Step 3, the low-power and low-computing-power exit predictor adjusts the computing frequency of the computing platform and the corresponding power supply voltage to the required level according to the current predicted exit starting position, the predicted early exit point position, and the total number of layers of the neural network.
[0034] Furthermore, Step 1 to Step 4 can be repeated to achieve continuous reasoning.
[0035] Compared with the prior art, the present invention has the following beneficial effects:
[0036] 1. The present invention predicts the possible exit points of the neural network in advance and calculates the remaining computing amount, adjusts the processor voltage frequency in real time, and ultimately achieves a neural network reasoning process with lower computing power and lower power consumption.
[0037] 2. In the exit layer part, the present invention solves the problem of complex calculations caused by the execution of multiple functions by computer processors, especially tensor processors and application-specific integrated circuit processors, by using the same exit layer design used in the network, thereby achieving the effect of reducing the additional computing overhead introduced by early exit.
[0038] 3. In the part of virtual placement of fine-grained exit layers, the present invention solves the extra computational overhead of running each exit layer by virtually placing the exit layer after each convolutional layer, thereby achieving the goal of reducing the amount of computation and making full use of early exit opportunities.
[0039] 4. In the low-power and low-computing-power exit predictor part of the present invention: by predicting possible early exit points of neural network reasoning and activating the early exit layer that responds to the predicted points, and calculating the remaining computing workload based on the expected exit points, the problem of selecting the appropriate computing frequency and voltage to run the inference is solved, thereby achieving the effect of further saving computing power and energy consumption. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Other features, objects and advantages of the present invention will become more apparent from the detailed description of non-limiting embodiments made with reference to the following drawings:
[0041] Figure 1A schematic diagram showing the structure comparison between the present invention and the existing neural network;
[0042] Figure 2 A schematic diagram of the logical structure of a low-power neural network system based on predictive exit according to the present invention;
[0043] Figure 3 A schematic diagram of the logic structure of the implementation method of the low-power and low-computing-power exit predictor of the present invention. DETAILED DESCRIPTION
[0044] The present invention is described in detail below in conjunction with specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but are not intended to limit the present invention in any form. It should be noted that, for those of ordinary skill in the art, several changes and improvements can also be made without departing from the concept of the present invention. These all belong to the protection scope of the present invention.
[0045] like Figure 1 As shown, in the process of completing reasoning, the neural network with an exit layer needs to traverse all the network layers set in the neural network (i.e., the network layers actually used in this reasoning) compared to the classical neural network. The neural network with an exit layer can end the reasoning in time after the exit layer is executed, without continuing the subsequent network layers (i.e., there are network layers that are not actually used in this reasoning). Compared with the classical neural network in (a) and the early exit neural network in (b), the core function of predicting exit of the present invention is shown in (c). Predicting exit can predict the reasoning process to obtain accurate reasoning results after passing through several layers of neural networks and start activating the exit layer at the corresponding position to realize exit. At the same time, the remaining neural network calculation amount is calculated according to the predicted exit position to adjust the processor voltage frequency, and finally ensure that the reasoning process is completed within a predetermined period while reducing the power consumption and energy consumption consumed by reasoning.
[0046] The present invention provides a low-power neural network system based on predicted exit, including a low-power and low-computing power exit predictor; the low-power and low-computing power exit predictor can predict the early exit point of neural network reasoning and activate the exit layer in response to the predicted early exit point, and at the same time, based on the remaining computing workload before the predicted early exit point, select the required computing frequency and voltage to run the neural network reasoning. Predicted exit can accurately predict the early exit point of reasoning during the process of neural network reasoning and start the early exit layer at this point while adjusting the processor voltage frequency, thereby reducing the amount of computing, power consumption and energy consumption consumed by the reasoning process while ensuring that the reasoning is completed before the deadline.
[0047] like Figure 2As shown, the structure of the neural network includes three new functional modules: exit layer, fine-grained exit layer virtual placement, and low-power and low-computing power exit predictor.
[0048] The exit layer can be arranged at each level in the neural network, and all exit layers share the same topology. In practical applications, the neural network runs on a graphics processing unit, a tensor processing unit, or an application-specific integrated circuit on a local or edge device. In order to reduce the burden of performing multi-function calculations, in the predictive exit design, all exit layers in the network will share the same topology. The topology can be any exit layer topology as long as the accuracy of the exit judgment is high.
[0049] This implementation adopts an optimized exit layer topology, which specifically includes a feature pooling group and a fully connected layer; the feature pooling group compresses the reasoning information to reduce redundancy; the fully connected layer classifies the information processed by the feature pooling group and outputs a classification evaluation value; when the evaluation value is greater than the threshold entered by the user, the reasoning result is obtained in advance and the neural network is exited, otherwise the neural network is returned to continue reasoning. In particular, the weights of the operators in the feature pooling group and the fully connected layer can be obtained through neural network pre-training. One of the preferred neural network training methods is to use cross entropy as the reinforcement learning target and train the neural network containing the exit layer using the back propagation algorithm. By using the same exit layer design in the network, the problem of complex calculations caused by the execution of multiple functions by computer processors, especially tensor processors and application-specific integrated circuit processors, is solved, and the effect of reducing the calculations introduced by early exit is achieved.
[0050] Fine-grained virtual placement of exit layers: In this embodiment, the placement of the exit layer in the neural network adopts a fine-grained virtual placement of exit layers; specifically, the exit layer is virtually placed after each convolutional layer, and the reasoning information of the neural network has the opportunity to execute the exit layer after each layer of neural network calculation. By virtually placing the exit layer after each convolutional layer, the additional computational overhead of running each exit layer is solved, and the amount of computation is reduced and the opportunity for early exit is fully utilized.
[0051] In particular, in order to avoid the computational overhead of running each exit layer but not successfully exiting, the inference process will only execute the exit layers at the exit points predicted by the low-power and low-computation exit predictor, that is, only the exit layers at the predicted exit points are activated.
[0052] Low-power and low-computing power exit predictor. The core of the low-power neural network system based on predicted exit is the low-power and low-computing power exit predictor, which predicts the possible early exit point of neural network reasoning and activates the early exit layer in response to the predicted point. Based on the remaining computing workload before the expected exit point, the appropriate computing frequency and voltage will be selected to run the inference and further save computing power and energy consumption. By predicting the possible early exit point of neural network reasoning and activating the early exit layer in response to the predicted point, and calculating the remaining computing workload based on the expected exit point, the problem of selecting the appropriate computing frequency and voltage to run the inference is solved, achieving the effect of further saving computing power and energy consumption.
[0053] This implementation adopts an optimized low-power and low-computing power exit predictor, which specifically includes a zero-filling module, a one-dimensional convolutional layer module, a prediction function module, an activation and start-up exit module, and a processor voltage and frequency adjustment module; the zero-filling module expands the dimension according to the result vector obtained from the predicted exit starting layer set in the neural network by the low-power and low-computing power exit predictor and the exit layer execution result vector corresponding to the result vector (i.e., making full use of the intermediate result information in the neural network), and obtains the expanded-dimensional output vector; the one-dimensional convolutional layer module performs a one-dimensional convolution operation on the expanded-dimensional output vector of the zero-filling module and obtains a new feature weight vector; the prediction function module determines whether the feature weight vector of the one-dimensional convolutional layer module meets the prediction function condition, and when the feature weight vector meets the prediction function condition, the neural network layer corresponding to the feature weight vector is the predicted early exit point; the activation and start-up exit module activates the positioning exit layer according to the early exit point of the prediction function module, and the positioning exit layer is the exit layer where the early exit point is located; the processor voltage and frequency adjustment module selects the calculation frequency and voltage to run the neural network reasoning before the positioning exit layer is executed according to the remaining calculation workload of the positioning exit layer activated by the activation and start-up exit module in the neural network.
[0054] Specifically, the working principle of the low-power and low-computing power exit predictor and its modules is as follows: Figure 3 As shown:
[0055] 1) Assume that the predictor starts predicting from the L0th layer of the neural network, and the predictor input is the execution result vector y0 of the L0th layer neural network and its corresponding exit layer execution result vector
[0056] 2) The zero-filling module expands the dimension of the above results by adding Add 0 at both ends to convert the vector Depend on Dimensions expanded to dimension, and recorded as Where N c For vector The length is also the number of inference result categories, and K is the number of feature vectors in the feature pooling group. The specific calculation process is shown in the following formula:
[0057]
[0058] 3) Output vector after zero padding module A one-dimensional convolution operation will be performed and a new feature weight vector will be obtained The convolution weight is a one-dimensional vector h of length K. The specific calculation process is shown in the following formula:
[0059]
[0060] 4) By repeating the above calculation process, we can get Until That is, the calculation steps of the prediction function are:
[0061] 5) When satisfy Condition, is the predicted early exit position, where Specifies the threshold value for the user. The larger it is, the more accurate the prediction will be, and the predicted position will be closer to the back of the neural network. When it is infinite, the prediction result is the last layer of the neural network. Users can choose the threshold value according to actual usage.
[0062] 6) Predictor start activation is located at the predicted exit position Exit layer.
[0063] 7) At the same time, the predictor adjusts the processor computing frequency (and the corresponding supply voltage) to where f max The maximum frequency supported by the processor, L total is the total number of layers in the neural network.
[0064] The implementation method of the low-power neural network system based on predicted exit in this embodiment adopts the above-mentioned low-power neural network system based on predicted exit, and also includes the following implementation steps:
[0065] Step 1 The neural network inputs the predicted exit starting position and related parameters specified by the user, and enters the first layer of the neural network to start calculation;
[0066] Step 2: When the neural network reaches the predicted exit starting position specified by the user, it starts to execute the predicted exit algorithm in the low-power and low-computing power exit predictor and obtains the predicted early exit point;
[0067] Step 3: The low-power and low-computing power exit predictor activates the early exit layer located at the early exit point; after executing the network layer located at the early exit point, the neural network inference will execute the early exit layer located at the network layer and exit early;
[0068] In Step 4, at the same time as Step 3, the low-power and low-computing-power exit predictor adjusts the computing frequency of the computing platform and the corresponding power supply voltage to the required level according to the current predicted exit starting position, the predicted early exit point position, and the total number of layers of the neural network.
[0069] In particular, Step 1 to Step 4 can be repeated to achieve continuous reasoning.
[0070] Specifically, the implementation method and principle of the low-power neural network system based on predictive exit are as follows:
[0071] Step 1 The input of the neural network (such as pictures, voice, etc.) is in the overall structure Figure 2 The left side input enters the first layer of the neural network to start calculation.
[0072] Step 2: When the neural network reaches the predicted exit starting position L0 specified by the user, it starts to execute the predicted exit algorithm in the low-power and low-computing power exit predictor and obtains the predicted early exit position.
[0073] Step3 Low power consumption, low computing power exit predictor activation start located at position The neural network inference will be executed after the first After the layer network, execution is at position The early exit layer is used to exit early.
[0074] Step 4 is simultaneous with Step 3. The low-power, low-computing-power exit predictor calculates the predicted early exit position based on the current predicted starting point L0 and the predicted early exit position. And the total number of layers of the neural network L total . Adjust the computing frequency (and corresponding power supply voltage) of the computing platform (such as CPU, GPU, etc.) to where f max The highest frequency supported by the processor.
[0075] The above steps are the specific steps to complete one reasoning (one picture, one voice). Continuous reasoning is the repetition of the above process.
[0076] Experiments have shown that, taking VGG-19 and ResNet-34 neural networks as examples, through the tests of CIFAR-10, CIFAR-100, SVHN and STL10 standard test sets, compared with traditional neural networks, the exit prediction method can reduce 96.2% of the computational workload and 72.9% of the energy consumption; compared with the neural network with only early exit, the exit prediction method can reduce 12.8% of the computational workload and 37.6% of the energy consumption.
[0077] In the description of the present application, it should be understood that the terms "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", etc., indicating orientations or positional relationships, are based on the orientations or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present application and simplifying the description, and do not indicate or imply that the referred device or element must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be understood as a limitation on the present application.
[0078] Those skilled in the art know that, in addition to realizing the system and its various devices, modules, and units provided by the present invention in a purely computer-readable program code, it is entirely possible to realize the same functions in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers by logically programming the method steps. Therefore, the system and its various devices, modules, and units provided by the present invention can be considered as a hardware component, and the devices, modules, and units included therein for realizing various functions can also be regarded as structures within the hardware component; the devices, modules, and units for realizing various functions can also be regarded as both software modules for realizing the method and structures within the hardware component.
[0079] The above describes the specific embodiments of the present invention. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art can make various changes or modifications within the scope of the claims, which does not affect the essence of the present invention. In the absence of conflict, the embodiments of the present application and the features in the embodiments can be combined with each other arbitrarily.
Claims
1. A low-power neural network system based on predictive exit, characterized in that: Includes low power and low computing power exit predictor; The low-power and low-computing power exit predictor is capable of predicting an early exit point of neural network reasoning and activating an exit layer in response to the predicted early exit point, while selecting a required computing frequency and voltage to run the neural network reasoning based on the remaining computing workload before the predicted early exit point; The low-power and low-computing power exit predictor includes a zero-filling module, a one-dimensional convolutional layer module, a prediction function module, an activation start exit module, and a processor voltage and frequency adjustment module; The zero filling module expands the dimension according to the result vector obtained by the low-power and low-computing power exit predictor from the predicted exit starting layer set in the neural network and the exit layer execution result vector corresponding to the result vector to obtain an expanded dimension output vector; The one-dimensional convolution layer module performs a one-dimensional convolution operation on the dimension-expanded output vector of the zero-filling module and obtains a new feature weight vector; The prediction function module determines whether the feature weight vector of the one-dimensional convolutional layer module satisfies the prediction function condition. When the feature weight vector satisfies the prediction function condition, the neural network level corresponding to the feature weight vector is the predicted early exit point. The activation start exit module activates the positioning exit layer according to the early exit point of the prediction function module, and the positioning exit layer is the exit layer where the early exit point is located; The processor voltage and frequency adjustment module selects a computing frequency and a voltage to run the neural network reasoning before the positioning exit layer is executed according to the remaining computing workload of the positioning exit layer activated by the activation start exit module in the neural network; The placement of the exit layer in the neural network adopts a fine-grained exit layer virtual placement method; The fine-grained exit layer virtual placement method is that the exit layer is virtually placed after each convolutional layer, and the reasoning information of the neural network has the opportunity to execute the exit layer after passing through the neural network calculation of each layer.
2. The low-power neural network system based on predictive exit according to claim 1, characterized in that: The exit layers are arranged at various levels in the neural network, and all the exit layers share the same topological structure.
3. The low-power neural network system based on predictive exit according to claim 1, characterized in that: The topological structure of the exit layer includes a feature pooling group and a fully connected layer; The feature pooling group compresses the reasoning information to reduce redundancy; The fully connected layer classifies the information processed by the feature pooling group and outputs a classification evaluation value; When the evaluation value is greater than the threshold value entered by the user, the inference result is obtained in advance and the neural network is exited; otherwise, the neural network is returned to continue reasoning.
4. The low-power neural network system based on predictive exit according to claim 1, characterized in that: The calculation process of the expanded dimension output vector of the zero padding module is: The predicted exit starting layer L0 in the neural network execution result vector y0 and the corresponding exit layer execution result vector To expand the dimension, by Add 0 at both ends to convert the vector Depend on Dimensions expanded to dimension, and recorded as Where N c For vector The length is also the number of inference result categories, and K is the number of feature vectors in the feature pooling group.
5. The low-power neural network system based on predictive exit according to claim 1, characterized in that: The calculation process of the feature weight vector of the one-dimensional convolutional layer module is: in is the feature weight vector, is the expanded output vector, the convolution weight is a one-dimensional vector h of length K, K is the number of feature vectors in the feature pooling group, i = [1, 2, 3...], k = [1, 2, 3...], and L0 is the prediction exit starting layer.
6. The low-power neural network system based on predictive exit according to claim 1, characterized in that: The prediction function condition is: When satisfy Condition, The layer is the network layer level where the predicted early exit point is located, where Specify level thresholds for users; is the feature weight vector, and L0 is the prediction exit starting layer.
7. A method for implementing a low-power neural network system based on predictive exit, characterized in that: The low-power neural network system based on predictive exit according to any one of claims 1 to 6 is adopted, and further comprises the following implementation steps: Step 1 The neural network inputs the predicted exit starting position and related parameters specified by the user, and enters the first layer of the neural network to start calculation; Step 2: When the neural network reaches the predicted exit starting position specified by the user, it starts to execute the predicted exit algorithm in the low-power and low-computing power exit predictor and obtains the predicted early exit point; Step 3: The low-power and low-computing power exit predictor activates the early exit layer located at the early exit point; after executing the network layer located at the early exit point, the neural network inference will execute the early exit layer located at the network layer and exit early; In Step 4, at the same time as Step 3, the low-power and low-computing-power exit predictor adjusts the computing frequency of the computing platform and the corresponding power supply voltage to the required level according to the current predicted exit starting position, the predicted early exit point position, and the total number of layers of the neural network.
8. The method for implementing a low-power neural network system based on predictive exit according to claim 7, characterized in that: Step 1 to Step 4 can be repeated to achieve continuous reasoning.
Citation Information
Patent Citations
Systems, methods, and devices for early-exit from convolution
EP3997621A1