Asynchronous training method and device for online learning AI chip

By using asynchronous training methods and 1LSB random quantization technology on online learning AI chips, the problem of high power consumption of edge-end AI chip training is solved, and a more efficient training and reasoning process is achieved.

CN119358700BActive Publication Date: 2025-05-16TSINGHUA UNIVERSITY
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411391271.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-30
Publication Date
2025-05-16
Estimated Expiration
2044-09-30

AI Technical Summary

Technical Problem

When online learning AI chips run on the edge for a long time and rely on battery power, the training power consumption is extremely high, limiting its development.

Method used

The asynchronous training method is adopted, and the calculation method of the calculation module is scheduled by setting an asynchronous learning engine, reducing the number of training times for updating all synaptic weights, and discarding the lowest data bit of the synaptic weights during storage, and multiplexing random quantized data during inference.

Benefits of technology

It effectively reduces the training power consumption and storage cost of online learning AI chips, and improves training efficiency and reasoning accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119358700B_ABST
    Figure CN119358700B_ABST
Patent Text Reader

Abstract

The present disclosure provides an asynchronous training method and device for online learning AI chips. The method includes: when the AI ​​chip is in training mode, initializing the network synaptic weights of the network model in the AI ​​chip; deploying an asynchronous learning engine according to a built-in algorithm, and the asynchronous learning engine is used to schedule the calculation method of the calculation module; inputting training sample data into the network model in the AI ​​chip so that the calculation module updates the network synaptic weights; determining whether the target synaptic weight has completed training based on the number of updates of the target synaptic weight in the network model and a preset number threshold. In this way, in this embodiment, an asynchronous learning engine is set to adjust the calculation method of the calculation module, which can reduce the number of training times for updating all synaptic weights.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular to an asynchronous training method and device for online learning AI chips. Background Art

[0002] With the rapid development of artificial intelligence technology, artificial intelligence has penetrated into all aspects of daily life, for example, from smartphones to self-driving cars. The early training of artificial intelligence often adopts offline learning mode, but the traditional offline learning mode is unable to meet the real-time and personalized requirements, which has prompted the birth of online learning AI chips.

[0003] Online learning AI chips have demonstrated their unique advantages in edge applications and are widely used in the fields of Internet of Things, smart home, smart health, etc. They can realize instant calculation and decision-making where the data is generated, without the need to transmit the data to the cloud, thus meeting the real-time requirements of scenarios such as automatic obstacle avoidance.

[0004] However, online learning AI chips have extremely high requirements for power consumption due to long-term operation at the edge and reliance on battery power. The power consumption of online learning AI chips is mainly divided into two parts: training power consumption and inference power consumption, of which training power consumption is usually much greater than inference power consumption. This is because the training phase involves a large number of computationally intensive operations, which require constant adjustment of model parameters, requiring not only a large amount of data processing, but also an increase in data transmission and storage. As the complexity of application scenarios increases, training power consumption also increases accordingly, limiting the development of online learning AI chips. Summary of the invention

[0005] The present invention provides an asynchronous training method and device for online learning of AI chips.

[0006] According to a first aspect of the present disclosure, there is provided an asynchronous training method for an online learning AI chip, the method comprising:

[0007] When the AI ​​chip is in training mode, the network synaptic weights of the network model in the AI ​​chip are initialized;

[0008] Deploy an asynchronous learning engine according to a built-in algorithm, wherein the asynchronous learning engine is used to schedule a computing mode of a computing module;

[0009] Inputting training sample data into the network model in the AI ​​chip so that the computing module updates the network synaptic weights;

[0010] Whether the target synaptic weight has completed training is determined according to the update times of the target synaptic weight in the network model and a preset times threshold.

[0011] Optionally, determining whether the target synaptic weight has completed training according to the update times of the target synaptic weight in the network model and a preset times threshold includes:

[0012] Comparing the update times of the target synaptic weight in the network model with the preset times threshold;

[0013] When the number of updates is less than or equal to the preset number threshold, determining that the training is not completed and jumping to the step of inputting the training sample data into the network model in the AI ​​chip;

[0014] When the update number is greater than the preset number threshold, it is determined that the target synapse has completed training and a first synaptic weight of the target synapse is obtained.

[0015] Optionally, the method further comprises:

[0016] discarding the lowest data bit of the first synaptic weight to obtain a second synaptic weight;

[0017] The second synaptic weight is stored in a specified location.

[0018] Optionally, the method further comprises:

[0019] When the AI ​​chip is in inference mode, random quantized data is obtained;

[0020] combining the second synaptic weight and the random quantized data to obtain a third synaptic weight;

[0021] The weight of the target synapse of the network model is updated to the third synaptic weight to obtain a target network model, and the target network model is used to infer the result according to the input data.

[0022] Optionally, random quantitative data is obtained, including:

[0023] The second synaptic weight is randomly quantized by 1LSB through an asynchronous random number generator, and the asynchronous random number generator works when it detects that the AI ​​chip reads the weight data, otherwise it is in a standby state;

[0024] The asynchronous random number generator is a multi-stage pipeline composed of a Click unit and a linear feedback shift register LSFR, which is used to generate different random numbers according to an input seed, and use the lowest bit data of the random number as the random quantization data.

[0025] Optionally, the method further comprises:

[0026] The training samples are input into the network model in the AI ​​chip through a bundled data asynchronous circuit based on the Click unit; when the request input end of the Click unit detects that the request signal is flipped, a fire pulse signal is output, and the fire pulse signal is used to control data input and data output.

[0027] Optionally, the asynchronous learning engine includes a control module and a calculation module, the control module generates a control code according to preset configuration data, and the control code is used to schedule a calculation method of the calculation module.

[0028] Optionally, the calculation method of the calculation module includes at least one of the following: addition, subtraction, multiplication, division, exponential operation, logarithmic operation, and trigonometric function operation.

[0029] Optionally, when the calculation mode of the calculation module is multiplication or division, the calculation module is implemented by a multi-stage pipeline structure consisting of a Click unit, a shifter and an adder;

[0030] or,

[0031] When the calculation mode of the calculation module is exponential operation or logarithmic operation, the calculation module is implemented by a Click unit and a lookup table LUT;

[0032] or,

[0033] When the calculation mode of the calculation module is trigonometric function operation, the calculation module is implemented by a Click unit, an adder, a shifter and a lookup table LUT.

[0034] According to a second aspect of the present disclosure, there is provided an asynchronous training device for online learning of an AI chip, the device comprising:

[0035] A weight initialization module, used to initialize the network synaptic weights of the network model in the AI ​​chip when the AI ​​chip is in training mode;

[0036] An engine configuration module, used to deploy an asynchronous learning engine according to a built-in algorithm, wherein the asynchronous learning engine is used to schedule a calculation method of a calculation module;

[0037] A training sample input module, used to input training sample data into the network model in the AI ​​chip so that the computing module updates the network synaptic weights;

[0038] The training completion determination module is used to determine whether the target synaptic weight has completed training according to the number of updates of the target synaptic weight in the network model and a preset number threshold.

[0039] The technical solution provided by the embodiments of the present disclosure may have the following beneficial effects:

[0040] The present embodiment provides an asynchronous training method for an online learning AI chip, including: when the AI ​​chip is in training mode, initializing the network synaptic weights of the network model in the AI ​​chip; then, deploying an asynchronous learning engine according to a built-in algorithm, the asynchronous learning engine is used to schedule the calculation method of the calculation module; thereafter, inputting the training sample data into the network model in the AI ​​chip so that the calculation module updates the network synaptic weights; finally, determining whether the target synaptic weight has completed training based on the number of updates of the target synaptic weight in the network model and a preset number threshold. In this way, in this embodiment, an asynchronous learning engine is set to adjust the calculation method of the calculation module, which can reduce the number of training times for updating all synaptic weights.

[0041] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 This is a schematic diagram of the structure of a multi-input multi-output Click unit with an enablement according to an embodiment of the present disclosure.

[0043] Figure 2 The present invention is a schematic diagram of the structure of an asynchronous learning engine according to an embodiment of the present invention.

[0044] Figure 3 The figure is a schematic diagram of the structure of an addition and subtraction operator according to an embodiment of the present disclosure.

[0045] Figure 4 The figure is a schematic diagram of the structure of a multiplication and division operator according to an embodiment of the present disclosure.

[0046] Figure 5 The present invention is a schematic diagram of the structure of an asynchronous random number generator according to an embodiment of the present invention.

[0047] Figure 6 This is a flowchart of an asynchronous training method for an online learning AI chip according to an embodiment of the present disclosure.

[0048] Figure 7 This is a flowchart of another asynchronous training method for an online learning AI chip according to an embodiment of the present disclosure.

[0049] Figure 8 This is a block diagram of an asynchronous training device for online learning AI chips according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0050] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of devices consistent with some aspects of the present disclosure as detailed in the appended claims.

[0051] In order to solve the above technical problems, the embodiments of the present disclosure provide an asynchronous training method and device for an online learning AI chip. The asynchronous training method for an online learning AI chip provided by the present disclosure has the following inventive concept: first, in terms of computing, an asynchronous learning engine is provided, and the asynchronous learning engine includes a control module, which can control the computing method of at least one computing module in the network model, thereby adjusting the weight of the target synapse corresponding to the computing module, which can reduce the power consumption of chip training. Second, in terms of storage, a 1LSB random quantization generator is provided, and the lowest data bit of the synaptic weight is discarded during storage, and the last bit of the random number generated by the 1LSB random quantization generator is reused during inference as the lowest data bit of the above synaptic weight, which can reduce the storage cost of the synaptic weight.

[0052] Considering that the asynchronous training method for online learning AI chips provided in the present disclosure involves some modules, the following embodiments first explain these modules and then describe the asynchronous training method.

[0053] See also Figure 1 This embodiment provides a multi-input multi-output Click unit with enable. The multi-input multi-output Click unit includes a control path 11 and a data path 20, and the data path 20 and the control path 11 are decoupled for processing.

[0054] In one example, the control path 11 is implemented using a multi-level chain structure with Click units as basic units, and the number of input and output channels and / or request response signal generation behavior of each level of Click units can be set according to specific scenarios.

[0055] Understandably, Figure 1The example shows that the control path 11 includes two Click units (i.e., Click111 and Click112) and a delay unit Delay_Unit. Each Click unit includes two input terminals and two output terminals. The request signal In_Req can be input to the first input terminal of the Click unit Click111, and then output to the delay unit Delay_Unit by the first output terminal; the input terminal of the Click unit Click112 can receive the request signal after the delay, and output the request signal Out_Req through the first output terminal. When the request signal Out_Req is processed, a response signal can be obtained. At this time, the second input terminal of the Click unit Click112 can receive the above-mentioned response signal Out_Ack, and transmit it to the second input terminal of the Click unit Click111 through the second output terminal, and the second output terminal of the Click unit Click112 can output the above-mentioned response signal In_Ack.

[0056] It should be noted that in this example, the Click unit adopts a two-phase handshake protocol, that is, each time the request signal received by the input terminal In_Req of the Click unit flips (such as high level to low level or low level to high level), it represents that a request signal has been received; the Click unit generates a fire pulse signal for each request signal. The fire pulse signal can be used as a clock signal to capture and store data, or to control the input or output data of the trigger. At the same time, it will also drive the D flip-flop to transmit internal signals. The D flip-flop will be controlled by the external input control signal Ctrl to generate the control input response signal In_Ack and the output request signal Out_Req signal flip, and determine whether the signal is transmitted to the next level Click for request signals.

[0057] It should be noted that, in this example, the delay duration of the delay unit Delay_Unit is set to be greater than the calculation duration of the Logic module in the data path, so as to output a normal control signal.

[0058] In one example, the data path may include a D flip-flop DFF and a computing module Logic. It should be noted that the number of the D flip-flop DFF and the computing module Logic may be set according to a specific scenario. Figure 1The control path is provided with two triggers, namely trigger DFF121 and trigger DFF123, and calculation module Logic 122. Trigger DFF1 can receive the fire pulse signal output by Click unit 01, and forward the input data Data_in to calculation module Logic 122. Calculation module Logic 122 processes the input data according to the preset calculation method to obtain the synaptic weight, and sends the synaptic weight to trigger DFF2. Trigger DFF2 can receive the fire pulse signal output by Click unit 02 and output the synaptic weight Data_out.

[0059] See also Figure 2 This embodiment provides an asynchronous learning engine. The asynchronous learning engine includes a control module 21 and a data module 22. The control path 21 is a cascade circuit composed of Click units, including Click units Click211, Click212 and Click213. The data module 22 includes a selector 221, an operator unit 222 and a register 223. Among them,

[0060] The click unit Click211 is connected to the selector 221, and the selector 221 provides an enable signal for the click unit Click211. After receiving the enable signal, the click unit Click211 can receive the input request signal In_Req or output the input response signal In_Ack, that is, the click unit Click211 controls the input and output data after receiving the enable signal. In addition, the click unit Click211 can also provide a fire pulse signal to the selector, and the fire signal will serve as a local clock signal to guide the selector to transfer data to different operator units under the control of the control code. The enable signal of the selector 221 is determined by the control code. When all calculations are completed, the enable signal is high, and the control click unit no longer generates a request signal. The loop chain stops, and the output of the last stage DFF is no longer an intermediate result but a final result.

[0061] Click unit Click212 is connected to operator unit 222. It is understandable that operator unit 222 includes multiple operators, such as Figure 2The scenario of N operators such as operator 1, operator 2, ..., operator N is illustrated, which can be set according to the function of operator unit 222, and the corresponding scheme falls within the data scope of the present disclosure. Considering that each Click unit provides a fire pulse signal for one operator, the number of Click units in Click212 is equal to the number of operators in operator unit 222. It is understandable that the request signals of both Click unit Click212 and Click unit Click211 may be delayed, and the delay duration can be set according to the calculation duration of operator unit 222, and the delay duration is greater than the calculation duration. For details, refer to Figure 1 The setting of the delay time length of the delay unit Delay_Unit will not be repeated here.

[0062] The Click unit in Click212 can generate a fire pulse signal when receiving a request signal, and the Click unit in Click212 provides a fire pulse signal to the (matched) operator. The operator can participate in data calculation in an enabled state, but not in a disabled state. Alternatively, the operator can output a calculation result in an enabled state, but not in a disabled state. That is, the effect of the fire pulse signal on the operator can be adjusted according to the specific scenario, and the corresponding solution falls within the protection scope of the present disclosure.

[0063] The first input terminal of the selector 221 can obtain the original data and the second input terminal can obtain the intermediate result, and then output data (original data and / or intermediate data) to the corresponding at least one operator according to the control code, so as to facilitate the operator calculation. The operator can calculate the received quantity in the enabled state and obtain the calculation result. After receiving the fire pulse signal provided by the Click unit, the calculation result can be output to the register.

[0064] Click unit Click213 is connected to register 223, and the number of the two can be set according to the specific scenario. Figure 2 The example of a scenario where both are set is shown in FIG. It is understandable that upon receiving the request information, the Click unit Click213 can provide a fire pulse signal to the register, thereby outputting the synaptic weight in the register to a specified location (such as a local cache, a memory, or a cloud, etc.).

[0065] It should be noted that the asynchronous learning engine provided in this embodiment is composed of a variety of Click units, selectors and triggers DFF and different asynchronous operators. The operator determines which calculation methods the asynchronous learning engine supports. The above-mentioned calculation methods may include at least one of the following: addition, subtraction, multiplication, division, exponential operation, logarithmic operation, trigonometric function operation. For example, when the calculation method of the calculation module is multiplication or division, the calculation module is implemented by a multi-stage pipeline structure composed of Click units, shifters and adders; or, when the calculation method of the calculation module is exponential operation or logarithmic operation, the calculation module is implemented by Click units and lookup tables LUT; or, when the calculation method of the calculation module is trigonometric function operation, the calculation module is implemented by Click units, adders, shifters and lookup tables LUT. The corresponding operation or combination of operations can be selected according to the specific scenario, and the corresponding scheme falls within the protection scope of this disclosure.

[0066] Taking the synaptic weight update method STDP (spike timing dependent plasticity) in the spiking neural network as an example, an example weight update calculation formula is Δ w =a×exp b-c . Where a, b and c are three input data, Δ w is the calculation result. Assume that a, b and c are coded as 01, 10, 11 respectively. Figure 2 , the control code CTRL[15:0] first enters 0101_0010_0110_0111, CTRL[15:14]=01 indicates the number of current calculations, CTRL[13:12]=01 indicates that the result of this calculation is stored in the first D flip-flop, CTRL[11:8]=0010 indicates the operator used in this calculation, such as 0010 indicates that the subtraction operation is used this time, CTRL[7:4] indicates the first calculation object of this calculation, CTRL[7:6]=01 indicates that the first calculation object is the original input data rather than the storage result, and CTRL[5:4]=10 corresponds to the code of variable b. CTRL[3:0] indicates the second calculation object of this calculation. Therefore, 0101_0010_0110_0111 indicates that this calculation is the first calculation, realizing the function of bc, and storing the result in the first flip-flop DFF, which is regarded as the storage result. Then, the control code CTRL[15:0] is input 1010_0101_1001_0000 to implement exp b-c The function of a×exp is used to achieve the complete calculation. b-c The control code is used to finally obtain the synaptic weight.

[0067] The following takes the example that the operator unit 222 includes addition and subtraction operators or multiplication and division algorithms to describe the structure and working principle of each algorithm.

[0068] Take the example that the operator unit 222 includes addition and subtraction operators, see Figure 3 , the addition and subtraction operator may include a complement unit 31, an adder unit 32, a click unit 33, a trigger DFF34 and a trigger DFF35. The addition and subtraction operator may obtain input data x[n:0] and input data y[n:0]. The complement unit 31 performs complement processing on the input data y[n:0] in the channel selection signal sel to obtain complement data. The adder 32 may perform addition processing on the above complement data and the input data x[n:0], and send the addition result to the trigger DFF34 and the trigger DFF35. After receiving the request information, the click unit 33 may generate a fire pulse signal to the trigger DFF34 and the trigger DFF35 respectively. When the trigger DFF34 or the trigger DFF35 receives the selection of the channel selection signal sel, one of the triggers outputs the addition result or the subtraction result. In this way, the operator unit 222 in this embodiment may realize the asynchronous addition or asynchronous subtraction function.

[0069] Take the example that the operator unit 222 includes multiplication and division operators, see Figure 4 , the multiplication and division operator may include a control path 41 and a data path 42. The control path 41 includes a click unit Click411 and a click unit Click412. The data path 42 includes a shift unit 421, an adder 422 and a trigger DFF423. The operator unit includes input data x[n:0] and input data y[n:0], wherein the input data y[n:0] is represented in the form of each data bit, namely y[n], y[n-1], ..., y[0]. The shift unit 421 may include a plurality of shift operators, which may be implemented using a linear feedback shift register LSFR. Each shift operator performs a shift process on the input data x[n:0], data bits y[n], y[n-1], ... and y[0] to obtain a shift result. After receiving a request signal, the click unit Click411 may generate a fire pulse signal and send it to each shift operator. After receiving the fire pulse signal, each shift operator outputs the shift result to the adder 422. The adder 422 adds each shift result and sends the calculation result to the trigger DFF423. When the trigger DFF423 receives the fire pulse signal output by the click unit Click412, it can output the above addition result as a multiplication or division result.

[0070] This embodiment also provides an asynchronous random number generator, which is a multi-stage pipeline composed of a Click unit and a linear feedback shift register LSFR, and is used to generate different random numbers according to an input seed, and use the lowest bit data of the random number as the random quantization data. Figure 5 The asynchronous random number generator includes a control path 51 and a data path 52. The control path 51 is a cascade circuit composed of multiple stages of Click units. Figure 5 The scenario of three cascaded Click units is illustrated in the figure. The data path 52 is a cascade circuit composed of a trigger and an adder. Each Click unit provides a fire pulse signal to the matching trigger DFF. The input seed g0 is output to the first adder through the first trigger DFF0. The first adder adds the input seed g0 and the input seed g1, and sends the result to the second trigger DFF1. After the second trigger DFF1 receives the fire pulse signal provided by the Click unit, it outputs the result to the second adder. The second adder adds the above result and the input seed g2, and adds the result to the input seed g3 to obtain an asynchronous random number. In this example, the asynchronous random number generator only works when the AI ​​chip reads the weight data, otherwise it processes the standby state, thereby reducing power consumption.

[0071] Based on the above-mentioned newly added modules, the present disclosure provides an asynchronous training method for online learning AI chips, see Figure 6 , including steps 61 to 64.

[0072] In step 61, when the AI ​​chip is in training mode, the network synaptic weights of the network model in the AI ​​chip are initialized.

[0073] In this step, a network model is pre-set in the AI ​​chip, and the network model includes a number of neurons, each of which is connected to other neurons in the network model through synapses. Before training, the initial values ​​of the network synaptic weights of the network model in the AI ​​chip can be pre-stored in a specified location, such as a local cache, a local memory, or the cloud.

[0074] In this step, during the training phase of the AI ​​chip, that is, when the AI ​​chip is in the training model, the synaptic weights of the network synapses in the network model can be initialized, that is, the initial value of each network synapse is read from the specified position and used as the original input data.

[0075] In step 62, an asynchronous learning engine is deployed according to a built-in algorithm, and the asynchronous learning engine is used to schedule the calculation mode of the calculation module.

[0076] In this step, a built-in algorithm can be pre-stored, which is used to generate Figure 2The control code is CTRL.

[0077] In this step, the above-mentioned built-in algorithm can be called to deploy the asynchronous learning engine, that is, the control code generated by the built-in algorithm is written into the asynchronous learning engine, and the asynchronous learning engine can schedule the calculation method of the calculation module. In other words, the asynchronous learning engine can enable the modules required this time, such as selecting different operators in the operator unit to combine different weight calculation formulas, etc., thereby forming corresponding control paths and data paths. In this way, the calculation methods of synaptic weights of different algorithms can be deployed through the asynchronous learning engine, and different input control codes are used for different algorithms to reasonably schedule the internal resources of the asynchronous learning engine to execute the correct operation sequence.

[0078] In step 63, the training sample data is input into the network model in the AI ​​chip so that the computing module updates the network synaptic weights.

[0079] In this step, training sample data of the network model can be obtained. The training sample data can be encoded data such as images, audio and text, and can be set according to specific scenarios. The corresponding solutions fall within the protection scope of this disclosure.

[0080] In this step, the training samples are input into the network model in the AI ​​chip through the bundled data asynchronous circuit based on the Click unit; when the request input end of the Click unit detects the request signal flipping, it outputs a fire pulse signal, and the fire pulse signal is used to control data input and data output.

[0081] In this step, the training sample data can be input into the network model in the AI ​​chip. It is understandable that after the training sample data is input into the network model, each computing module can perform calculations based on the input data and the calculation method to obtain the synaptic weights of this training.

[0082] In step 64, it is determined whether the target synaptic weight has completed training according to the update times of the target synaptic weight in the network model and a preset times threshold.

[0083] In this step, a preset number threshold Nth may be stored, and the preset number threshold is used as a threshold for the number of synapse zero weight updates, that is, an upper limit of the number of times the synaptic weight of the target synapse is updated to zero.

[0084] In this step, after the asynchronous learning engine calculates the synaptic weight (i.e., synaptic update value) of the target synapse during this training process, it is necessary to judge this weight update value. When the weight update value is 0, the update count N is incremented by 1 and the update count N is recorded. When the weight update value is not 0, the weight update value is cleared. Then, the update count of the target synaptic weight in the network model is compared with the above preset count threshold. When the update count is less than or equal to the preset count threshold (i.e., N < Nth or N = Nth), it can be determined that the training is not completed and the process jumps to step 63. When the update count is greater than the preset count threshold (i.e., N > Nth), it is determined that the training is completed and the synaptic weight of the target synapse is obtained, which is hereinafter referred to as the first synaptic weight to distinguish it from the subsequent second synaptic weight. When the synaptic weights of all synapses are completed in training, the training of the network model can be ended. In this way, in this step, the target synapse for which the training is completed is locked by the update count, and the synaptic weight update process can be directly skipped in the subsequent training process, which can reduce the working time of the asynchronous learning engine; moreover, the above target synapse is no longer referenced in the subsequent training process until all synapses are completed in training, which can reduce the data processing volume in the subsequent training process and is beneficial to improving the training efficiency.

[0085] The solution of this embodiment ensures that the AI chip maintains a low-power standby state without external input and has no dynamic power consumption caused by the clock through the event-triggering characteristic of the asynchronous circuit. On the premise of not affecting or slightly reducing the recognition accuracy of the AI chip, it can effectively reduce the computing power consumption during the chip training process.

[0086] After the synaptic weight training of a target synapse is completed, that is, after the first synaptic weight of the target synapse is obtained, it can be stored at a specified location. Considering that directly storing the first synaptic weight will occupy a large storage space, in one embodiment, the lowest data bit of the first synaptic weight can be discarded to obtain the second synaptic weight. For example, for the first synaptic weight w[x - 1:0], the lowest data bit w[0] is discarded this time, and w[x - 1:1] is retained and stored at the specified location. In this way, by reducing the number of bits of the synaptic weight, the storage cost can be reduced.

[0087] In the inference mode of the network model, that is, the usage mode after the network model is trained, the AI ​​chip reads the second synaptic weight w[x-1:1] of each synapse from the specified position, and uses the random quantization generator to perform 1LSB random quantization on the second synaptic weight to obtain the random quantization data w[0]', that is, to generate the lowest data bit for the second synaptic weight. Then, the second synaptic weight w[x-1:1] and the above-mentioned random quantization data w[0]' are combined to obtain the third synaptic weight W'={W[x-1:1], W[0]'}. The difference between the third synaptic weight and the first synaptic weight is that the lowest data bit may be different, and the other data bits are the same. Finally, the weight of the target synapse of the network model can be updated to the third synaptic weight to obtain the target network model. At this time, the target network model can perform inference results based on the input data. In this way, by restoring the lowest data bit of the stored second synaptic weight, the storage cost can be reduced without affecting the use of the target model.

[0088] In combination with the contents of the above embodiments, the present disclosure provides an asynchronous training method for online learning AI chips, see Figure 7 ,include:

[0089] 701, the AI ​​chip is in training mode.

[0090] 702, initializing the network synaptic weights of the network model, and reading the initial values ​​of the weights from the specified positions.

[0091] 703, adjusting the asynchronous learning engine configuration according to the built-in algorithm.

[0092] 704, the training sample data is input into the network model.

[0093] At 705 , the asynchronous learning engine updates the synaptic weight of the target synapse.

[0094] 706 , determining whether the synaptic update weight is 0.

[0095] 707, when the synaptic update weight is not 0, the update times N of the target synapse is changed to 0.

[0096] 708. When the synaptic update weight is 0, increase the update times N of the target synapse by 1.

[0097] 709 , determine whether the update number N is greater than a preset number threshold Nth; if the update number is less than or equal to the preset number threshold, jump to step 704 , otherwise jump to step 710 .

[0098] At 710 , the lowest data bit of the first synaptic weight of the target synapse is discarded to obtain a second synaptic weight, and the second synaptic weight is stored.

[0099] 711 , in the inference mode, perform 1LSB random quantization on the second synaptic weight to obtain a third synaptic weight.

[0100] 712, determine whether all synaptic weights have been stored, if not, jump to step 704, if yes, end this training.

[0101] Based on the asynchronous training method of an online learning AI chip provided in the embodiment of the present disclosure, the embodiment of the present disclosure also provides an asynchronous training device for an online learning AI chip, see Figure 8 , the device comprises:

[0102] A weight initialization module 81 is used to initialize the network synaptic weights of the network model in the AI ​​chip when the AI ​​chip is in training mode;

[0103] An engine configuration module 82, used to deploy an asynchronous learning engine according to a built-in algorithm, wherein the asynchronous learning engine is used to schedule a calculation method of a calculation module;

[0104] A training sample input module 83, used to input training sample data into the network model in the AI ​​chip so that the computing module updates the network synaptic weights;

[0105] The training completion determination module 84 is used to determine whether the target synaptic weight has completed training according to the update times of the target synaptic weight in the network model and a preset times threshold.

[0106] It should be noted that the device shown in this embodiment matches the contents of the method embodiment, and reference may be made to the contents of the above method embodiment, which will not be repeated here.

[0107] In some possible embodiments, a non-transitory computer-readable storage medium is provided, and when an executable computer program in the storage medium is executed by a processor, the asynchronous training method of the online learning AI chip as described above can be implemented.

[0108] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the disclosure disclosed herein. The present disclosure is intended to cover any variations, uses, or adaptations that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art that are not disclosed in the present disclosure. The description and examples are to be considered exemplary only, and the true scope and spirit of the present disclosure are indicated by the following claims.

[0109] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.

Claims

1. An asynchronous training method for online learning AI chip, characterized in that: The method comprises: When the AI ​​chip is in training mode, the network synaptic weights of the network model in the AI ​​chip are initialized; Deploy an asynchronous learning engine according to a built-in algorithm, wherein the asynchronous learning engine is used to schedule a computing mode of a computing module; Inputting training sample data into the network model in the AI ​​chip so that the computing module updates the network synaptic weights; Determining whether the target synaptic weight has completed training according to the update number of the target synaptic weight in the network model and a preset number threshold, and obtaining a first synaptic weight of the target synapse after the training is completed; The method further comprises: The lowest data bit of the first synaptic weight is discarded to obtain a second synaptic weight; and the second synaptic weight is stored in a designated location; When the AI ​​chip is in inference mode, random quantized data is obtained; the second synaptic weight and the random quantized data are combined to obtain a third synaptic weight; the random quantized data is used as the lowest data bit of the third synaptic weight; the weight of the target synapse of the network model is updated to the third synaptic weight to obtain a target network model, and the target network model is used to infer the result based on the input data.

2. The asynchronous training method according to claim 1, characterized in that: Determining whether the target synaptic weight has completed training according to the update times of the target synaptic weight in the network model and a preset times threshold includes: Comparing the update times of the target synaptic weight in the network model with the preset times threshold; When the number of updates is less than or equal to the preset number threshold, determining that the training is not completed and jumping to the step of inputting the training sample data into the network model in the AI ​​chip; When the update number is greater than the preset number threshold, it is determined that the target synapse has completed training and a first synaptic weight of the target synapse is obtained.

3. The asynchronous training method according to claim 1, characterized in that: Get random quantitative data, including: The second synaptic weight is randomly quantized by 1LSB through an asynchronous random number generator, and the asynchronous random number generator works when it detects that the AI ​​chip reads the weight data, otherwise it is in a standby state; The asynchronous random number generator is a multi-stage pipeline composed of Click units and linear feedback shift registers LSFR, which is used to generate different random numbers according to an input seed and use the lowest bit data of the random number as the random quantization data.

4. The asynchronous training method according to claim 1, characterized in that: The method further comprises: The training samples are input into the network model in the AI ​​chip through a bundled data asynchronous circuit based on the Click unit; when the request input end of the Click unit detects that the request signal is flipped, a fire pulse signal is output, and the fire pulse signal is used to control data input and data output.

5. The asynchronous training method according to claim 1, characterized in that: The asynchronous learning engine includes a control module and a calculation module. The control module generates a control code according to preset configuration data, and the control code is used to schedule the calculation mode of the calculation module.

6. The asynchronous training method according to claim 1, characterized in that: The calculation method of the calculation module includes at least one of the following: addition, subtraction, multiplication, division, exponential operation, logarithmic operation, and trigonometric function operation.

7. The asynchronous training method according to claim 6, characterized in that: When the calculation mode of the calculation module is multiplication or division, the calculation module is implemented by a multi-stage pipeline structure consisting of a click unit, a shifter and an adder; or, When the calculation mode of the calculation module is exponential operation or logarithmic operation, the calculation module is implemented by a Click unit and a lookup table LUT; or, When the calculation mode of the calculation module is trigonometric function operation, the calculation module is implemented by a Click unit, an adder, a shifter and a lookup table LUT.

8. An asynchronous training device for online learning AI chip, characterized in that: The device comprises: A weight initialization module, used to initialize the network synaptic weights of the network model in the AI ​​chip when the AI ​​chip is in training mode; An engine configuration module, used to deploy an asynchronous learning engine according to a built-in algorithm, wherein the asynchronous learning engine is used to schedule a calculation method of a calculation module; A training sample input module, used to input training sample data into the network model in the AI ​​chip so that the computing module updates the network synaptic weights; A training completion determination module, used to determine whether the target synaptic weight has completed training according to the number of updates of the target synaptic weight in the network model and a preset number threshold, and obtain the first synaptic weight of the target synapse after the training is completed; The device is also used for: discarding the lowest data bit of the first synaptic weight to obtain a second synaptic weight; and storing the second synaptic weight in a designated location; When the AI ​​chip is in inference mode, random quantized data is obtained; the second synaptic weight and the random quantized data are combined to obtain a third synaptic weight; the random quantized data is used as the lowest data bit of the third synaptic weight; the weight of the target synapse of the network model is updated to the third synaptic weight to obtain a target network model, and the target network model is used to infer the result based on the input data.

Citation Information

Patent Citations

  • Deep reinforcement learning algorithm acceleration circuit and method based on asynchronous circuit

    CN118114730A