Normalization operation circuit and system
By designing a normalized computing circuit including configuration and data cache modules and arithmetic logic modules, the problems of real-time and low power consumption in the edge computing environment in the prior art are solved, and efficient normalized computing is achieved.
Patent Information
- Application Number
- CN202510477314.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-04-16
Smart Images

Figure CN120011729A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of normalization technology, and in particular to a normalization operation circuit and system. Background Art
[0002] With the rapid development of artificial intelligence, various neural network models continue to emerge, and more and more operators are widely used in the calculation process of neural networks. Normalization is an operator of neural networks and a data preprocessing technology. The purpose is to adjust the data to a specific range or distribution to make it more suitable for subsequent calculations or analysis. The core idea of normalization is to scale or transform the data so that data of different dimensions or features are in the same scale range in terms of numerical value.
[0003] Edge computing is a distributed computing model whose core idea is to process data on devices close to the data source or user, rather than sending all data to a remote central server or cloud computing platform for processing. By moving some computing, storage, and data processing work to the "edge" of the network (such as routers, gateways, edge servers, etc.), latency and bandwidth requirements can be reduced, thereby improving real-time performance. Therefore, if normalized operations can be implemented using the edge computing model, system performance can be further improved.
[0004] However, the normalization operation in the related technology is usually implemented using CPU (Central Processing Unit) and GPU (Graphics Processing Unit). The CPU does not have a specific circuit for calculating the normalization operator. Due to its ability to execute programs serially, the normalization operator requires a longer instruction stream for calculation, and its real-time performance cannot meet the requirements of edge computing. The GPU consumes a lot of hardware resources and has high power consumption, resulting in poor energy efficiency, which does not meet the low power consumption and cost requirements of edge computing. Summary of the invention
[0005] The present invention aims to solve one of the technical problems in the related art at least to a certain extent. To this end, the first object of the present invention is to propose a normalization operation circuit to meet the needs of edge computing.
[0006] The second objective of the present invention is to provide a normalization operation system.
[0007] To achieve the above-mentioned purpose, the first aspect of the present invention proposes a normalization operation circuit, which includes: a configuration and data cache module, which is used to determine the data to be operated, the normalization function type and the operation parameters; an arithmetic logic module, which is connected to the configuration and data cache module, and the arithmetic logic module includes multiple operation sub-modules, which are used to determine the target operation sub-module from the multiple operation sub-modules according to the normalization function type, and call the target operation sub-module according to the normalization function type and the operation parameters to perform normalization operation on the data to be operated, wherein the multiple operation sub-modules include an addition sub-module, a multiplication sub-module, a division sub-module and a square root sub-module; the first input end of the addition sub-module is designed to be the arithmetic logic module. The input end of the logic module is used to receive the data to be operated, the output end of the addition submodule is connected to the first input end of the division submodule, the second input end of the division submodule is connected to the first input end of the addition submodule, and the output end of the division submodule is designed to be the output end of the arithmetic logic module; wherein, when the normalization function type is the first preset type, the data to be operated is a vector, the target operation submodule includes the addition submodule and the division submodule, the addition submodule is used to sum the absolute values of all elements of the data to be operated to obtain a first summation result, and the division submodule is used to divide the data to be operated by the first summation result to obtain a normalized operation result of the data to be operated.
[0008] In addition, the normalization operation circuit according to the embodiment of the present invention may also have the following additional technical features: In one embodiment of the present invention, the addition submodule includes: a first parameter output device, an inverter, a first multiplexer, an adder, a first register, a second multiplexer, and a second register; wherein the input end of the inverter is designed to be the first input end of the addition submodule; the first input end of the first multiplexer is connected to the output end of the inverter, and the second input end of the first multiplexer is connected to the input end of the inverter; the first input end of the adder is connected to the output end of the first multiplexer; the input end of the first register is connected to the output end of the adder; the first input end of the second multiplexer is connected to the output end of the first parameter output device, the second input end of the second multiplexer is connected to the output end of the first register, and the second input end of the second multiplexer is connected to the output end of the first register. The output end is connected to the second input end of the adder; the input end of the second register is connected to the output end of the first register, and the output end of the second register is connected to the input end of the inverter; the output end of the first register is also set as the output end of the addition submodule; the division submodule includes: a divider, a left shifter, a third register, and a right shifter; wherein the first input end of the divider is designed as the first input end of the division submodule, and the second input end of the divider is connected to the output end of the third register; the input end of the left shifter is designed as the second input end of the division submodule; the input end of the third register is connected to the output end of the left shifter; the input end of the right shifter is connected to the output end of the divider, and the output end of the right shifter is designed as the output end of the division submodule.
[0009] In one embodiment of the present invention, when the normalization function type is the first preset type, for each element in the data to be operated, when the element is input into the addition submodule, the first multiplexer outputs the element to the adder when the element is a positive number, and outputs the element inverted by the inverter to the adder when the element is a negative number, and the adder adds the output data of the first multiplexer to the operation parameter output by the first parameter output device to obtain the first added data; when no data is stored in the second register and there are elements that are not input into the data to be operated, the first added data is stored in the second register via the first register; when no data is stored in the second register and all elements of the data to be operated have been input into the addition submodule, it is determined that the first added data is the first added data. a summation result; when data has been stored in the second register, the first added data is output to the adder via the first register and the second multiplexer, and the first stored data in the second register is output to the adder via the first multiplexer, the adder adds the first added data to the first stored data to obtain second added data; if there are elements that are not input into the data to be operated, the second added data is stored as new first stored data into the second register; if all elements of the data to be operated have been input into the addition submodule, it is determined that the second added data is the first summation result; the left shifter left-shifts the data to be operated, the divider divides the left-shifted data to be operated by the first summation result, and the right shifter right-shifts the output data of the divider to obtain the normalized operation result of the data to be operated.
[0010] In one embodiment of the present invention, the input end of the second register is designed as the second input end of the addition submodule, the first input end and the second input end of the multiplication submodule are both connected to the first input end of the addition submodule, the output end of the multiplication submodule is connected to the second input end of the addition submodule, the input end of the square root submodule is connected to the output end of the addition submodule, and the output end of the square root submodule is connected to the first input end of the division submodule; the multiplication submodule includes: a second parameter output device, a third multiplexer, and a multiplier; wherein the third multiplexer The first input end is connected to the output end of the second parameter output device, and the second input end of the third multiplexer is designed as the second input end of the multiplication sub-module; the first input end of the multiplier is designed as the first input end of the multiplication sub-module, the second input end of the multiplier is connected to the output end of the third multiplexer, and the output end of the multiplier is designed as the output end of the multiplication sub-module; the square root sub-module includes a square root generator, the input end of the square root generator is designed as the input end of the square root sub-module, and the output end of the square root generator is designed as the output end of the square root sub-module.
[0011] In one embodiment of the present invention, when the normalization function type is the second preset type, the data to be operated is a vector, and the target operation submodule includes the addition submodule, the multiplication submodule, the division submodule and the square root submodule; wherein, for each element in the data to be operated, the element is input to the first input terminal and the second input terminal of the multiplier, so that the multiplier multiplies the element with itself to obtain a first multiplication result, and stores the first multiplication result to the second register; if no data is stored in the first register, the second register outputs the first multiplication result to the adder through the first multiplexer, and the adder adds the first multiplication result to the operation parameter output by the first parameter output device, and adds the first multiplication result to the operation parameter output by the first parameter output device, and stores the result to the second register; The addition result is stored in the first register; if data has been stored in the first register, the first register outputs the data stored therein to the adder through the second multiplexer, and the adder adds the data stored in the first register and the first multiplication result, and stores the addition result in the first register; after the addition result corresponding to each element in the data to be calculated is stored in the first register, the data in the first register is used as the third addition result, and the third addition result is output to the square root generator, the square root generator takes the square root of the third addition result to obtain a square root result, and outputs the square root result to the divider, and the divider divides the data to be calculated by the square root result to obtain a normalized calculation result of the data to be calculated.
[0012] In one embodiment of the present invention, the arithmetic logic module further includes a fourth multiplexer, a first input end of the fourth multiplexer is connected to the output end of the right shifter, an output end of the fourth multiplexer is connected to the output end of the first register, and the output end of the fourth multiplexer is designed to be the output end of the arithmetic logic module.
[0013] In one embodiment of the present invention, when the normalization function type is the third preset type, the target operator module includes the addition submodule and the multiplication submodule; wherein the third multiplexer outputs the operation parameter output by the second parameter output device to the multiplier, and after receiving the data to be operated, the multiplier multiplies the data to be operated with the operation parameter output by the second parameter output device to obtain a second multiplication result, and outputs the second multiplication result to the adder through the second register and the first multiplexer, and the adder adds the second multiplication result with the operation parameter output by the first parameter output device to obtain a normalized operation result of the data to be operated.
[0014] In one embodiment of the present invention, the input end of the square root submodule is also connected to the input end of the addition submodule and the second input end of the division submodule. When the normalization function type is the fourth preset type, the data to be operated is a vector, and the target operation submodule includes the addition submodule, the multiplication submodule, the division submodule and the square root submodule; wherein the addition submodule is used to sum all elements of the data to be operated to obtain a fourth addition result, and the division submodule is used to divide the fourth addition result by the total number of elements in the data to be operated to obtain a first initial operation result; the multiplication submodule is used to multiply all elements of the data to be operated with themselves to obtain multiple third multiplication results, and the addition submodule is also used to add all the third multiplication results to obtain a fifth multiplication result. The division submodule is used to divide the fifth addition result by the total number of elements in the data to be calculated to obtain a second initial calculation result; the square root submodule is used to take the square root of the second initial calculation result to obtain a square root result; the addition submodule is used to invert the first initial calculation result and add it to each element in the data to be calculated to obtain a sixth addition result; the division submodule is used to divide the sixth addition result by the square root result to obtain a third initial calculation result; the multiplication submodule is used to multiply the third initial calculation result by the calculation parameter output by the second parameter output device to obtain a fourth multiplication result; the addition submodule is used to add the fourth multiplication result to the calculation parameter output by the first parameter output device to obtain a normalized calculation result of the data to be calculated.
[0015] To achieve the above-mentioned purpose, a second embodiment of the present invention proposes a normalization operation system, including the above-mentioned normalization operation circuit.
[0016] According to the normalization operation circuit and system of the embodiment of the present invention, it includes: a configuration and data cache module, which is used to determine the data to be operated, the normalization function type and the operation parameters; an arithmetic logic module, which is connected to the configuration and data cache module, and the arithmetic logic module includes multiple operation submodules, which are used to determine the target operation submodule from multiple operation submodules according to the normalization function type, and call the target operation submodule according to the normalization function type and the operation parameters to perform normalization operation on the data to be operated, wherein the multiple operation submodules include an addition submodule, a multiplication submodule, a division submodule and a square root submodule. Therefore, by using an additionally set arithmetic logic module to perform normalization operation, there is no need for device parameters such as CPU and GPU, thereby meeting the needs of edge computing.
[0017] Additional aspects and advantages of the present invention will be given in part in the following description and in part will be obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 is a structural block diagram of a normalization operation circuit according to an embodiment of the present invention; Figure 2 is a working schematic diagram of a normalization operation circuit according to an embodiment of the present invention; Figure 3 is a circuit diagram of a normalization operation circuit of a specific embodiment of the present invention; Figure 4 It is a structural block diagram of a normalization operation system according to an embodiment of the present invention. DETAILED DESCRIPTION
[0019] The normalization operation circuit and system of the embodiment of the present invention are described below with reference to the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements with the same or similar functions. The embodiments described with reference to the accompanying drawings are exemplary and should not be construed as limiting the present invention.
[0020] Figure 1 4 is a structural block diagram of a normalization operation circuit according to an embodiment of the present invention.
[0021] like Figure 1As shown, the normalization operation circuit 100 includes: a configuration and data cache module 200, which is used to determine the data to be operated, the normalization function type and the operation parameters; an arithmetic logic module 300, which is connected to the configuration and data cache module 200, and the arithmetic logic module 300 includes multiple operation sub-modules, which are used to determine the target operation sub-module from the multiple operation sub-modules according to the normalization function type, and call the target operation sub-module according to the normalization function type and the operation parameters to perform a normalization operation on the data to be operated, wherein the multiple operation sub-modules include an addition sub-module, a multiplication sub-module, a division sub-module and a square root sub-module.
[0022] Common normalization methods include L1 normalization, L2 normalization, batch normalization, and layer normalization. There are three reasons for using normalization: (1) It can avoid differences in the magnitude of different features. In a data set, the numerical ranges of different features may vary greatly. If normalization is not performed, this feature may dominate the calculation process and affect the learning effect of the model. (2) It can speed up convergence. In machine learning, normalization can speed up model training. Especially when using the stochastic gradient descent method, the normalized data will make the update of model parameters more stable and avoid some features being updated too quickly or too slowly. (3) Algorithms such as the Single Shot Multi-Box Detector algorithm are sensitive to the relative distance between features. Normalization can make the weights between different features more balanced and improve the accuracy of the model.
[0023] Therefore, in order to realize normalization operation based on multiple normalization methods and meet the edge computing requirements, the normalization operation circuit 100 as described above is set.
[0024] The normalization operation circuit 100 includes a configuration and data cache module 200 and an arithmetic logic module 300. The configuration and data cache module 200 is used to determine the data to be operated, the normalization function type and the operation parameters. The normalization function type includes L1 normalization, L2 normalization, batch normalization and layer normalization.
[0025] The above-mentioned arithmetic logic module 300 is a specific operation module, including an addition sub-module, a multiplication sub-module, a division sub-module and a square root sub-module. After the configuration and data cache module 200 determines the normalization function type and outputs the operation parameters, the arithmetic logic module 300 can call the operation sub-module according to the determined normalization function type and perform normalization operation on the operation data based on the operation parameters.
[0026] That is, in order to achieve normalization in various forms and meet the requirements of edge computing, a configuration and data cache module 200 and an arithmetic logic module 300 are provided, see Figure 2The configuration and data cache module 200 is used to receive the shaped data and commands. The shaped data includes the quantized shaped data of the neural network layer, and the shaped data is the data to be operated. The command is a command including the type, size, and parameters of the neural network layer. After receiving the shaped data and commands, the configuration and data cache module 200 converts the shaped data into a form that can be processed by the arithmetic logic module 300, sends it to the arithmetic logic module 300, and determines which normalization to be performed according to the command, and generates the operation parameters corresponding to the normalization to be performed, and sends the specific normalization function type to be performed and the corresponding operation parameters to the arithmetic logic module 300. The configuration and data cache module 200 can be integrated in the circuit using a digital logic circuit.
[0027] After receiving the data to be operated, the normalization function type and the operation parameters, the arithmetic logic module 300 uses the operation submodules contained therein to perform normalization operation on the data to be operated to obtain a normalization operation result. Figure 2 The result in is the normalized operation result. Moreover, since the arithmetic logic module 300 uses the multiple operation submodules contained therein to perform normalized operations, the control of the arithmetic logic module 300 only needs to control which specific operation submodules work and how each operation submodule works, so that the control of the arithmetic logic module 300 can be implemented using digital logic circuits without the participation of devices such as CPU and GPU, thereby meeting the needs of edge computing.
[0028] In some embodiments of the present invention, the first input end of the addition submodule is designed as the input end of the arithmetic logic module 300, which is used to receive the data to be operated, the output end of the addition submodule is connected to the first input end of the division submodule, the second input end of the division submodule is connected to the first input end of the addition submodule, and the output end of the division submodule is designed as the output end of the arithmetic logic module 300; wherein, when the normalization function type is the first preset type, the data to be operated is a vector, the target operation submodule includes an addition submodule and a division submodule, the addition submodule is used to sum the absolute values of all elements of the data to be operated to obtain a first summation result, and the division submodule is used to divide the data to be operated by the first summation result to obtain a normalized operation result of the data to be operated.
[0029] In some embodiments of the present invention, the adding submodule includes: a first parameter output device, an inverter, a first multiplexer, an adder, a first register, a second multiplexer, and a second register; wherein the input end of the inverter is designed as the first input end of the adding submodule; the first input end of the first multiplexer is connected to the output end of the inverter, and the second input end of the first multiplexer is connected to the input end of the inverter; the first input end of the adder is connected to the output end of the first multiplexer; the input end of the first register is connected to the output end of the adder; the first input end of the second multiplexer is connected to the output end of the first parameter output device, the second input end of the second multiplexer is connected to the output end of the first register, and the output end of the second multiplexer is connected to the second input end of the adder; the input end of the second register is connected to the output end of the first register, and the output end of the second register is connected to the input end of the inverter; the output end of the first register is also set as the output end of the adding submodule; The division submodule comprises: a divider, a left shifter, a third register, and a right shifter; wherein the first input end of the divider is designed as the first input end of the division submodule, and the second input end of the divider is connected to the output end of the third register; the input end of the left shifter is designed as the second input end of the division submodule; the input end of the third register is connected to the output end of the left shifter; the input end of the right shifter is connected to the output end of the divider, and the output end of the right shifter is designed as the output end of the division submodule.
[0030] In some embodiments of the present invention, when the normalization function type is the first preset type, for each element in the data to be operated, when the element is input into the addition submodule, the first multiplexer outputs the element to the adder when the element is a positive number, and outputs the element after the inverter is inverted to the adder when the element is a negative number, and the adder adds the output data of the first multiplexer to the operation parameter output by the first parameter output device to obtain the first added data; when no data is stored in the second register and there are elements in which the data to be operated is not input, the first added data is stored in the second register via the first register; when no data is stored in the second register, When all elements of the data to be operated have been input into the addition submodule, the first added data is determined to be the first summation result; when data has been stored in the second register, the first added data is output to the adder through the first register and the second multiplexer, and the first stored data in the second register is output to the adder through the first multiplexer, and the adder adds the first added data to the first stored data to obtain the second added data; if there are elements that have not been input into the data to be operated, the second added data is stored in the second register as the new first stored data; if all elements of the data to be operated have been input into the addition submodule, the second added data is determined to be the first summation result; The left shifter shifts the data to be calculated leftwards, the divider divides the left-shifted data to be calculated by the first summation result, and the right shifter shifts the output data of the divider rightwards to obtain a normalized calculation result of the data to be calculated.
[0031] In some embodiments of the present invention, the input end of the second register is designed as the second input end of the addition submodule, the first input end and the second input end of the multiplication submodule are both connected to the first input end of the addition submodule, the output end of the multiplication submodule is connected to the second input end of the addition submodule, the input end of the square root submodule is connected to the output end of the addition submodule, and the output end of the square root submodule is connected to the first input end of the division submodule; A multiplication submodule, comprising: a second parameter output device, a third multiplexer, and a multiplier; wherein a first input end of the third multiplexer is connected to an output end of the second parameter output device, and a second input end of the third multiplexer is designed as a second input end of the multiplication submodule; a first input end of the multiplier is designed as a first input end of the multiplication submodule, a second input end of the multiplier is connected to an output end of the third multiplexer, and an output end of the multiplier is designed as an output end of the multiplication submodule; The square root submodule comprises a square root generator, the input end of the square root generator is designed as the input end of the square root submodule, and the output end of the square root generator is designed as the output end of the square root submodule.
[0032] In some embodiments of the present invention, when the normalization function type is the second preset type, the data to be operated is a vector, and the target operator module includes an addition submodule, a multiplication submodule, a division submodule and a square root submodule; Wherein, for each element in the data to be operated, the element is input to the first input terminal and the second input terminal of the multiplier, so that the multiplier multiplies the element with itself to obtain a first multiplication result, and the first multiplication result is stored in the second register; if the first register does not store data, the second register outputs the first multiplication result to the adder through the first multiplexer, the adder adds the first multiplication result and the operation parameter output by the first parameter output device, and stores the addition result in the first register; if the first register has data stored, the first register outputs the data stored therein to the adder through the second multiplexer, the adder adds the data stored in the first register and the first multiplication result, and stores the addition result in the first register; After the addition results corresponding to each element in the data to be calculated are stored in the first register, the data in the first register is used as the third addition result, and the third addition result is output to the square root generator. The square root generator takes the square root of the third addition result to obtain a square root result, and the square root result is output to the divider. The divider divides the data to be calculated by the square root result to obtain a normalized calculation result of the data to be calculated.
[0033] In some embodiments of the present invention, the arithmetic logic module 300 further includes a fourth multiplexer, the first input end of the fourth multiplexer is connected to the output end of the right shifter, the output end of the fourth multiplexer is connected to the output end of the first register, and the output end of the fourth multiplexer is designed to be the output end of the arithmetic logic module 300.
[0034] In some embodiments of the present invention, when the normalization function type is a third preset type, the target operator module includes an addition submodule and a multiplication submodule; Among them, the third multiplexer outputs the operation parameter output by the second parameter output device to the multiplier. After receiving the data to be operated, the multiplier multiplies the data to be operated with the operation parameter output by the second parameter output device to obtain a second multiplication result, and outputs the second multiplication result to the adder through the second register and the first multiplexer. The adder adds the second multiplication result and the operation parameter output by the first parameter output device to obtain a normalized operation result of the data to be operated.
[0035] In some embodiments of the present invention, the input end of the square root submodule is also connected to the input end of the addition submodule and the second input end of the division submodule. When the normalization function type is the fourth preset type, the data to be operated is a vector, and the target operation submodule includes an addition submodule, a multiplication submodule, a division submodule and a square root submodule. The addition submodule is used to sum all elements of the data to be operated to obtain a fourth addition result, and the division submodule is used to divide the fourth addition result by the total number of elements in the data to be operated to obtain a first initial operation result; the multiplication submodule is used to multiply all elements of the data to be operated by themselves to obtain multiple third multiplication results, and the addition submodule is also used to add all the third multiplication results to obtain a fifth addition result, and the division submodule is also used to divide the fifth addition result by the total number of elements in the data to be operated to obtain a second initial operation result; the square root submodule is used to multiply the second The initial operation result is squared to obtain a square root result. The addition submodule is also used to invert the first initial operation result and add it to each element in the data to be operated to obtain a sixth addition result. The division submodule is also used to divide the sixth addition result by the square root result to obtain a third initial operation result. The multiplication submodule is also used to multiply the third initial operation result with the operation parameter output by the second parameter output device to obtain a fourth multiplication result. The addition submodule is also used to add the fourth multiplication result with the operation parameter output by the first parameter output device to obtain a normalized operation result of the data to be operated.
[0036] Combine the following Figure 3 The specific embodiment shown is used to illustrate the arithmetic logic module 300 .
[0037] exist Figure 3In the specific embodiment shown, the control of the arithmetic logic module 300 is implemented by a multiplexer in the module, and the multiplexer and register are controlled by a digital logic circuit.
[0038] exist Figure 3 301 is a first parameter output device, 302 is an inverter, 303 is a first multiplexer, 304 is an adder, 305 is a first register, 306 is a second multiplexer, 307 is a second register, 308 is a divider, 309 is a left shifter, 310 is a third register, 311 is a right shifter, 312 is a second parameter output device, 313 is a third multiplexer, 314 is a multiplier, 315 is a square root extractor, 316 is a fourth multiplexer, 318 is a fifth multiplexer, 319 is a sixth multiplexer, 320 is a fourth register, the input is the input end of the arithmetic logic module 300, and the output is the output end of the arithmetic logic module 300. The adder 304 is an integer adder, the divider 308 is an integer divider, the multiplier 314 is an integer multiplier, and the square root extractor 315 is an integer square root extractor.
[0039] The first preset type is L1 normalization, the second preset type is L2 normalization, the third preset type is batch normalization, and the fourth preset type is layer normalization.
[0040] Specifically, when the normalization function type is the first preset type, the data to be operated is a vector.
[0041] The fourth multiplexer 316 controls the conduction of the path from the right shifter 311 to the output end, and the fifth multiplexer 318 controls the conduction of the path from the input end to the input end of the left shifter 309 .
[0042] The configuration and data cache module 200 inputs the elements of the data to be operated one by one. Assuming that the data to be operated is a vector , the configuration and data cache module 200 will first input the element through the input terminal , in the arithmetic logic module 300 pairs of elements After the operation is completed, the configuration and data cache module 200 inputs the element , in the arithmetic logic module 300 pairs of elements After the operation is completed, the configuration and data cache module 200 inputs the element , and the process repeats until all elements of vector X are input through the input terminal.
[0043] The target operator module includes an addition submodule and a division submodule. At this time, the multiplier 314 and the square root extractor 315 can be controlled not to work. For example, the power supply to the multiplier 314 and the square root extractor 315 can be stopped. For example, a multiplication register can also be set. The input end of the multiplication register is designed to be the first input end of the multiplication submodule, and the output end of the multiplication register is connected to the first input end of the multiplier 314. When the normalization function type is the first preset type, the multiplication register directly deletes the data after receiving the data. For example, a control signal can be sent to the control end of the multiplier 314 and the square root extractor 315 to stop the multiplier 314 and the square root extractor 315 from working.
[0044] After the configuration and data cache module 200 inputs an element through the input terminal, it first determines whether the first digit of the element is 0 or 1. If the first digit of the element is 1, that is, the element is a negative number, the first multiplexer 303 controls the path from the output terminal of the inverter 302 to the first input terminal of the adder 304 to be connected, and the output parameter of the first parameter output device 301 is 1; if the first digit of the element is 0, that is, the element is a positive number, the first multiplexer 303 controls the path from the input terminal of the inverter 302 to the first input terminal of the adder 304 to be connected, and the output parameter of the first parameter output device 301 is 0. In addition, the element is also input into the third register 310.
[0045] Assuming that the data to be operated is a vector [3, -5, 7], the configuration and data cache module 200 will first input the data 00000011 (binary 3) through the input terminal. At this time, the second multiplexer 306 controls the path from the first parameter output device 301 to the second input terminal of the adder 304 to be turned on. Since the first bit of 00000011 is 0, the first multiplexer 303 controls the path from the input terminal of the inverter 302 to the first input terminal of the adder 304 to be turned on. The first parameter is 0, and 00000011 is output to the first terminal of the adder 304. The second multiplexer 306 outputs 0 to the second terminal of the adder 304. The adder 304 adds 0 and 00000011 to obtain 00000011, and stores 00000011 to the first register 305, and then stores it to the second register 307 from the first register 305.
[0046] At the same time, 00000011 is input to the left shifter 309 via the fifth multiplexer 318 . The left shifter 309 left-shifts 00000011 and outputs the left-shift result to the third register 310 . The third register 310 stores the data output by the left shifter 309 .
[0047] Afterwards, the configuration and data cache module 200 inputs the data 11111011 (binary -5) through the input terminal. At this time, the second multiplexer 306 controls the path from the first parameter output device 301 to the second input terminal of the adder 304 to be turned on. Since the first digit of 11111011 is 1, the first multiplexer 303 controls the path from the output terminal of the inverter 302 to the first input terminal of the adder 304 to be turned on. The first parameter is 1, that is, after the configuration and data cache module 200 inputs the data 11111011 through the input terminal, 11111011 will first be inverted by the inverter 302 to obtain 00000100. The adder 304 adds 00000100 and 1 to obtain 00000101, which is 5 in binary, thereby converting the element -5 to 5. Afterwards, the adder 304 outputs 00000101 to the first register 305, and the second multiplexer 306 controls the path from the output end of the first register 305 to the second input end of the adder 304. Further, the second register 307 outputs 00000011 stored therein to the first input end of the adder 304, and the first register 305 outputs 00000101 stored therein to the first input end of the adder 304. The adder 304 adds 00000011 and 00000101 to obtain 00001000, which is 8 in binary. The adder 304 outputs 00001000 to the first register 305, and the first register 305 outputs 00001000 to the second register 307.
[0048] At the same time, the left shifter 309 shifts the above 11111011 to the left and outputs it to the third register 310 after the left shift.
[0049] Finally, the configuration and data cache module 200 inputs the data 00000111 (binary 7) through the input terminal. At this time, the second multiplexer 306 controls the path from the first parameter output device 301 to the second input terminal of the adder 304 to be turned on. Since the first bit of 00000111 is 0, the first multiplexer 303 controls the path from the input terminal of the inverter 302 to the first input terminal of the adder 304 to be turned on. The first parameter is 0, that is, after the configuration and data cache module 200 inputs the data 00000111 through the input terminal, the adder 304 adds 00000111 to 0 to obtain 00000111. After that, the adder 304 outputs 00000111 to the first register 305, and the second multiplexer 306 controls the path from the output terminal of the first register 305 to the second input terminal of the adder 304 to be turned on. Further, the second register 307 outputs 1000 stored therein to the first input terminal of the adder 304, and the first register 305 outputs 00000111 stored therein to the first input terminal of the adder 304. The adder 304 adds 00001000 and 00000111 to obtain 00001111, which is 15 in binary.
[0050] At the same time, the left shifter 309 shifts the above 00000111 left and outputs it to the third register 310 after the left shift.
[0051] Since at this time, all elements of the data to be calculated have been input into the addition submodule, it is determined that 00001111 currently stored in the first register 305 is the first summation result. Thus, it can be achieved that for each element in the data to be calculated, when the element is input into the addition submodule, when the element is a positive number, the element is output to the adder, and when the element is a negative number, the inverted element is output to the adder 304, and the adder 304 adds the output data of the first multiplexer 303 and the operation parameter output by the first parameter output device 301 to obtain the first added data; when no data is stored in the second register 307, and there are elements that are not input into the data to be calculated, the first added data is stored in the second register 307 via the first register 305; when no data is stored in the second register 307, and all elements of the data to be calculated have been input into the addition submodule, the first added data is stored in the second register 307 via the first register 305; when no data is stored in the second register 307, and all elements of the data to be calculated have been input into the addition submodule, the first added data is stored in the second register 307 via the first register 305; module, determine that the first added data is the first summation result; when data has been stored in the second register 307, output the first added data to the adder 304 via the first register 305 and the second multiplexer 306, and output the first stored data in the second register 307 to the adder 304 via the first multiplexer 303, and the adder 304 adds the first added data to the first stored data to obtain the second added data; if there are elements that have not been input into the data to be calculated, store the second added data as the new first stored data in the second register 307; if all elements of the data to be calculated have been input into the addition submodule, determine that the second added data is the first summation result.
[0052] The first register 305 outputs 00001111, and the third register 310 sequentially outputs the three data stored therein. The divider 308 sequentially divides the three data output by the third register 310 by the first summation result output by the first register 305, that is, first divides 3 by the first summation result, then divides -5 by the first summation result, and then divides 7 by the first summation result. The right shifter 311 shifts the data output by the divider 308 backward to obtain the normalized operation results of 3, -5, and 7.
[0053] Assuming that the data to be calculated is a one-dimensional vector [-6], the configuration and data cache module 200 first inputs the data 11111010 (binary -6) through the input terminal. At this time, the second multiplexer 306 controls the path from the first parameter output device 301 to the second input end of the adder 304 to be turned on. Since the first digit of 11111010 is 1, the first multiplexer 303 controls the path from the output end of the inverter 302 to the first input end of the adder 304 to be turned on. The first parameter is 1. The inverter 302 inverts 11111010 to obtain 00000101, and outputs 00000101 to the first end of the adder 304. The second multiplexer 306 outputs 1 to the second end of the adder 304. The adder 304 adds 1 and 00000101 to obtain 00000110, which is 6 in binary, and stores 00000110 in the first register 305.
[0054] At the same time, 1010 is input to the left shifter 309 via the fifth multiplexer 318 . The left shifter 309 left-shifts 11111010 and outputs the left-shift result to the third register 310 . The third register 310 stores the data output by the left shifter 309 .
[0055] Since at this time, all elements of the data to be operated have been input into the addition submodule, it is determined that 00000110 currently stored in the first register 305 is the first summation result.
[0056] The first register 305 outputs 00000110, and the third register 310 outputs the data stored therein, that is, the binary form of the above-mentioned vector [-6]. The divider 308 divides each element in the vector [-6] by the first summation result output by the first register 305, and shifts the output data of the divider 308 backward to obtain the normalized operation result of the vector [-6].
[0057] In this way, L1 normalization can be performed on the data to be operated. Assume that the data to be operated is =[ , , , …, ], and use the elements For example, the specific calculation formula can be shown as follows: , in, For the elements Perform L1 normalization, , , , …, is the element in the data to be calculated X, N is the number of elements in the data to be calculated X, = , which is the sum of the absolute values of all elements in the operation data.
[0058] When the normalization function type is the second preset type, the data to be operated is a vector.
[0059] The third multiplexer 313 controls the path connection from the input end to the second input end of the multiplier 314, the fourth multiplexer 316 controls the path connection from the right shifter 311 to the output end, and the fifth multiplexer 318 controls the path connection from the input end to the input end of the left shifter 309.
[0060] The configuration and data cache module 200 inputs the elements of the data to be operated one by one. Assuming that the data to be operated is a vector , the configuration and data cache module 200 will first input the element through the input terminal , in the arithmetic logic module 300 pairs of elements After the operation is completed, the configuration and data cache module 200 inputs the element , in the arithmetic logic module 300 pairs of elements After the operation is completed, the configuration and data cache module 200 inputs the element , and the process repeats until all elements of vector X are input through the input terminal.
[0061] The target operator submodule includes an addition submodule, a multiplication submodule, a division submodule and a square root submodule.
[0062] The first multiplexer 303 controls the path from the input terminal of the inverter 302 to the first input terminal of the adder 304 to be conductive.
[0063] Assuming that the above data to be operated is a vector [3, -5, 7], the configuration and data cache module 200 will first input 3 in binary form through the input terminal. At this time, the adder 304 ignores the data. For example, the adder 304 can be set not to work at this time. For another example, an addition register can be set, the input end of the addition register is connected to the output end of the first multiplexer 303, and the output end of the addition register is connected to the first input end of the adder 304, and the addition register is set to delete the data. For another example, a control signal can be sent to the adder 304 to stop the adder 304 before the multiplier 314 outputs a signal and after the calculated data is sent to the first register 305. 3 will be input to the first input terminal and the second input terminal of the multiplier 314, so that the multiplier 314 multiplies 3 by 3 to obtain 9, and outputs 9 to the second register 307. The second multiplexer 306 controls the path from the first parameter output device 301 to the adder 304 , the first parameter is 0, and the second register 307 outputs 9. The adder 304 adds 9 and 0 to obtain 9, and stores 9 in the first register 305 .
[0064] At the same time, the binary form of 3 is input to the left shifter 309 via the fifth multiplexer 318 . The left shifter 309 left-shifts the data and outputs the left-shift result to the third register 310 . The third register 310 stores the data output by the left shifter 309 .
[0065] Further, the configuration and data cache module 200 inputs the data -5 in binary form through the input terminal. At this time, the adder 304 ignores the data, and -5 is input to the first input terminal and the second input terminal of the multiplier 314, so that the multiplier 314 multiplies -5 by -5 to obtain 25, and outputs 25 to the second register 307. The second multiplexer 306 controls the conduction of the path from the output terminal of the first register 305 to the second input terminal of the adder 304, the first register 305 outputs 9 to the second input terminal of the adder 304, the second register 307 outputs 25 to the first input terminal of the adder 304, the adder 304 adds 25 and 9 to obtain 34, and stores 34 to the first register 305.
[0066] At the same time, -5 in binary form is input to the left shifter 309 , which left-shifts the data and outputs the left-shift result to the third register 310 , which stores the data output by the left shifter 309 .
[0067] Further, the configuration and data cache module 200 inputs the data 7 in binary form through the input terminal. At this time, the adder 304 ignores the data, and 7 is input to the first input terminal and the second input terminal of the multiplier 314, so that the multiplier 314 multiplies 7 by 7 to obtain 49, and outputs 49 to the second register 307. The second multiplexer 306 controls the conduction of the path from the output terminal of the first register 305 to the second input terminal of the adder 304, the first register 305 outputs 34 to the second input terminal of the adder 304, the second register 307 outputs 49 to the first input terminal of the adder 304, and the adder 304 adds 34 and 49 to obtain 83, and stores 83 to the first register 305.
[0068] At the same time, the binary form of 7 is input to the left shifter 309 , which shifts the data left and outputs the left shift result to the third register 310 , which stores the data output by the left shifter 309 .
[0069] At this time, since the addition results corresponding to all elements of the data to be operated have been stored in the first register 305, it is determined that the binary form of 83 currently stored in the first register 305 is the third addition result.
[0070] Furthermore, the first register 305 outputs the third addition result to the input end of the square root generator 315 , and the square root generator 315 performs square root operation on the third addition result to obtain a square root result.
[0071] Furthermore, the square root generator 315 outputs the square root result to the first input terminal of the divider 308, the third register 310 outputs the three data stored therein to the second input terminal of the divider 308 in sequence, and the divider 308 divides the data output by the third register 310 by the square root result in sequence to obtain a normalized operation result of the data to be operated.
[0072] In this way, L2 normalization can be performed on the data to be operated. Assume that the data to be operated is =[ , , , …, ], and use the elements For example, the specific calculation formula can be shown as follows: , in, For the elements Perform L2 normalization, , , , …, is the element in the data to be calculated X, N is the number of elements in the data to be calculated X, = , which is to sum the squares of all elements in the operation data and then take the square root.
[0073] When the normalization function type is the third preset type, the batch normalization function includes BatchNorm1d, BatchNorm2d, BatchNorm3d and other functions. BatchNorm1d processes one-dimensional data and is suitable for fully connected layers or sequence data. The size of the input data is (batch size N, number of channels C) or (batch size N, number of channels C, length L). BatchNorm2d processes two-dimensional data and is usually used for convolutional layers. The size of the input data is (batch size N, number of channels C, length H, width W). BatchNorm3d processes three-dimensional data and is suitable for three-dimensional convolutional layers. The size of the input data is (batch size N, number of channels C, depth D, height H, width W). The output and input shapes are consistent. For hardware, different functions have different amounts of input and output data each time, and the configuration and data cache module 200 needs to be used to configure the parameters, but the calculation algorithm for the data is the same. Assuming that the data to be calculated is x at this time, the formulas of each function can be unified into the following formula: .
[0074] is the batch normalization function, E1(x), Var1(x), , are the parameters in this formula, which can be generated by training and are known. is a minimum value, preventing division by zero, so it can be converted into the following formula: .
[0075] k and b are parameters in this formula.
[0076] At this time, the third multiplexer 313 controls the path connection from the second parameter output device 312 to the second input terminal of the multiplier 314, the first multiplexer 303 controls the path connection from the input terminal of the inverter 302 to the first input terminal of the adder 304, and the second multiplexer 306 controls the path connection from the first parameter output device 301 to the second input terminal of the adder 304.
[0077] The target operator module includes an addition module and a multiplication module. At this time, the division module and the square root submodule stop working. For example, the divider 308 and the square root extractor 315 can be powered off. For another example, the square root register can be set, and the square root register and the third register 310 can be controlled to directly delete the data when receiving the data. For another example, a control signal can be sent to the divider 308 and the square root extractor 315 to stop the divider 308 and the square root extractor 315.
[0078] The configuration and data cache module 200 sends the data to be calculated through the input terminal. At this time, the control adder 304 also stops working. For the specific method, please refer to the above method. Assuming that the data to be calculated is the above x, the second parameter output device 312 outputs the calculation parameter k, and the first parameter output device 301 outputs the calculation parameter b. The multiplier 314 receives x through its first input terminal and receives k through its second input terminal, calculates x multiplied by k, and uses the calculation result as the second multiplication result, outputs the second multiplication result to the second register 307, and the second register 307 outputs the second multiplication result to the adder 304. The adder 304 adds the second multiplication result to b to obtain the normalized calculation result of the parameter to be calculated, and controls the fourth multiplexer 316 to connect the path from the output terminal of the first register 305 to the output terminal, and outputs the normalized calculation result.
[0079] When the normalization function type is the fourth preset type, the data to be operated is a vector, and the target operation submodule includes an addition submodule, a multiplication submodule, a division submodule and a square root submodule.
[0080] Assume that the above input data is a vector [3, -5, 7].
[0081] In the first step, the addition submodule is used to sum all elements of the data to be calculated to obtain a fourth addition result, and the division submodule is used to divide the fourth addition result by the total number of elements in the data to be calculated to obtain a first initial calculation result.
[0082] Specifically, the configuration and data cache module 200 first inputs the binary data 3 through the input terminal. At this time, the second multiplexer 306 controls the path from the first parameter output device 301 to the second input terminal of the adder 304 to be turned on, the first multiplexer 303 controls the path from the input terminal of the inverter 302 to the first input terminal of the adder 304 to be turned on, the first parameter is 0, 3 is output to the first terminal of the adder 304, the second multiplexer 306 outputs 0 to the second terminal of the adder 304, the adder 304 adds 0 and 3 to obtain 3, and stores 3 in the first register 305.
[0083] Afterwards, the configuration and data cache module 200 inputs the binary data -5 through the input terminal. At this time, the second multiplexer 306 controls the path from the output terminal of the first register 305 to the second input terminal of the adder 304 to be turned on, and the first multiplexer 303 controls the path from the input terminal of the inverter 302 to the first input terminal of the adder 304 to be turned on. The first parameter is 0. At this time, the adder 304 adds -5 to 3 to obtain -2, and the adder 304 outputs -2 to the first register 305.
[0084] Finally, the configuration and data cache module 200 inputs the binary data 7 through the input terminal. At this time, the second multiplexer 306 controls the path from the output terminal of the first register 305 to the second input terminal of the adder 304 to be turned on, and the first multiplexer 303 controls the path from the input terminal of the inverter 302 to the first input terminal of the adder 304 to be turned on. The adder 304 adds -2 and 7 to obtain 5. After that, the adder 304 outputs 5 to the first register 305.
[0085] The binary 5 is the fourth addition result. After the binary 5 is input to the first register 305, the first register 305 outputs the binary 5, and the data 5 is shifted left by the left shifter 309 and reaches the third register 310. Further, the fifth multiplexer 318 is controlled so that the second parameter output device outputs the number of elements 3 in the vector [3, -5, 7], the third register 310 outputs the left-shifted data 5, and the data 3 and the left-shifted data 5 are output to the two input ends of the divider 308, so that the divider 308 divides 5 by 3, and the output data of the divider 308 is right-shifted by the right shifter 311 to obtain the first initial operation result, which is the mean value of the elements in the vector.
[0086] Furthermore, the first initial operation result is stored in the fourth register 320 via the sixth multiplexer 319 .
[0087] In the second step, all elements of the data to be calculated are multiplied by themselves to obtain multiple third multiplication results. The addition submodule is also used to add all the third multiplication results to obtain a fifth addition result. The division submodule is also used to divide the fifth addition result by the total number of elements in the data to be calculated to obtain a second initial calculation result.
[0088] Specifically, the configuration and data cache module 200 first inputs 3 in binary form through the input terminal. At this time, the third multiplexer 313 controls the path from the input terminal to the second input terminal of the multiplier 314, and the adder 304 ignores the data. 3 is input to the first input terminal and the second input terminal of the multiplier 314, so that the multiplier 314 multiplies 3 by 3 to obtain 9, and outputs 9 to the second register 307. The second multiplexer 306 controls the path from the first parameter output device 301 to the adder 304, the first parameter is 0, and the second register 307 outputs 9. The adder 304 adds 9 to 0 to obtain 9, and stores 9 in the first register 305.
[0089] Further, the configuration and data cache module 200 inputs the binary data -5 in binary form through the input terminal. At this time, the adder 304 ignores the data, and -5 is input to the first input terminal and the second input terminal of the multiplier 314, so that the multiplier 314 multiplies -5 by -5 to obtain 25, and outputs 25 to the second register 307. The second multiplexer 306 controls the conduction of the path from the output terminal of the first register 305 to the second input terminal of the adder 304, the first register 305 outputs 9 to the second input terminal of the adder 304, the second register 307 outputs 25 to the first input terminal of the adder 304, the adder 304 adds 25 and 9 to obtain 34, and stores 34 to the first register 305.
[0090] Further, the configuration and data cache module 200 inputs the binary data 7 in binary form through the input terminal. At this time, the adder 304 ignores the data, and 7 is input to the first input terminal and the second input terminal of the multiplier 314, so that the multiplier 314 multiplies 7 by 7 to obtain 49, and outputs 49 to the second register 307. The second multiplexer 306 controls the conduction of the path from the output terminal of the first register 305 to the second input terminal of the adder 304, the first register 305 outputs 34 to the second input terminal of the adder 304, the second register 307 outputs 49 to the first input terminal of the adder 304, and the adder 304 adds 34 and 49 to obtain 83, and stores 83 to the first register 305.
[0091] The binary value 83 is the fifth addition result. The first register 305 outputs the data 83 stored therein. The binary data 83 is left-shifted by the left shifter 309 and reaches the third register 310. The third register 310 outputs the data stored therein, and the fifth multiplexer enables the second parameter output device 312 to output the number of elements 3 in the binary vector [3, -5, 7]. The left-shifted data 83 and data 3 are output to the two input terminals of the divider 308. The divider 308 divides 83 by 3, and right-shifts the output data of the divider 308 by the right shifter 311 to obtain the second initial operation result.
[0092] Moreover, the second initial operation result is also stored in the fourth register 320 via the sixth multiplexer 319 .
[0093] In the third step, the square root of the second initial operation result is obtained to obtain the square root result. The addition submodule is also used to invert the first initial operation result and add it to each element in the data to be operated to obtain the sixth addition result. The division submodule is also used to divide the sixth addition result by the square root result to obtain the third initial operation result.
[0094] Specifically, the fourth register 320 outputs the second initial operation result, the fifth multiplexer 318 controls the path from the fourth register 320 to the output end of the fifth multiplexer 318 to be turned on, the adder 304 ignores the received data, and the third register 310 deletes the received data. The square rooter 315 performs square root operation on the second initial operation result to obtain a square root result.
[0095] The fourth register 320 outputs the first initial operation result. At this time, the multiplier 314 ignores the data. The first multiplexer 303 controls the conduction of the path from the output end of the inverter 302 to the first input end of the adder 304. After being inverted by the inverter 302, the first initial operation result is input to the first input end of the adder 304. At the same time, the first parameter output device 301 outputs 1 to the second input end of the adder 304. The adder 304 adds the first initial operation result to 1 and stores the addition result in the first register 305.
[0096] Afterwards, the first multiplexer 303 controls the conduction of the path from the input end of the inverter 302 to the first input end of the adder 304, and the second multiplexer 306 controls the conduction of the path from the output end of the first register 305 to the second input end of the adder 304. By inputting the element 3 in the data vector [3, -5, 7] to be operated to the first input end of the adder 304, the first register 305 outputs the data stored therein, and the adder 304 adds the data received at the first input end and the data received at the second input end to obtain a sixth addition result, and stores the sixth addition result in the first register 305.
[0097] The first register 305 outputs data to the second input terminal of the divider, and the square root result is output to the first input terminal of the divider 308. The divider 308 divides the data output by the first register 305 by the data output by the third register 310. The right shifter 311 right shifts the data output by the divider 308 to obtain a third initial operation result.
[0098] For each element in the data vector [3, -5, 7] to be operated, operation is performed according to the above method, and three third initial operation results are obtained in total.
[0099] Furthermore, the three third initial operation results are stored in the fourth register 320 via the sixth multiplexer 319 .
[0100] The fourth step is to multiply the third initial operation result with the operation parameter output by the second parameter output device 312 to obtain a fourth multiplication result. The addition submodule is also used to add the fourth multiplication result with the operation parameter output by the first parameter output device 301 to obtain a normalized operation result of the data to be operated.
[0101] Specifically, the third multiplexer 313 controls the conduction of the path from the output end of the second parameter output device 312 to the multiplier 314, the first multiplexer 303 controls the conduction of the path from the input end of the inverter 302 to the adder 304, the second multiplexer 306 controls the conduction of the path from the first parameter output device 301 to the adder 304, and the fourth multiplexer 316 controls the conduction of the path from the output end to the output end of the first register 305. The fourth register 320 outputs a third initial operation result, the adder 304 ignores the data, the multiplier 314 multiplies the third initial operation result with the operation parameter output by the second parameter output device 312 to obtain a fourth multiplication result, and outputs the fourth multiplication result to the second register 307, the second register 307 outputs the fourth multiplication result, and the adder 304 adds the fourth multiplication result to the operation parameter output by the first parameter output device 301 to obtain a normalized operation result corresponding to the third initial operation result.
[0102] The above processing is performed on the above three third initial operation results to obtain the normalized operation result of the data to be operated [3, -5, 7].
[0103] In this way, layer normalization can be achieved for the data to be operated. Assume that the data to be operated is =[ , , , …, ], and use the elements For example, the specific calculation formula can be shown as follows: , in, For the elements Perform L2 normalization, because is a minimum value, so it can be ignored in the calculation. is the second initial operation result, which is the sum of all elements of the data to be operated and then divided by the number of elements. is the first initial operation result, is the mean value of the elements in the data to be operated, is the operation parameter output by the second parameter output device 312, It is the operation parameter output by the first parameter output device 301.
[0104] It should be noted that the operation parameters in the first parameter output device 301 and the second parameter output device 312 are parameters written by the configuration and data cache module 200. For example, the first parameter output device 301 can be set to be implemented by multiple registers and multiplexers, the number of registers is the same as the number of data that can be output by the first parameter output device 301, and the multiplexer selects the corresponding operation parameters for output.
[0105] It should also be noted that since neural network quantization is an optimization technology, it converts the model's weights, activations and other parameters from floating point numbers (such as single-precision floating point numbers Float32, half-precision floating point numbers FLoat16) to integers (such as Int8) to reduce the model's storage requirements and computational complexity, thereby accelerating the model's reasoning speed and reducing power consumption. As shown in the following formula, it is the formula for quantizing the Float32 parameter Weight into INT8's new_wegiht and the corresponding IN8 power exponent exp. The exp of each layer of the neural network is constant.
[0106] .
[0107] The quantized data is input into the normalization function. Since the normalization functions of L1 normalization, L2 normalization, and layer normalization contain integer division, and normalization converts data into a value between (0, 1), the output will be 0, which seriously affects the accuracy. Converting data from integers to floating points requires additional hardware circuits, resulting in excessive area and waste of resources.
[0108] In order to solve this problem, the left shifter 309 and the right shifter 311 are provided, that is, the numerator of the dividend is first shifted left, for example, by 10 bits, and the number of bits of the left shift is determined according to the accuracy requirement. After the division is completed, the numerator is right-shifted according to the power exponent exp of the data at this layer. Compared with the original integer method, the accuracy is improved, and valid data can be obtained, and the accuracy loss of the result obtained by the floating-point operation is not large. Since the adder 304 is an integer adder, the multiplier 314 is an integer multiplier, the divider 308 is an integer divider, and the square root extractor 315 is an integer square root extractor, the circuit uses integer calculation instead of floating-point calculation, which reduces power consumption and area within an acceptable accuracy loss range and improves performance.
[0109] In summary, the normalization operation circuit of the embodiment of the present invention includes: a configuration and data cache module, which is used to determine the data to be operated, the normalization function type and the operation parameters; an arithmetic logic module, which is connected to the configuration and data cache module, and the arithmetic logic module includes multiple operation sub-modules, which are used to determine the target operation sub-module from multiple operation sub-modules according to the normalization function type, and call the target operation sub-module according to the normalization function type and the operation parameters to perform normalization operation on the data to be operated, wherein the multiple operation sub-modules include an addition sub-module, a multiplication sub-module, a division sub-module and a square root sub-module. Therefore, by adopting an additionally set arithmetic logic module to perform normalization operation, there is no need for device parameters such as CPU and GPU, thereby meeting the needs of edge computing.
[0110] Furthermore, the present invention proposes a normalization operation system.
[0111] Figure 4 It is a structural block diagram of a normalization operation system according to an embodiment of the present invention.
[0112] like Figure 4 As shown, the normalization operation system 10 includes the normalization operation circuit 100 mentioned above.
[0113] The normalized operation system of the embodiment of the present invention can realize normalized operation through the normalized operation circuit of the above embodiment without the participation of device parameters such as CPU and GPU, thereby meeting the needs of edge computing.
[0114] It should be noted that the logic and / or steps represented in the flowchart or described in other ways herein can be considered as a sequenced list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, device or equipment (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, device or equipment and execute instructions), or in combination with these instruction execution systems, devices or equipment. For the purpose of this specification, "computer-readable medium" can be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, device or equipment, or in combination with these instruction execution systems, devices or equipment. More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection portion with one or more wirings (electronic device), a portable computer disk box (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and editable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium and then editing, interpreting or processing in other suitable ways if necessary, and then stored in a computer memory.
[0115] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiment, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. If implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0116] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.
[0117] In the description of this specification, the terms "center", "longitudinal", "lateral", "length", "width", "thickness", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", "clockwise", "counterclockwise", "axial", "radial", "circumferential" and the like indicate orientations or positional relationships based on the orientations or positional relationships shown in the accompanying drawings, and do not indicate or imply that the referred device or element must have a specific orientation, be constructed and operated in a specific orientation, and cannot be understood as a limitation on the present invention.
[0118] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of the features. In the description of the present invention, the meaning of "plurality" is at least two, such as two, three, etc., unless otherwise clearly and specifically defined.
[0119] In the description of this specification, unless otherwise specified, the terms "installed", "connected", "connected", "fixed" and the like should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium, it can be the internal connection of two elements or the interaction relationship between two elements, unless otherwise clearly defined. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0120] In the present invention, unless otherwise clearly specified and limited, a first feature being "above" or "below" a second feature may mean that the first and second features are in direct contact, or the first and second features are in indirect contact through an intermediate medium. Moreover, a first feature being "above", "above" or "above" a second feature may mean that the first feature is directly above or obliquely above the second feature, or simply means that the first feature is higher in level than the second feature. A first feature being "below", "below" or "below" a second feature may mean that the first feature is directly below or obliquely below the second feature, or simply means that the first feature is lower in level than the second feature.
[0121] Although the embodiments of the present invention have been shown and described above, it is to be understood that the above embodiments are exemplary and are not to be construed as limitations of the present invention. A person skilled in the art may change, modify, replace and vary the above embodiments within the scope of the present invention.
Claims
1. A normalization operation circuit, characterized in that: The circuit comprises: Configuration and data cache module, used to determine the data to be calculated, the normalization function type and the calculation parameters; an arithmetic logic module connected to the configuration and data cache module, the arithmetic logic module comprising a plurality of operator modules, for determining a target operator module from the plurality of operator modules according to the normalization function type, and calling the target operator module to perform a normalization operation on the data to be operated according to the normalization function type and the operation parameters, wherein the plurality of operator modules comprise an addition submodule, a multiplication submodule, a division submodule and a square root submodule; The first input end of the addition submodule is designed as the input end of the arithmetic logic module, and is used to receive the data to be operated. The output end of the addition submodule is connected to the first input end of the division submodule, the second input end of the division submodule is connected to the first input end of the addition submodule, and the output end of the division submodule is designed as the output end of the arithmetic logic module; wherein, When the normalization function type is the first preset type, the data to be operated is a vector, and the target operation submodule includes the addition submodule and the division submodule. The addition submodule is used to sum the absolute values of all elements of the data to be operated to obtain a first summation result, and the division submodule is used to divide the data to be operated by the first summation result to obtain a normalized operation result of the data to be operated.
2. The normalization operation circuit according to claim 1, characterized in that: The adding submodule comprises: a first parameter output device, an inverter, a first multiplexer, an adder, a first register, a second multiplexer, and a second register; wherein the input end of the inverter is designed as the first input end of the adding submodule; the first input end of the first multiplexer is connected to the output end of the inverter, and the second input end of the first multiplexer is connected to the input end of the inverter; the first input end of the adder is connected to the output end of the first multiplexer; the input end of the first register is connected to the output end of the adder; the first input end of the second multiplexer is connected to the output end of the first parameter output device, the second input end of the second multiplexer is connected to the output end of the first register, and the output end of the second multiplexer is connected to the second input end of the adder; the input end of the second register is connected to the output end of the first register, and the output end of the second register is connected to the input end of the inverter; the output end of the first register is also set as the output end of the adding submodule; The division submodule includes: a divider, a left shifter, a third register, and a right shifter; wherein the first input end of the divider is designed as the first input end of the division submodule, and the second input end of the divider is connected to the output end of the third register; the input end of the left shifter is designed as the second input end of the division submodule; the input end of the third register is connected to the output end of the left shifter; the input end of the right shifter is connected to the output end of the divider, and the output end of the right shifter is designed as the output end of the division submodule.
3. The normalization operation circuit according to claim 2, characterized in that: When the normalization function type is the first preset type, for each element in the data to be operated, when the element is input into the addition submodule, the first multiplexer outputs the element to the adder when the element is a positive number, and outputs the element inverted by the inverter to the adder when the element is a negative number, and the adder adds the output data of the first multiplexer to the operation parameter output by the first parameter output device to obtain the first added data; when no data is stored in the second register and there are elements that are not input into the data to be operated, the first added data is stored in the second register via the first register; when no data is stored in the second register and the data to be operated is When all elements of the data to be calculated have been input into the addition submodule, it is determined that the first added data is the first summation result; when data has been stored in the second register, the first added data is output to the adder via the first register and the second multiplexer, and the first stored data in the second register is output to the adder via the first multiplexer, and the adder adds the first added data to the first stored data to obtain second added data; if there are elements that are not input into the data to be calculated, the second added data is stored as new first stored data in the second register; if all elements of the data to be calculated have been input into the addition submodule, it is determined that the second added data is the first summation result; The left shifter shifts the data to be calculated leftwards, the divider divides the left-shifted data to be calculated by the first summation result, and the right shifter shifts the output data of the divider rightwards to obtain a normalized calculation result of the data to be calculated.
4. The normalization operation circuit according to claim 2, characterized in that: The input end of the second register is designed as the second input end of the addition submodule, the first input end and the second input end of the multiplication submodule are both connected to the first input end of the addition submodule, the output end of the multiplication submodule is connected to the second input end of the addition submodule, the input end of the square root submodule is connected to the output end of the addition submodule, and the output end of the square root submodule is connected to the first input end of the division submodule; The multiplication submodule comprises: a second parameter output device, a third multiplexer, and a multiplier; wherein the first input end of the third multiplexer is connected to the output end of the second parameter output device, and the second input end of the third multiplexer is designed as the second input end of the multiplication submodule; the first input end of the multiplier is designed as the first input end of the multiplication submodule, the second input end of the multiplier is connected to the output end of the third multiplexer, and the output end of the multiplier is designed as the output end of the multiplication submodule; The square root submodule includes a square root generator, the input end of the square root generator is designed as the input end of the square root submodule, and the output end of the square root generator is designed as the output end of the square root submodule.
5. The normalization operation circuit according to claim 4, characterized in that: When the normalization function type is the second preset type, the data to be operated is a vector, and the target operation submodule includes the addition submodule, the multiplication submodule, the division submodule and the square root submodule; Wherein, for each element in the data to be operated, the element is input to the first input terminal and the second input terminal of the multiplier, so that the multiplier multiplies the element with itself to obtain a first multiplication result, and the first multiplication result is stored in the second register; if the first register does not store data, the second register outputs the first multiplication result to the adder through the first multiplexer, and the adder adds the first multiplication result to the operation parameter output by the first parameter output device, and stores the addition result in the first register; if the first register has data stored, the first register outputs the data stored therein to the adder through the second multiplexer, and the adder adds the data stored in the first register to the first multiplication result, and stores the addition result in the first register; After the addition results corresponding to each element in the data to be calculated are stored in the first register, the data in the first register is used as the third addition result, and the third addition result is output to the square root generator. The square root generator takes the square root of the third addition result to obtain a square root result, and outputs the square root result to the divider. The divider divides the data to be calculated by the square root result to obtain a normalized calculation result of the data to be calculated.
6. The normalization operation circuit according to claim 4, characterized in that: The arithmetic logic module also includes a fourth multiplexer, a first input end of the fourth multiplexer is connected to the output end of the right shifter, an output end of the fourth multiplexer is connected to the output end of the first register, and an output end of the fourth multiplexer is designed to be the output end of the arithmetic logic module.
7. The normalization operation circuit according to claim 6, characterized in that: When the normalization function type is a third preset type, the target operator module includes the addition submodule and the multiplication submodule; Among them, the third multiplexer outputs the operation parameters output by the second parameter output device to the multiplier. After receiving the data to be operated, the multiplier multiplies the data to be operated by the operation parameters output by the second parameter output device to obtain a second multiplication result, and outputs the second multiplication result to the adder through the second register and the first multiplexer. The adder adds the second multiplication result and the operation parameters output by the first parameter output device to obtain a normalized operation result of the data to be operated.
8. The normalization operation circuit according to claim 4, characterized in that: The input end of the square root submodule is also connected to the input end of the addition submodule and the second input end of the division submodule. When the normalization function type is the fourth preset type, the data to be operated is a vector, and the target operation submodule includes the addition submodule, the multiplication submodule, the division submodule and the square root submodule. Among them, the addition submodule is used to sum all elements of the data to be calculated to obtain a fourth addition result, and the division submodule is used to divide the fourth addition result by the total number of elements in the data to be calculated to obtain a first initial calculation result; the multiplication submodule is used to multiply all elements of the data to be calculated by themselves to obtain multiple third multiplication results, and the addition submodule is also used to add all the third multiplication results to obtain a fifth addition result, and the division submodule is also used to divide the fifth addition result by the total number of elements in the data to be calculated to obtain a second initial calculation result; the square root submodule is used to multiply the third multiplication results by all the third multiplication results. The second initial operation result is square rooted to obtain a square root result. The addition submodule is also used to invert the first initial operation result and add it to each element in the data to be operated to obtain a sixth addition result. The division submodule is also used to divide the sixth addition result by the square root result to obtain a third initial operation result. The multiplication submodule is also used to multiply the third initial operation result with the operation parameter output by the second parameter output device to obtain a fourth multiplication result. The addition submodule is also used to add the fourth multiplication result with the operation parameter output by the first parameter output device to obtain a normalized operation result of the data to be operated.
9. A normalization operation system, characterized in that: The method comprises a normalization operation circuit according to any one of claims 1 to 8.
Citation Information
Patent Citations
Resource reuse type neural network hardware acceleration circuit based on fast convolution
CN112862091A
Calibration method and device of analog circuit for executing neural network calculation
CN114819051A
Neural network model compiling method and device and computer readable storage medium
CN115640017A
Hardware architecture and method for approximately calculating batch normalization layer
CN116205287A
Arithmetic logic unit, operation processing method, chip and electronic equipment
CN117873427A