Normalization operation circuit, system

By designing a normalized computing circuit including configuration and data cache modules and arithmetic logic modules, the problems of real-time and low power consumption in the edge computing environment in the prior art are solved, and efficient normalized computing is achieved.

CN120011729BActive Publication Date: 2025-06-27HEFEI XINCHE INFINITY SEMICONDUCTOR TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510477314.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-06-27
Estimated Expiration
2045-04-16

AI Technical Summary

Technical Problem

In the prior art, normalized computing is difficult to meet the needs of real-time and low power consumption in an edge computing environment. The CPU lacks special circuit support, and GPU resources are consumed and energy efficiency is not good.

Method used

A normalized operation circuit is designed, including configuration and data cache modules and arithmetic logic modules. The normalized operation of the data to be operated by multiple operation submodules (such as addition submodules, multiplication submodules, division submodules and root-digit submodules) is called according to the normalized function type.

Benefits of technology

This design does not require CPU or GPU support, and can effectively meet the real-time and low-power requirements of edge computing, improving system performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011729B_ABST
    Figure CN120011729B_ABST
Patent Text Reader

Abstract

The present invention discloses a normalization operation circuit and system, relating to the technical field of normalization. Among them, the circuit includes: a configuration and data caching module, which is used to determine the data to be operated, the type of normalization function, and the operation parameters; an arithmetic logic module, which is connected to the configuration and data caching module. The arithmetic logic module includes a plurality of operation sub-modules, which are used to determine a target operation sub-module from the plurality of operation sub-modules according to the type of normalization function, and call the target operation sub-module to perform normalization operation on the data to be operated according to the type of normalization function and the operation parameters. Among them, the plurality of operation sub-modules include an addition sub-module, a multiplication sub-module, a division sub-module, and a square root sub-module.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of normalization, and in particular, to a normalization operation circuit and system. Background Art

[0002] With the rapid development of artificial intelligence, various neural network models have emerged continuously, and more and more operators are widely used in the calculation process of neural networks. Normalization is an operator of neural networks and a data preprocessing technology, aiming to adjust data to a specific range or distribution to make it more suitable for subsequent calculations or analyses. The core idea of normalization is to scale or transform the data so that data in different dimensions or features are numerically within the same scale range.

[0003] Edge computing is a distributed computing mode, and its core idea is to perform data processing on devices close to the data source or users, rather than sending all data to a remote central server or cloud computing platform for processing. By moving part of the computing, storage, and data processing work to the "edge" of the network (such as routers, gateways, edge servers, etc.), latency can be reduced and bandwidth requirements can be lowered, thereby improving real-time performance. Therefore, if the normalization operation can be implemented in the mode of edge computing, the system performance can be further improved.

[0004] However, in the related art, the normalization operation usually uses a CPU (Central Processing Unit) or a GPU (Graphics Processing Unit) to implement. The CPU does not have a specific circuit for operating the normalization operator. Limited by its ability to execute programs serially, the normalization operator needs to use a long instruction stream for calculation, and the real-time performance cannot meet the requirements of edge computing. The GPU consumes a large amount of hardware resources and has high power consumption, resulting in poor energy efficiency and not meeting the low-power and cost requirements of edge computing. Summary of the Invention

[0005] The present invention aims to solve at least one of the technical problems in the related art to some extent. For this purpose, the first object of the present invention is to provide a normalization operation circuit to meet the requirements of edge computing.

[0006] The second object of the present invention is to provide a normalization operation system.

[0007] To achieve the above object, an embodiment of the first aspect of the present invention provides a normalization operation circuit, which includes: a configuration and data caching module for determining data to be operated, a normalization function type, and operation parameters; an arithmetic logic module connected to the configuration and data caching module, the arithmetic logic module including a plurality of operation sub-modules, for determining a target operation sub-module from the plurality of operation sub-modules according to the normalization function type, and calling the target operation sub-module to perform a normalization operation on the data to be operated according to the normalization function type and the operation parameters, wherein the plurality of operation sub-modules include an addition sub-module, a multiplication sub-module, a division sub-module, and a square root sub-module; a first input end of the addition sub-module is designed as an input end of the arithmetic logic module for receiving the data to be operated, an output end of the addition sub-module is connected to a first input end of the division sub-module, a second input end of the division sub-module is connected to the first input end of the addition sub-module, and an output end of the division sub-module is designed as an output end of the arithmetic logic module; wherein, when the normalization function type is a first preset type, the data to be operated is a vector, the target operation sub-module includes the addition sub-module and the division sub-module, the addition sub-module is used to sum the absolute values of all elements of the data to be operated to obtain a first summation result, and the division sub-module is used to divide the data to be operated by the first summation result to obtain a normalization operation result of the data to be operated.

[0008] In addition, the normalization operation circuit according to the embodiment of the present invention may further have the following additional technical features:

[0009] In one embodiment of the present invention, the addition sub-module includes: a first parameter output device, an inverter, a first multiplexer, an adder, a first register, a second multiplexer, and a second register; wherein, the input end of the inverter is designed as the first input end of the addition sub-module; the first input end of the first multiplexer is connected to the output end of the inverter, and the second input end of the first multiplexer is connected to the input end of the inverter; the first input end of the adder is connected to the output end of the first multiplexer; the input end of the first register is connected to the output end of the adder; the first input end of the second multiplexer is connected to the output end of the first parameter output device, the second input end of the second multiplexer is connected to the output end of the first register, and the output end of the second multiplexer is connected to the second input end of the adder; the input end of the second register is connected to the output end of the first register, and the output end of the second register is connected to the input end of the inverter; the output end of the first register is further set as the output end of the addition sub-module; the division sub-module includes: a divider, a left shifter, a third register, and a right shifter; wherein, the first input end of the divider is designed as the first input end of the division sub-module, and the second input end of the divider is connected to the output end of the third register; the input end of the left shifter is designed as the second input end of the division sub-module; the input end of the third register is connected to the output end of the left shifter; the input end of the right shifter is connected to the output end of the divider, and the output end of the right shifter is designed as the output end of the division sub-module.

[0010] In one embodiment of the present invention, when the normalization function type is the first preset type, for each element in the data to be operated on, when the element is input to the adder sub-module, the first multiplexer outputs the element to the adder when the element is positive, and outputs the element after being inverted by the inverter to the adder when the element is negative. The adder adds the output data of the first multiplexer and the operation parameter output by the first parameter output device to obtain first addition data. When no data is stored in the second register and there are elements in the data to be operated on that have not been input, the first addition data is stored in the second register through the first register. When no data is stored in the second register and all elements of the data to be operated on have been input to the adder sub-module, the first addition data is determined as the first summation result. When data is already stored in the second register, the first addition data is output to the adder through the first register and the second multiplexer, and the first stored data in the second register is output to the adder through the first multiplexer. The adder adds the first addition data and the first stored data to obtain second addition data. If there are elements in the data to be operated on that have not been input, the second addition data is stored in the second register as the new first stored data. If all elements of the data to be operated on have been input to the adder sub-module, the second addition data is determined as the first summation result. The left shifter shifts the data to be operated on to the left, the divider divides the data to be operated on after being shifted to the left by the first summation result, and the right shifter shifts the data output by the divider to the right to obtain the normalization operation result of the data to be operated on.

[0011] In an embodiment of the present invention, the input end of the second register is designed as the second input end of the addition sub-module. The first input end and the second input end of the multiplication sub-module are both connected to the first input end of the addition sub-module. The output end of the multiplication sub-module is connected to the second input end of the addition sub-module. The input end of the square root extraction sub-module is connected to the output end of the addition sub-module. The output end of the square root extraction sub-module is connected to the first input end of the division sub-module. The multiplication sub-module includes: a second parameter output device, a third multiplexer, and a multiplier. Among them, the first input end of the third multiplexer is connected to the output end of the second parameter output device. The second input end of the third multiplexer is designed as the second input end of the multiplication sub-module. The first input end of the multiplier is designed as the first input end of the multiplication sub-module. The second input end of the multiplier is connected to the output end of the third multiplexer. The output end of the multiplier is designed as the output end of the multiplication sub-module. The square root extraction sub-module includes a square root extractor. The input end of the square root extractor is designed as the input end of the square root extraction sub-module. The output end of the square root extractor is designed as the output end of the square root extraction sub-module.

[0012] In an embodiment of the present invention, when the normalization function type is the second preset type, the data to be operated on is a vector. The target operation sub-module includes the addition sub-module, the multiplication sub-module, the division sub-module, and the square root extraction sub-module. Among them, for each element in the data to be operated on, the element is input to the first input end and the second input end of the multiplier, so that the multiplier multiplies the element by itself to obtain a first multiplication result, and stores the first multiplication result in the second register. If no data is stored in the first register, the second register outputs the first multiplication result to the adder through the first multiplexer. The adder adds the first multiplication result to the operation parameter output by the first parameter output device and stores the added result in the first register. If data is already stored in the first register, the first register outputs the data stored therein to the adder through the second multiplexer. The adder adds the data stored in the first register to the first multiplication result and stores the added result in the first register. After the added results corresponding to each element in the data to be operated on are stored in the first register, the data in the first register is used as the third added result, and the third added result is output to the square root extractor. The square root extractor takes the square root of the third added result to obtain a square root result, and outputs the square root result to the divider. The divider divides the data to be operated on by the square root result to obtain the normalization operation result of the data to be operated on.

[0013] In one embodiment of the present invention, the arithmetic logic module further includes a fourth multiplexer. The first input end of the fourth multiplexer is connected to the output end of the right shifter. The output end of the fourth multiplexer is connected to the output end of the first register. The output end of the fourth multiplexer is designed as the output end of the arithmetic logic module.

[0014] In one embodiment of the present invention, when the normalization function type is the third preset type, the target operation sub-module includes the addition sub-module and the multiplication sub-module; wherein, the third multiplexer outputs the operation parameter output by the second parameter output device to the multiplier. After receiving the data to be operated, the multiplier multiplies the data to be operated by the operation parameter output by the second parameter output device to obtain a second multiplication result, and outputs the second multiplication result to the adder through the second register and the first multiplexer. The adder adds the second multiplication result to the operation parameter output by the first parameter output device to obtain the normalization operation result of the data to be operated.

[0015] In one embodiment of the present invention, the input end of the square root sub-module is further connected to the input end of the addition sub-module and the second input end of the division sub-module. When the normalization function type is the fourth preset type, the data to be operated is a vector, and the target operation sub-module includes the addition sub-module, the multiplication sub-module, the division sub-module and the square root sub-module; wherein, the addition sub-module is used to sum all elements of the data to be operated to obtain a fourth addition result. The division sub-module is used to divide the fourth addition result by the total number of elements in the data to be operated to obtain a first initial operation result. The multiplication sub-module is used to multiply all elements of the data to be operated by themselves to obtain a plurality of third multiplication results. The addition sub-module is further used to add all the third multiplication results to obtain a fifth addition result. The division sub-module is further used to divide the fifth addition result by the total number of elements in the data to be operated to obtain a second initial operation result. The square root sub-module is used to take the square root of the second initial operation result to obtain a square root result. The addition sub-module is further used to add the negated first initial operation result to each element in the data to be operated to obtain a sixth addition result. The division sub-module is further used to divide the sixth addition result by the square root result to obtain a third initial operation result. The multiplication sub-module is further used to multiply the third initial operation result by the operation parameter output by the second parameter output device to obtain a fourth multiplication result. The addition sub-module is further used to add the fourth multiplication result to the operation parameter output by the first parameter output device to obtain the normalization operation result of the data to be operated.

[0016] To achieve the above object, an embodiment of the second aspect of the present invention provides a normalization operation system, including the above normalization operation circuit.

[0017] According to the normalization operation circuit and system of the embodiments of the present invention, it includes: a configuration and data cache module, configured to determine the data to be operated, the type of normalization function, and operation parameters; an arithmetic logic module, connected to the configuration and data cache module, the arithmetic logic module includes a plurality of operation sub-modules, configured to determine a target operation sub-module from the plurality of operation sub-modules according to the type of normalization function, and call the target operation sub-module to perform normalization operation on the data to be operated according to the type of normalization function and operation parameters, wherein the plurality of operation sub-modules include an addition sub-module, a multiplication sub-module, a division sub-module, and a square root sub-module. Thus, by using the additionally provided arithmetic logic module for normalization operation, no CPU, GPU and other device parameters are required, thereby meeting the requirements of edge computing.

[0018] The additional aspects and advantages of the present invention will be given in part in the following description, become apparent in part from the following description, or be understood through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 is a structural block diagram of the normalization operation circuit according to an embodiment of the present invention;

[0020] Figure 2 is a working schematic diagram of the normalization operation circuit according to an embodiment of the present invention;

[0021] Figure 3 is a circuit diagram of the normalization operation circuit according to a specific embodiment of the present invention;

[0022] Figure 4 is a structural block diagram of the normalization operation system according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0023] The normalization operation circuit and system according to the embodiments of the present invention will be described below with reference to the drawings, where the same or similar reference numerals represent the same or similar elements or elements having the same or similar functions throughout. The embodiments described with reference to the drawings are exemplary and should not be construed as limiting the present invention.

[0024] Figure 1 is a structural block diagram of the normalization operation circuit according to an embodiment of the present invention.

[0025] As Figure 1As shown in the figure, the normalization operation circuit 100 includes: a configuration and data caching module 200, which is used to determine the data to be operated on, the type of normalization function, and the operation parameters; an arithmetic logic module 300, which is connected to the configuration and data caching module 200. The arithmetic logic module 300 includes multiple operation sub-modules, which are used to determine the target operation sub-module from multiple operation sub-modules according to the type of normalization function, and call the target operation sub-module to perform normalization operation on the data to be operated on according to the type of normalization function and the operation parameters. Among them, the multiple operation sub-modules include an addition sub-module, a multiplication sub-module, a division sub-module, and a square root sub-module.

[0026] Common normalization methods include L1 normalization, L2 normalization, batch normalization, and layer normalization. There are three reasons for using normalization: (1) It can avoid the difference in the magnitude of different features. In a dataset, the numerical ranges of different features may vary greatly. If normalization is not performed, this feature may dominate the calculation process and affect the learning effect of the model. (2) It speeds up the convergence rate. In machine learning, normalization can speed up the training speed of the model. Especially when using the stochastic gradient descent method, the normalized data will make the update of model parameters more stable, avoiding some features from being updated too fast or too slow. (3) Algorithms such as the Single Shot Multi-Box Detector algorithm are sensitive to the relative distance between features. Normalization can make the weights between different features more balanced and improve the accuracy of the model.

[0027] Therefore, in order to implement normalization operations based on multiple normalization methods and meet the requirements of edge computing, the above-mentioned normalization operation circuit 100 is set up.

[0028] The above-mentioned normalization operation circuit 100 includes a configuration and data caching module 200 and an arithmetic logic module 300. The configuration and data caching module 200 is used to determine the data to be operated on, the type of normalization function, and the operation parameters. The above-mentioned types of normalization functions include L1 normalization, L2 normalization, batch normalization, and layer normalization.

[0029] The above-mentioned arithmetic logic module 300 is a specific operation module, including an addition sub-module, a multiplication sub-module, a division sub-module, and a square root sub-module. After the configuration and data caching module 200 determines the type of normalization function and outputs the operation parameters, the arithmetic logic module 300 can call the operation sub-module according to the determined type of normalization function and perform normalization operation on the data to be operated on based on the operation parameters.

[0030] That is to say, in order to implement various forms of normalization and meet the requirements of edge computing, a configuration and data caching module 200 and an arithmetic logic module 300 are set up. See Figure 2, the configuration and data cache module 200 is used to receive integer data and commands. The integer data includes the integer data after quantization of the neural network layer, and this integer data is the data to be operated on. The command is a command including the type, size, and parameters of the neural network layer. After receiving the integer data and the command, the configuration and data cache module 200 converts the integer data into a form that can be processed by the arithmetic logic module 300, sends it to the arithmetic logic module 300, determines which normalization to perform specifically according to the command, and generates operation parameters corresponding to the normalization to be performed, and sends the specific normalization function type to be performed and the corresponding operation parameters to the arithmetic logic module 300. The configuration and data cache module 200 can be integrated into the circuit using digital logic circuits.

[0031] After receiving the data to be operated on, the normalization function type, and the operation parameters, the arithmetic logic module 300 uses the operation sub-modules included therein to perform a normalization operation on the data to be operated on, and obtains a normalization operation result. Figure 2 The result in is the normalization operation result. Moreover, since the arithmetic logic module 300 uses multiple operation sub-modules included therein to perform the normalization operation, the control of the arithmetic logic module 300 only needs to control which specific operation sub-modules work and how each operation sub-module works specifically, so that the control of the arithmetic logic module 300 can be implemented using digital logic circuits without the participation of devices such as CPUs and GPUs, thus meeting the requirements of edge computing.

[0032] In some embodiments of the present invention, the first input end of the addition sub-module is designed as the input end of the arithmetic logic module 300 for receiving the data to be operated on. The output end of the addition sub-module is connected to the first input end of the division sub-module, and the second input end of the division sub-module is connected to the first input end of the addition sub-module. The output end of the division sub-module is designed as the output end of the arithmetic logic module 300; wherein, when the normalization function type is the first preset type, the data to be operated on is a vector, and the target operation sub-modules include the addition sub-module and the division sub-module. The addition sub-module is used to sum the absolute values of all elements of the data to be operated on to obtain a first summation result, and the division sub-module is used to divide the data to be operated on by the first summation result to obtain the normalization operation result of the data to be operated on.

[0033] In some embodiments of the present invention, the addition sub-module includes: a first parameter output device, an inverter, a first multiplexer, an adder, a first register, a second multiplexer, and a second register; wherein, the input end of the inverter is designed as the first input end of the addition sub-module; the first input end of the first multiplexer is connected to the output end of the inverter, and the second input end of the first multiplexer is connected to the input end of the inverter; the first input end of the adder is connected to the output end of the first multiplexer; the input end of the first register is connected to the output end of the adder; the first input end of the second multiplexer is connected to the output end of the first parameter output device, the second input end of the second multiplexer is connected to the output end of the first register, and the output end of the second multiplexer is connected to the second input end of the adder; the input end of the second register is connected to the output end of the first register, and the output end of the second register is connected to the input end of the inverter; the output end of the first register is further set as the output end of the addition sub-module;

[0034] The division sub-module includes: a divider, a left shifter, a third register, and a right shifter; wherein, the first input end of the divider is designed as the first input end of the division sub-module, and the second input end of the divider is connected to the output end of the third register; the input end of the left shifter is designed as the second input end of the division sub-module; the input end of the third register is connected to the output end of the left shifter; the input end of the right shifter is connected to the output end of the divider, and the output end of the right shifter is designed as the output end of the division sub-module.

[0035] In some embodiments of the present invention, when the normalization function type is the first preset type, for each element in the data to be operated, when the element is input to the addition sub-module, the first multiplexer outputs the element to the adder when the element is positive, and outputs the element after being inverted by the inverter to the adder when the element is negative. The adder adds the output data of the first multiplexer and the operation parameter output by the first parameter output device to obtain the first added data; when there is no data stored in the second register and there are elements in the data to be operated that have not been input, the first added data is stored in the second register through the first register; when there is no data stored in the second register and all elements of the data to be operated have been input to the addition sub-module, the first added data is determined as the first summation result; when there is data stored in the second register, the first added data is output to the adder through the first register and the second multiplexer, and the first stored data in the second register is output to the adder through the first multiplexer. The adder adds the first added data and the first stored data to obtain the second added data; if there are elements in the data to be operated that have not been input, the second added data is stored in the second register as the new first stored data; if all elements of the data to be operated have been input to the addition sub-module, the second added data is determined as the first summation result;

[0036] The left shifter shifts the data to be operated on to the left, the divider divides the data to be operated on after left shifting by the first summation result, and the right shifter shifts the data output by the divider to the right to obtain the normalized operation result of the data to be operated on.

[0037] In some embodiments of the present invention, the input end of the second register is designed as the second input end of the addition sub-module. The first input end and the second input end of the multiplication sub-module are both connected to the first input end of the addition sub-module. The output end of the multiplication sub-module is connected to the second input end of the addition sub-module. The input end of the square root extraction sub-module is connected to the output end of the addition sub-module. The output end of the square root extraction sub-module is connected to the first input end of the division sub-module;

[0038] The multiplication sub-module includes: a second parameter output device, a third multiplexer, and a multiplier; wherein, the first input end of the third multiplexer is connected to the output end of the second parameter output device, and the second input end of the third multiplexer is designed as the second input end of the multiplication sub-module; the first input end of the multiplier is designed as the first input end of the multiplication sub-module, the second input end of the multiplier is connected to the output end of the third multiplexer, and the output end of the multiplier is designed as the output end of the multiplication sub-module;

[0039] The square root extraction sub-module includes a square root extractor. The input end of the square root extractor is designed as the input end of the square root extraction sub-module, and the output end of the square root extractor is designed as the output end of the square root extraction sub-module.

[0040] In some embodiments of the present invention, when the normalization function type is the second preset type, the data to be operated on is a vector, and the target operation sub-module includes an addition sub-module, a multiplication sub-module, a division sub-module, and a square root extraction sub-module;

[0041] Among them, for each element in the data to be operated on, the element is input to the first input end and the second input end of the multiplier, so that the multiplier multiplies the element by itself to obtain a first multiplication result, and stores the first multiplication result in the second register; if there is no data stored in the first register, the second register outputs the first multiplication result to the adder through the first multiplexer, and the adder adds the first multiplication result to the operation parameter output by the first parameter output device, and stores the added result in the first register; if there is already data stored in the first register, the first register outputs the data stored therein to the adder through the second multiplexer, and the adder adds the data stored in the first register to the first multiplication result, and stores the added result in the first register;

[0042] After the addition results corresponding to each element in the data to be operated are stored in the first register, the data in the first register is used as the third addition result, and the third addition result is output to the square root calculator. The square root calculator takes the square root of the third addition result to obtain the square root result, and outputs the square root result to the divider. The divider divides the data to be operated by the square root result to obtain the normalized operation result of the data to be operated.

[0043] In some embodiments of the present invention, the arithmetic logic module 300 further includes a fourth multiplexer. The first input terminal of the fourth multiplexer is connected to the output terminal of the right shifter, the output terminal of the fourth multiplexer is connected to the output terminal of the first register, and the output terminal of the fourth multiplexer is designed as the output terminal of the arithmetic logic module 300.

[0044] In some embodiments of the present invention, when the normalization function type is the third preset type, the target operation sub-module includes an addition sub-module and a multiplication sub-module;

[0045] Among them, the third multiplexer outputs the operation parameters output by the second parameter output device to the multiplier. After receiving the data to be operated, the multiplier multiplies the data to be operated by the operation parameters output by the second parameter output device to obtain the second multiplication result, and outputs the second multiplication result to the adder through the second register and the first multiplexer. The adder adds the second multiplication result and the operation parameters output by the first parameter output device to obtain the normalized operation result of the data to be operated.

[0046] In some embodiments of the present invention, the input terminal of the square root sub-module is further connected to the input terminal of the addition sub-module and the second input terminal of the division sub-module. When the normalization function type is the fourth preset type, the data to be operated is a vector, and the target operation sub-module includes an addition sub-module, a multiplication sub-module, a division sub-module, and a square root sub-module;

[0047] Among them, the addition sub-module is used to sum all elements of the data to be operated to obtain a fourth addition result, and the division sub-module is used to divide the fourth addition result by the total number of elements in the data to be operated to obtain a first initial operation result; the multiplication sub-module is used to multiply all elements of the data to be operated by themselves to obtain a plurality of third multiplication results, and the addition sub-module is also used to sum all the third multiplication results to obtain a fifth addition result, and the division sub-module is also used to divide the fifth addition result by the total number of elements in the data to be operated to obtain a second initial operation result; the square root extraction sub-module is used to extract the square root of the second initial operation result to obtain a square root extraction result, and the addition sub-module is also used to add the negated first initial operation result to each element in the data to be operated to obtain a sixth addition result, and the division sub-module is also used to divide the sixth addition result by the square root extraction result to obtain a third initial operation result; the multiplication sub-module is also used to multiply the third initial operation result by the operation parameter output by the second parameter output device to obtain a fourth multiplication result, and the addition sub-module is also used to add the fourth multiplication result to the operation parameter output by the first parameter output device to obtain a normalized operation result of the data to be operated.

[0048] The following will describe the arithmetic logic module 300 in conjunction with Figure 3 the specific embodiments shown below.

[0049] In Figure 3 the specific embodiment shown, the control of the arithmetic logic module 300 is implemented by the multiplexer in this module, and the multiplexer and register are controlled by digital logic circuits.

[0050] In Figure 3 , 301 is the first parameter output device, 302 is the inverter, 303 is the first multiplexer, 304 is the adder, 305 is the first register, 306 is the second multiplexer, 307 is the second register, 308 is the divider, 309 is the left shifter, 310 is the third register, 311 is the right shifter, 312 is the second parameter output device, 313 is the third multiplexer, 314 is the multiplier, 315 is the square root extractor, 316 is the fourth multiplexer, 318 is the fifth multiplexer, 319 is the sixth multiplexer, 320 is the fourth register, the input is the input end of the arithmetic logic module 300, and the output is the output end of the arithmetic logic module 300. The above adder 304 is an integer adder, the divider 308 is an integer divider, the multiplier 314 is an integer multiplier, and the square root extractor 315 is an integer square root extractor.

[0051] The above first preset type is L1 normalization, the second preset type is L2 normalization, the third preset type is batch normalization, and the fourth preset type is layer normalization.

[0052] Specifically, when the normalization function type is the first preset type, the data to be operated on is a vector.

[0053] The fourth multiplexer 316 controls the conduction of the path from the right shifter 311 to the output terminal, and the fifth multiplexer 318 controls the conduction of the path from the input terminal to the input terminal of the left shifter 309.

[0054] The configuration and data cache module 200 inputs the elements of the data to be operated on one by one. Assuming the data to be operated on is a vector , the configuration and data cache module 200 first inputs the element through the input terminal , after the arithmetic logic module 300 completes the operation on the element , the configuration and data cache module 200 then inputs the element through the input terminal , after the arithmetic logic module 300 completes the operation on the element , the configuration and data cache module 200 then inputs the element through the input terminal , and this process is repeated until all elements of the vector X are input through the input terminal.

[0055] The target operation sub-module includes an addition sub-module and a division sub-module. At this time, the multiplier 314 and the square root extractor 315 can be controlled not to work. For example, the power supply to the multiplier 314 and the square root extractor 315 can be controlled to stop. For another example, a multiplication register can also be set. The input terminal of the multiplication register is designed as the first input terminal of the multiplication sub-module, and the output terminal of the multiplication register is connected to the first input terminal of the multiplier 314. When the normalization function type is the first preset type, after receiving the data, the multiplication register directly deletes the data. For another example, a control signal can be sent to the control terminals of the multiplier 314 and the square root extractor 315 to stop the multiplier 314 and the square root extractor 315 from working.

[0056] After the configuration and data cache module 200 inputs an element through the input terminal, it first determines whether the first bit of the element is 0 or 1. If the first bit of the element is 1, that is, the element is negative, the first multiplexer 303 controls the conduction of the path from the output terminal of the inverter 302 to the first input terminal of the adder 304, and the output parameter of the first parameter output device 301 is 1; if the first bit of the element is 0, that is, the element is positive, the first multiplexer 303 controls the conduction of the path from the input terminal of the inverter 302 to the first input terminal of the adder 304, and the output parameter of the first parameter output device 301 is 0. Moreover, the element is also input to the third register 310.

[0057] Assume that the data to be operated on above is the vector [3, -5, 7]. The configuration and data cache module 200 will first input the data 00000011 (3 in binary) through the input terminal. At this time, the second multiplexer 306 controls the path from the first parameter output device 301 to the second input terminal of the adder 304 to be conducting. Since the first bit of 00000011 is 0, the first multiplexer 303 controls the path from the input terminal of the inverter 302 to the first input terminal of the adder 304 to be conducting. The first parameter is 0, and 00000011 is output to the first terminal of the adder 304. The second multiplexer 306 outputs 0 to the second terminal of the adder 304. The adder 304 adds 0 and 00000011 to get 00000011, and stores 00000011 in the first register 305, and then stores it in the second register 307 by the first register 305.

[0058] Meanwhile, 00000011 will be input to the left shifter 309 through the fifth multiplexer 318. The left shifter 309 shifts 00000011 to the left and outputs the shifted result to the third register 310. The third register 310 stores the data output by the left shifter 309.

[0059] After that, the configuration and data cache module 200 inputs the data 11111011 (-5 in binary) through the input terminal. At this time, the second multiplexer 306 controls the path from the first parameter output device 301 to the second input terminal of the adder 304 to be conducting. Since the first bit of 11111011 is 1, the first multiplexer 303 controls the path from the output terminal of the inverter 302 to the first input terminal of the adder 304 to be conducting. The first parameter is 1. That is to say, after the configuration and data cache module 200 inputs the data 11111011 through the input terminal, 11111011 will first be inverted by the inverter 302 to get 00000100. The adder 304 adds 00000100 and 1 to get 00000101, which is 5 in binary, thus realizing the conversion of the element -5 to 5. Then, the adder 304 outputs 00000101 to the first register 305. The second multiplexer 306 then controls the path from the output terminal of the first register 305 to the second input terminal of the adder 304 to be conducting. Further, the second register 307 outputs the stored 00000011 to the first input terminal of the adder 304, and the first register 305 outputs the stored 00000101 to the first input terminal of the adder 304. The adder 304 adds 00000011 and 00000101 to get 00001000, which is 8 in binary. The adder 304 outputs 00001000 to the first register 305, and the first register 305 then outputs 00001000 to the second register 307.

[0060] Meanwhile, the left shifter 309 shifts the above 11111011 to the left and outputs the result to the third register 310 after the shift.

[0061] Finally, the configuration and data cache module 200 inputs the data 00000111 (7 in binary) through the input terminal. At this time, the second multiplexer 306 controls the path from the first parameter output device 301 to the second input terminal of the adder 304 to be turned on. Since the first bit of 00000111 is 0, the first multiplexer 303 controls the path from the input terminal of the inverter 302 to the first input terminal of the adder 304 to be turned on, and the first parameter is 0. That is to say, after the configuration and data cache module 200 inputs the data 00000111 through the input terminal, the adder 304 adds 00000111 and 0 to get 00000111. Then, the adder 304 outputs 00000111 to the first register 305, and the second multiplexer 306 controls the path from the output terminal of the first register 305 to the second input terminal of the adder 304 to be turned on. Further, the second register 307 outputs the stored 1000 to the first input terminal of the adder 304, and the first register 305 outputs the stored 00000111 to the first input terminal of the adder 304. The adder 304 adds 00001000 and 00000111 to get 00001111, which is 15 in binary.

[0062] Meanwhile, the left shifter 309 shifts the above 00000111 to the left and outputs the result to the third register 310 after the shift.

[0063] At this time, since all elements of the data to be operated on have been input to the addition sub-module, it is determined that the 00001111 currently stored in the first register 305 is the first summation result. Thus, for each element in the data to be operated on, when the element is input to the addition sub-module, if the element is positive, the element is output to the adder, and if the element is negative, the element after taking the inverse is output to the adder 304. The adder 304 adds the output data of the first multiplexer 303 and the operation parameter output by the first parameter output device 301 to obtain the first added data. When there is no data stored in the second register 307 and there are elements of the data to be operated on that have not been input, the first added data is stored in the second register 307 via the first register 305. When there is no data stored in the second register 307 and all elements of the data to be operated on have been input to the addition sub-module, the first added data is determined to be the first summation result. When there is data stored in the second register 307, the first added data is output to the adder 304 via the first register 305 and the second multiplexer 306, and the first stored data in the second register 307 is output to the adder 304 via the first multiplexer 303. The adder 304 adds the first added data and the first stored data to obtain the second added data. If there are elements of the data to be operated on that have not been input, the second added data is stored in the second register 307 as the new first stored data. If all elements of the data to be operated on have been input to the addition sub-module, the second added data is determined to be the first summation result.

[0064] The first register 305 outputs 00001111, and the third register 310 sequentially outputs the three data stored therein. The divider 308 sequentially divides the three data output by the third register 310 by the first summation result output by the first register 305, that is, first divides 3 by the first summation result, then divides -5 by the first summation result, and then divides 7 by the first summation result. The shifter 311 shifts the data output by the divider 308 to the right to obtain the normalized operation results of 3, -5, and 7.

[0065] Suppose the above data to be operated on is a one-dimensional vector [-6]. The configuration and data cache module 200 will first input the data 11111010 (binary -6) through the input terminal. At this time, the second multiplexer 306 controls the path from the first parameter output device 301 to the second input terminal of the adder 304 to be turned on. Since the first bit of 11111010 is 1, the first multiplexer 303 controls the path from the output terminal of the inverter 302 to the first input terminal of the adder 304 to be turned on. The first parameter is 1, and the inverter 302 inverts 11111010 to obtain 00000101, and outputs 00000101 to the first terminal of the adder 304. The second multiplexer 306 outputs 1 to the second terminal of the adder 304. The adder 304 adds 1 and 00000101 to obtain 00000110, which is binary 6, and stores 00000110 in the first register 305.

[0066] Meanwhile, 1010 will be input into the left shifter 309 through the fifth multiplexer 318. The left shifter 309 shifts 11111010 to the left and outputs the shifted result to the third register 310. The third register 310 stores the data output by the left shifter 309.

[0067] Since at this time, all elements of the data to be operated on have been input into the addition sub-module, it is determined that the 00000110 currently stored in the first register 305 is the first summation result.

[0068] The first register 305 outputs 00000110, and the third register 310 outputs the data stored therein, that is, the binary form of the above vector [-6]. The divider 308 divides each element in the vector [-6] by the first summation result output by the first register 305, and shifts the data output by the divider 308 to the right to obtain the normalized operation result of the vector [-6].

[0069] Thus, L1 normalization of the data to be operated on can be achieved. Suppose the data to be operated on is = , , , …, , and taking the element as an example, the specific calculation formula can be as follows:

[0070] ,

[0071] where, is the L1 normalization of the element , , , , …, is an element in the data X to be operated, and N is the number of elements in the data X to be operated. = , which is to sum the absolute values of all elements in the data to be operated.

[0072] When the normalization function type is the second preset type, the above data to be operated is a vector.

[0073] The third multiplexer 313 controls the connection of the path from the input end to the second input end of the multiplier 314, the fourth multiplexer 316 controls the conduction of the path from the right shifter 311 to the output end, and the fifth multiplexer 318 controls the conduction of the path from the input end to the input end of the left shifter 309.

[0074] The configuration and data cache module 200 inputs the elements of the data to be operated one by one. Assuming the data to be operated is a vector , then the configuration and data cache module 200 will first input the element through the input end , after the arithmetic logic module 300 completes the operation on the element , the configuration and data cache module 200 then inputs the element through the input end , after the arithmetic logic module 300 completes the operation on the element , the configuration and data cache module 200 then inputs the element through the input end , and loops this process until all elements of the vector X are input through the input end.

[0075] The target operation sub-module includes an addition sub-module, a multiplication sub-module, a division sub-module, and a square root sub-module.

[0076] The first multiplexer 303 controls the conduction of the path from the input end of the inverter 302 to the first input end of the adder 304.

[0077] Suppose the data to be operated on above is the vector [3, -5, 7]. The configuration and data cache module 200 will first input 3 in binary form through the input terminal. At this time, the adder 304 ignores this data. For example, the adder 304 can be set not to work at this time. Another example is that an adder register can be set. The input terminal of the adder register is connected to the output terminal of the first multiplexer 303, and the output terminal of the adder register is connected to the first input terminal of the adder 304. And the adder register is set to delete this data. Another example is that a control signal can be sent to the adder 304 to make the adder 304 stop working before the multiplier 314 outputs a signal and after the calculated data is sent to the first register 305. 3 will be input to the first input terminal and the second input terminal of the multiplier 314, so that the multiplier 314 multiplies 3 by 3 to get 9 and outputs 9 to the second register 307. The second multiplexer 306 controls the conduction of the path from the first parameter output device 301 to the adder 304. The first parameter is 0. At the same time, the second register 307 outputs 9. The adder 304 adds 9 and 0 to get 9 and stores 9 in the first register 305.

[0078] At the same time, 3 in binary form will be input to the left shifter 309 through the fifth multiplexer 318. The left shifter 309 shifts this data to the left and outputs the left shift result to the third register 310. The third register 310 stores the data output by the left shifter 309.

[0079] Furthermore, the configuration and data cache module 200 inputs the data -5 in binary form through the input terminal. At this time, the adder 304 ignores this data. -5 will be input to the first input terminal and the second input terminal of the multiplier 314, so that the multiplier 314 multiplies -5 by -5 to get 25 and outputs 25 to the second register 307. The second multiplexer 306 controls the conduction of the path from the output terminal of the first register 305 to the second input terminal of the adder 304. The first register 305 outputs 9 to the second input terminal of the adder 304, and the second register 307 outputs 25 to the first input terminal of the adder 304. The adder 304 adds 25 and 9 to get 34 and stores 34 in the first register 305.

[0080] At the same time, -5 in binary form will be input to the left shifter 309. The left shifter 309 shifts this data to the left and outputs the left shift result to the third register 310. The third register 310 stores the data output by the left shifter 309.

[0081] Further, the configuration and data cache module 200 inputs the data 7 in binary form through the input end. At this time, the adder 304 ignores this data, and 7 is input to the first input end and the second input end of the multiplier 314, so that the multiplier 314 multiplies 7 by 7 to obtain 49, and outputs 49 to the second register 307. The second multiplexer 306 controls the conduction of the path from the output end of the first register 305 to the second input end of the adder 304. The first register 305 outputs 34 to the second input end of the adder 304, and the second register 307 outputs 49 to the first input end of the adder 304. The adder 304 adds 34 and 49 to obtain 83, and stores 83 in the first register 305.

[0082] Meanwhile, the binary-form 7 is input to the left shifter 309. The left shifter 309 shifts this data to the left and outputs the left-shifted result to the third register 310. The third register 310 stores the data output by the left shifter 309.

[0083] At this time, since the addition results corresponding to all elements of the data to be operated on have been stored in the first register 305, it is determined that the binary-form 83 currently stored in the first register 305 is the third addition result.

[0084] Further, the first register 305 outputs the third addition result to the input end of the square root extractor 315. The square root extractor 315 takes the square root of the third addition result to obtain the square root result.

[0085] Further, the square root extractor 315 outputs the square root result to the first input end of the divider 308. The third register 310 sequentially outputs the three data stored therein to the second input end of the divider 308. The divider 308 sequentially divides the data output by the third register 310 by the square root result to obtain the normalized operation result of the data to be operated on.

[0086] Thus, L2 normalization of the data to be operated on can be achieved. Assume the data to be operated on is = , , ,…, , and taking the element as an example, the specific calculation formula can be shown as follows:

[0087] ,

[0088] Among them, is the L2 normalization of the element , 、 、 、…、 is an element in the data X to be operated on, and N is the number of elements in the data X to be operated on. = , which is to take the square root after summing the squares of all elements in the data to be operated on.

[0089] When the normalization function type is the third preset type, the batch normalization functions include functions such as BatchNorm1d, BatchNorm2d, and BatchNorm3d. BatchNorm1d processes one-dimensional data and is suitable for fully connected layers or sequence data. The size of the input data is (batch size N, number of channels C) or (batch size N, number of channels C, length L). BatchNorm2d processes two-dimensional data and is usually used in convolutional layers. The size of the input data is (batch size N, number of channels C, length H, width W). BatchNorm3d processes three-dimensional data and is suitable for three-dimensional convolutional layers. The size of the input data is (batch size N, number of channels C, depth D, height H, width W). The output shape is the same as the input. For hardware, the number of input and output data for different functions is different each time, and the configuration and data cache module 200 needs to be used to configure the parameters properly. However, the calculation algorithm for the data is the same. Assuming the data to be operated on at this time is x, the function formulas can be unified into the following formula:

[0090] .

[0091] is the batch normalization function, E1(x), Var1(x), , are the parameters in this formula and can all be generated by training and are known. is a very small value to prevent division by zero, so it can be converted into the following formula:

[0092] .

[0093] Both k and b are the parameters in this formula.

[0094] At this time, the third multiplexer 313 controls the path connection from the second parameter output device 312 to the second input terminal of the multiplier 314, the first multiplexer 303 controls the path conduction from the input terminal of the inverter 302 to the first input terminal of the adder 304, and the second multiplexer 306 controls the path conduction from the first parameter output device 301 to the second input terminal of the adder 304.

[0095] The target operator module includes an addition sub-module and a multiplication sub-module. At this time, the division sub-module and the square root sub-module stop working. For example, the divider 308 and the square root unit 315 can be powered off. For another example, the above square root register can be set, and when the square root register and the third register 310 receive data, they can be directly deleted. For yet another example, a control signal can be sent to the divider 308 and the square root unit 315 to make the divider 308 and the square root unit 315 stop working.

[0096] The configuration and data cache module 200 sends the data to be operated through the input terminal. At this time, the adder 304 is also controlled to stop working. The specific method can refer to the above method. Assume that the data to be operated is the above x, the second parameter output device 312 outputs the operation parameter k, the first parameter output device 301 outputs the operation parameter b, the multiplier 314 receives x through its first input terminal, receives k through its second input terminal, calculates x multiplied by k, takes the calculation result as the second multiplication result, outputs the second multiplication result to the second register 307, the second register 307 outputs the second multiplication result to the adder 304, the adder 304 adds the second multiplication result and b to obtain the normalized operation result of the data to be operated, and controls the fourth multiplexer 316 to make the path from the output terminal of the first register 305 to the output terminal conductive, and outputs the normalized operation result.

[0097] When the normalization function type is the fourth preset type, the above data to be operated is a vector, and the target operator module includes an addition sub-module, a multiplication sub-module, a division sub-module, and a square root sub-module.

[0098] Assume that the above data to be input is the vector [3, -5, 7].

[0099] In the first step, the addition sub-module is used to sum all elements of the data to be operated to obtain the fourth addition result, and the division sub-module is used to divide the fourth addition result by the total number of elements in the data to be operated to obtain the first initial operation result.

[0100] Specifically, the configuration and data cache module 200 first inputs the binary data 3 through the input terminal. At this time, the second multiplexer 306 controls the path from the first parameter output device 301 to the second input terminal of the adder 304 to be conductive, the first multiplexer 303 controls the path from the input terminal of the inverter 302 to the first input terminal of the adder 304 to be conductive, the first parameter is 0, 3 is output to the first terminal of the adder 304, the second multiplexer 306 outputs 0 to the second terminal of the adder 304, the adder 304 adds 0 and 3 to obtain 3, and stores 3 in the first register 305.

[0101] After that, the configuration and data cache module 200 inputs the binary data -5 through the input terminal. At this time, the second multiplexer 306 controls the path from the output terminal of the first register 305 to the second input terminal of the adder 304 to conduct, and the first multiplexer 303 controls the path from the input terminal of the inverter 302 to the first input terminal of the adder 304 to conduct. The first parameter is 0. At this time, the adder 304 adds -5 and 3 to get -2, and the adder 304 outputs -2 to the first register 305.

[0102] Finally, the configuration and data cache module 200 inputs the binary data 7 through the input terminal. At this time, the second multiplexer 306 controls the path from the output terminal of the first register 305 to the second input terminal of the adder 304 to conduct, and the first multiplexer 303 controls the path from the input terminal of the inverter 302 to the first input terminal of the adder 304 to conduct. The adder 304 adds -2 and 7 to get 5. After that, the adder 304 outputs 5 to the first register 305.

[0103] The binary 5 is the above-mentioned fourth addition result. After the binary 5 is input to the first register 305, the first register 305 outputs the binary 5. After the data 5 is shifted left by the left shifter 309, it reaches the third register 310. Further, the fifth multiplexer 318 is controlled so that the second parameter output device outputs the number of elements 3 in the vector [3, -5, 7]. The third register 310 outputs the left-shifted data 5. The data 3 and the left-shifted data 5 are output to the two input terminals of the divider 308, so that the divider 308 divides 5 by 3, and the output data of the divider 308 is shifted right by the right shifter 311 to obtain the first initial operation result, which is the mean value of the elements in the vector.

[0104] Further, the above-mentioned first initial operation result is stored in the fourth register 320 through the sixth multiplexer 319.

[0105] In the second step, all elements of the data to be operated are multiplied by themselves to obtain a plurality of third multiplication results. The addition sub-module is further used to add all the third multiplication results to obtain a fifth addition result. The division sub-module is further used to divide the fifth addition result by the total number of elements in the data to be operated to obtain a second initial operation result.

[0106] Specifically, the configuration and data cache module 200 first inputs 3 in binary form through the input terminal. At this time, the third multiplexer 313 controls the conduction of the path from the input terminal to the second input terminal of the multiplier 314. The adder 304 ignores this data, and 3 is input to the first input terminal and the second input terminal of the multiplier 314, so that the multiplier 314 multiplies 3 by 3 to obtain 9, and outputs 9 to the second register 307. The second multiplexer 306 controls the conduction of the path from the first parameter output device 301 to the adder 304. The first parameter is 0, and at the same time the second register 307 outputs 9. The adder 304 adds 9 and 0 to obtain 9, and stores 9 in the first register 305.

[0107] Further, the configuration and data cache module 200 inputs the binary data -5 in binary form through the input terminal. At this time, the adder 304 ignores this data, and -5 is input to the first input terminal and the second input terminal of the multiplier 314, so that the multiplier 314 multiplies -5 by -5 to obtain 25, and outputs 25 to the second register 307. The second multiplexer 306 controls the conduction of the path from the output terminal of the first register 305 to the second input terminal of the adder 304. The first register 305 outputs 9 to the second input terminal of the adder 304, and the second register 307 outputs 25 to the first input terminal of the adder 304. The adder 304 adds 25 and 9 to obtain 34, and stores 34 in the first register 305.

[0108] Further, the configuration and data cache module 200 inputs the binary data 7 in binary form through the input terminal. At this time, the adder 304 ignores this data, and 7 is input to the first input terminal and the second input terminal of the multiplier 314, so that the multiplier 314 multiplies 7 by 7 to obtain 49, and outputs 49 to the second register 307. The second multiplexer 306 controls the conduction of the path from the output terminal of the first register 305 to the second input terminal of the adder 304. The first register 305 outputs 34 to the second input terminal of the adder 304, and the second register 307 outputs 49 to the first input terminal of the adder 304. The adder 304 adds 34 and 49 to obtain 83, and stores 83 in the first register 305.

[0109] The 83 in binary is the above-mentioned fifth addition result. The first register 305 outputs the data 83 stored therein. The binary data 83 reaches the third register 310 after being shifted left by the left shifter 309. The third register 310 outputs the data stored therein, and the fifth multiplexer enables the second parameter output device 312 to output the number 3 of elements in the binary vector [3, -5, 7]. The shifted data 83 and the data 3 are output to the two input terminals of the divider 308. The divider 308 divides 83 by 3 and shifts the data output by the divider 308 to the right by the right shifter 311 to obtain the second initial operation result.

[0110] Moreover, the above-mentioned second initial operation result is also stored in the fourth register 320 through the sixth multiplexer 319.

[0111] In the third step, take the square root of the second initial operation result to obtain the square root result. The addition sub-module is further used to add the first initial operation result after taking the inverse to each element in the data to be operated to obtain the sixth addition result. The division sub-module is further used to divide the sixth addition result by the square root result to obtain the third initial operation result.

[0112] Specifically, the fourth register 320 outputs the second initial operation result. The fifth multiplexer 318 controls the conduction of the path from the fourth register 320 to the output terminal of the fifth multiplexer 318. The adder 304 ignores the received data, and the third register 310 deletes the received data. The square rooter 315 takes the square root of the second initial operation result to obtain the square root result.

[0113] The fourth register 320 outputs the first initial operation result. At this time, the multiplier 314 ignores this data. The first multiplexer 303 controls the conduction of the path from the output terminal of the inverter 302 to the first input terminal of the adder 304. After the first initial operation result is inverted by the inverter 302, it is input to the first input terminal of the adder 304. At the same time, the first parameter output device 301 outputs 1 to the second input terminal of the adder 304. The adder 304 adds the first initial operation result and 1 and stores the added result in the first register 305.

[0114] After that, the first multiplexer 303 controls the conduction of the path from the input terminal of the inverter 302 to the first input terminal of the adder 304, and the second multiplexer 306 controls the conduction of the path from the output terminal of the first register 305 to the second input terminal of the adder 304. By inputting the element 3 in the data vector [3, -5, 7] to be operated to the first input terminal of the adder 304, the first register 305 outputs the data stored therein. The adder 304 adds the data received at the first input terminal and the data received at the second input terminal to obtain a sixth addition result and stores the sixth addition result in the first register 305.

[0115] The first register 305 outputs data to the second input terminal of the divider, and the square root result is output to the first input terminal of the divider 308. The divider 308 divides the data output by the first register 305 by the data output by the third register 310. The right shifter 311 right-shifts the data output by the divider 308 to obtain a third initial operation result.

[0116] For each element in the data vector to be operated on [3, -5, 7], the operation is performed according to the above method, and a total of three third initial operation results are obtained.

[0117] Further, the above three third initial operation results are stored in the fourth register 320 via the sixth multiplexer 319.

[0118] In the fourth step, the third initial operation result is multiplied by the operation parameter output by the second parameter output device 312 to obtain a fourth multiplication result. The adder sub-module is further configured to add the fourth multiplication result to the operation parameter output by the first parameter output device 301 to obtain a normalized operation result of the data to be operated on.

[0119] Specifically, the third multiplexer 313 controls the conduction of the path from the output terminal of the second parameter output device 312 to the multiplier 314. The first multiplexer 303 controls the conduction of the path from the input terminal of the inverter 302 to the adder 304. The second multiplexer 306 controls the conduction of the path from the first parameter output device 301 to the adder 304. The fourth multiplexer 316 controls the conduction of the path from the output terminal of the first register 305 to the output terminal. The fourth register 320 outputs a third initial operation result, and the adder 304 ignores this data. The multiplier 314 multiplies the third initial operation result by the operation parameter output by the second parameter output device 312 to obtain a fourth multiplication result, and outputs the fourth multiplication result to the second register 307. The second register 307 outputs the fourth multiplication result, and the adder 304 adds the fourth multiplication result to the operation parameter output by the first parameter output device 301 to obtain a normalized operation result corresponding to this third initial operation result.

[0120] The above processing is performed on all the three third initial operation results to obtain a normalized operation result of the data to be operated on [3, -5, 7].

[0121] Thus, layer normalization of the data to be operated on can be achieved. Assume the data to be operated on is = , , , …, , and taking the element as an example, the specific calculation formula can be as follows:

[0122] ,

[0123] Among them, is to perform L2 normalization on the element . Since is a minimum value, it can be ignored during the operation. is the above-mentioned second initial operation result, which is the sum of the self-multiplications of all elements of the data to be operated divided by the number of elements. is the above-mentioned first initial operation result, which is the mean value of the elements in the data to be operated. is the operation parameter output by the second parameter output device 312. is the operation parameter output by the first parameter output device 301.

[0124] It should be noted that the operation parameters in the above-mentioned first parameter output device 301 and second parameter output device 312 are the parameters written by the configuration and data cache module 200. For example, it can be set that the above-mentioned first parameter output device 301 is implemented by multiple registers and a multiplexer. The number of registers is the same as the number of data that the first parameter output device 301 may output, and the corresponding operation parameter is selected and output by the multiplexer.

[0125] It should also be noted that since neural network quantization is an optimization technology, it reduces the storage requirements and computational complexity of the model by converting parameters such as the weights and activations of the model from floating-point numbers (such as single-precision floating-point number Float32, half-precision floating-point number FLoat16) to integer (such as Int8) form, thereby accelerating the inference speed of the model and reducing power consumption. As shown in the following formula, it is the formula for quantizing the parameter Weight from Float32 to INT8 new_wegiht and the corresponding power exponent exp of IN8. The exp of each layer of the neural network is unchanged.

[0126] .

[0127] Using the quantized data and inputting it into the normalization function, since the normalization functions of L1 normalization, L2 normalization, and layer normalization contain integer division, and normalization is to convert the data into a value between (0,1), the output will be 0, seriously affecting the accuracy. And converting the data from integer to floating-point requires additional hardware circuits, resulting in problems such as excessive area and waste of resources.

[0128] To solve this problem, the above-mentioned left shifter 309 and right shifter 311 are provided. That is, first shift the numerator of the dividend to the left, for example, shift it 10 bits to the left. The number of bits shifted to the left is determined according to the precision requirement. After the division is completed, then shift it to the right according to the power exponent exp of the data at this layer. Compared with the original integer method, the precision is improved, and effective data can be obtained, and the precision loss of the result obtained by the floating-point operation is not large. Since the above adder 304 is an integer adder, the above multiplier 314 is an integer multiplier, the above divider 308 is an integer divider, and the above square root extractor 315 is an integer square root extractor, this circuit uses integer calculations instead of floating-point calculations, reducing power consumption and area within an acceptable precision loss range and improving performance.

[0129] In summary, the normalization operation circuit of the embodiment of the present invention includes: a configuration and data cache module, which is used to determine the data to be operated, the normalization function type, and the operation parameters; an arithmetic logic module, which is connected to the configuration and data cache module. The arithmetic logic module includes multiple operation sub-modules, which are used to determine the target operation sub-module from multiple operation sub-modules according to the normalization function type, and call the target operation sub-module to perform normalization operation on the data to be operated according to the normalization function type and operation parameters. Among them, the multiple operation sub-modules include an addition sub-module, a multiplication sub-module, a division sub-module, and a square root extraction sub-module. Thus, by using the additionally provided arithmetic logic module for normalization operation, no device parameters such as CPU and GPU are required, thereby meeting the requirements of edge computing.

[0130] Furthermore, the present invention proposes a normalization operation system.

[0131] Figure 4 It is the structural block diagram of the normalization operation system of the embodiment of the present invention.

[0132] As Figure 4 shown, the normalization operation system 10 includes the above-mentioned normalization operation circuit 100.

[0133] The normalization operation system of the embodiment of the present invention can implement normalization operation without the participation of device parameters such as CPU and GPU through the normalization operation circuit of the above embodiment, thereby meeting the requirements of edge computing.

[0134] It should be noted that the logic and / or steps represented in the flowchart or described otherwise herein can be considered as a definite sequence list of executable instructions for implementing logical functions, which can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in combination with these instruction execution systems, apparatuses, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in combination with an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of the computer-readable medium include the following: an electrical connection portion with one or more wirings (electronic device), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or other suitable processing when necessary, and then stored in a computer memory.

[0135] It should be understood that various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. If implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGA), field programmable gate arrays (FPGA), and the like.

[0136] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.

[0137] In the description of this specification, the orientation or positional relationship indicated by terms such as "center", "longitudinal", "transverse", "length", "width", "thickness", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", "clockwise", "counterclockwise", "axial", "radial", "circumferential", etc. is based on the orientation or positional relationship shown in the drawings, and does not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and should not be construed as a limitation on the present invention.

[0138] In addition, the terms "first" and "second" are used only for descriptive purposes and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of the present invention, the meaning of "a plurality of" is at least two, such as two, three, etc., unless otherwise specifically and clearly defined.

[0139] In the description of this specification, unless otherwise stated, terms such as "mounted", "connected", "joined", "fixed", etc. should be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or integrated; it may be a mechanical connection or an electrical connection; it may be directly connected or indirectly connected through an intermediate medium, and it may be the communication inside two elements or the interaction relationship between two elements, unless otherwise clearly defined. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0140] In the present invention, unless otherwise clearly specified and limited, the first feature being "on" or "under" the second feature may be that the first and second features are in direct contact, or the first and second features are indirectly in contact through an intermediate medium. Moreover, the first feature being "above", "over" and "on top of" the second feature may be that the first feature is directly above or obliquely above the second feature, or merely indicates that the first feature has a higher horizontal height than the second feature. The first feature being "under", "beneath" and "underneath" the second feature may be that the first feature is directly below or obliquely below the second feature, or merely indicates that the first feature has a lower horizontal height than the second feature.

[0141] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as a limitation on the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention.

Claims

1. A normalization operation circuit, characterized in that: The circuit comprises: Configuration and data cache module, used to determine the data to be calculated, the normalization function type and the calculation parameters; an arithmetic logic module connected to the configuration and data cache module, the arithmetic logic module comprising a plurality of operator modules, for determining a target operator module from the plurality of operator modules according to the normalization function type, and calling the target operator module to perform a normalization operation on the data to be operated according to the normalization function type and the operation parameters, wherein the plurality of operator modules comprise an addition submodule, a multiplication submodule, a division submodule and a square root submodule; The first input end of the addition submodule is designed as the input end of the arithmetic logic module, and is used to receive the data to be operated. The output end of the addition submodule is connected to the first input end of the division submodule, the second input end of the division submodule is connected to the first input end of the addition submodule, and the output end of the division submodule is designed as the output end of the arithmetic logic module; wherein, When the normalization function type is the first preset type, the data to be operated is a vector, and the target operation submodule includes the addition submodule and the division submodule. The addition submodule is used to sum the absolute values ​​of all elements of the data to be operated to obtain a first summation result, and the division submodule is used to divide the data to be operated by the first summation result to obtain a normalized operation result of the data to be operated.

2. The normalization operation circuit according to claim 1, characterized in that: The adding submodule comprises: a first parameter output device, an inverter, a first multiplexer, an adder, a first register, a second multiplexer, and a second register; wherein the input end of the inverter is designed as the first input end of the adding submodule; the first input end of the first multiplexer is connected to the output end of the inverter, and the second input end of the first multiplexer is connected to the input end of the inverter; the first input end of the adder is connected to the output end of the first multiplexer; the input end of the first register is connected to the output end of the adder; the first input end of the second multiplexer is connected to the output end of the first parameter output device, the second input end of the second multiplexer is connected to the output end of the first register, and the output end of the second multiplexer is connected to the second input end of the adder; the input end of the second register is connected to the output end of the first register, and the output end of the second register is connected to the input end of the inverter; the output end of the first register is also set as the output end of the adding submodule; The division submodule includes: a divider, a left shifter, a third register, and a right shifter; wherein the first input end of the divider is designed as the first input end of the division submodule, and the second input end of the divider is connected to the output end of the third register; the input end of the left shifter is designed as the second input end of the division submodule; the input end of the third register is connected to the output end of the left shifter; the input end of the right shifter is connected to the output end of the divider, and the output end of the right shifter is designed as the output end of the division submodule.

3. The normalization operation circuit according to claim 2, characterized in that: When the normalization function type is the first preset type, for each element in the data to be operated, when the element is input into the addition submodule, the first multiplexer outputs the element to the adder when the element is a positive number, and outputs the element inverted by the inverter to the adder when the element is a negative number, and the adder adds the output data of the first multiplexer to the operation parameter output by the first parameter output device to obtain the first added data; when no data is stored in the second register and there are elements that are not input into the data to be operated, the first added data is stored in the second register via the first register; when no data is stored in the second register and the data to be operated is When all elements of the data to be calculated have been input into the addition submodule, it is determined that the first added data is the first summation result; when data has been stored in the second register, the first added data is output to the adder via the first register and the second multiplexer, and the first stored data in the second register is output to the adder via the first multiplexer, and the adder adds the first added data to the first stored data to obtain second added data; if there are elements that are not input into the data to be calculated, the second added data is stored as new first stored data in the second register; if all elements of the data to be calculated have been input into the addition submodule, it is determined that the second added data is the first summation result; The left shifter shifts the data to be calculated leftwards, the divider divides the left-shifted data to be calculated by the first summation result, and the right shifter shifts the output data of the divider rightwards to obtain a normalized calculation result of the data to be calculated.

4. The normalization operation circuit according to claim 2, characterized in that: The input end of the second register is designed as the second input end of the addition submodule, the first input end and the second input end of the multiplication submodule are both connected to the first input end of the addition submodule, the output end of the multiplication submodule is connected to the second input end of the addition submodule, the input end of the square root submodule is connected to the output end of the addition submodule, and the output end of the square root submodule is connected to the first input end of the division submodule; The multiplication submodule comprises: a second parameter output device, a third multiplexer, and a multiplier; wherein the first input end of the third multiplexer is connected to the output end of the second parameter output device, and the second input end of the third multiplexer is designed as the second input end of the multiplication submodule; the first input end of the multiplier is designed as the first input end of the multiplication submodule, the second input end of the multiplier is connected to the output end of the third multiplexer, and the output end of the multiplier is designed as the output end of the multiplication submodule; The square root submodule includes a square root generator, the input end of the square root generator is designed as the input end of the square root submodule, and the output end of the square root generator is designed as the output end of the square root submodule.

5. The normalization operation circuit according to claim 4, characterized in that: When the normalization function type is the second preset type, the data to be operated is a vector, and the target operation submodule includes the addition submodule, the multiplication submodule, the division submodule and the square root submodule; Wherein, for each element in the data to be operated, the element is input to the first input terminal and the second input terminal of the multiplier, so that the multiplier multiplies the element with itself to obtain a first multiplication result, and the first multiplication result is stored in the second register; if the first register does not store data, the second register outputs the first multiplication result to the adder through the first multiplexer, and the adder adds the first multiplication result to the operation parameter output by the first parameter output device, and stores the addition result in the first register; if the first register has data stored, the first register outputs the data stored therein to the adder through the second multiplexer, and the adder adds the data stored in the first register to the first multiplication result, and stores the addition result in the first register; After the addition results corresponding to each element in the data to be calculated are stored in the first register, the data in the first register is used as the third addition result, and the third addition result is output to the square root generator. The square root generator takes the square root of the third addition result to obtain a square root result, and outputs the square root result to the divider. The divider divides the data to be calculated by the square root result to obtain a normalized calculation result of the data to be calculated.

6. The normalization operation circuit according to claim 4, characterized in that: The arithmetic logic module also includes a fourth multiplexer, a first input end of the fourth multiplexer is connected to the output end of the right shifter, an output end of the fourth multiplexer is connected to the output end of the first register, and an output end of the fourth multiplexer is designed to be the output end of the arithmetic logic module.

7. The normalization operation circuit according to claim 6, characterized in that: When the normalization function type is a third preset type, the target operator module includes the addition submodule and the multiplication submodule; Among them, the third multiplexer outputs the operation parameters output by the second parameter output device to the multiplier. After receiving the data to be operated, the multiplier multiplies the data to be operated by the operation parameters output by the second parameter output device to obtain a second multiplication result, and outputs the second multiplication result to the adder through the second register and the first multiplexer. The adder adds the second multiplication result and the operation parameters output by the first parameter output device to obtain a normalized operation result of the data to be operated.

8. The normalization operation circuit according to claim 4, characterized in that: The input end of the square root submodule is also connected to the input end of the addition submodule and the second input end of the division submodule. When the normalization function type is the fourth preset type, the data to be operated is a vector, and the target operation submodule includes the addition submodule, the multiplication submodule, the division submodule and the square root submodule. Among them, the addition submodule is used to sum all elements of the data to be calculated to obtain a fourth addition result, and the division submodule is used to divide the fourth addition result by the total number of elements in the data to be calculated to obtain a first initial calculation result; the multiplication submodule is used to multiply all elements of the data to be calculated by themselves to obtain multiple third multiplication results, and the addition submodule is also used to add all the third multiplication results to obtain a fifth addition result, and the division submodule is also used to divide the fifth addition result by the total number of elements in the data to be calculated to obtain a second initial calculation result; the square root submodule is used to multiply the third multiplication results by all the third multiplication results. The second initial operation result is square rooted to obtain a square root result. The addition submodule is also used to invert the first initial operation result and add it to each element in the data to be operated to obtain a sixth addition result. The division submodule is also used to divide the sixth addition result by the square root result to obtain a third initial operation result. The multiplication submodule is also used to multiply the third initial operation result with the operation parameter output by the second parameter output device to obtain a fourth multiplication result. The addition submodule is also used to add the fourth multiplication result with the operation parameter output by the first parameter output device to obtain a normalized operation result of the data to be operated.

9. A normalization operation system, characterized in that: The method comprises a normalization operation circuit according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Resource reuse type neural network hardware acceleration circuit based on fast convolution

    CN112862091A

  • Calibration method and device of analog circuit for executing neural network calculation

    CN114819051A