A data processing device, method and related equipment
By designing a scalar splitter and computing module in the data processing device, combined with the use of bucket memory, the problem of large storage resource overhead in multi-scalar multiplication is solved, and efficient computing and storage is achieved.
Patent Information
- Application Number
- CN202510161418.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-13
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-02-13
AI Technical Summary
The prior art requires huge memory when performing multiscalar multiplication, resulting in large overhead of storage resources.
A data processing device is designed, including a scalar splitter, a computing module and a bucket memory. By dividing the scalar of the to-processed data from low to high into multiple sub-scalars, and accumulating and storing the points accumulation results in each segmentation stage, multiplying with the depth address as the weight, the multiscalar multiplication results are gradually calculated, and the iteration results are stored in multiple depth addresses.
Effectively reduces the waste of storage resources, avoids additional memory requirements, and simplifies the reading and computing process of the next segmentation stage.
Smart Images

Figure CN119646377B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of computer technology, and more specifically, to a data processing apparatus, method, and related equipment. Background Art
[0002] Multi-scalar multiplication is widely used in fields such as zero-knowledge proof and elliptic curve cryptography. Multi-scalar multiplication refers to the following operations:
[0003]
[0004] in, is a vector of points on the elliptic curve, is the scalar corresponding to each point on the elliptic curve, is the result of multi-scalar multiplication, and n represents n points on the elliptic curve. The bit width is m, which can reach 200 bits or more, that is, The bit width is large, so the amount of calculation is large. Split to facilitate calculation, split as follows:
[0005]
[0006] Among them, m is The bit width of , c is the bit width of each sub-scalar after splitting, and r is the number of sub-scalars after the scalar with the highest bit width is split.
[0007] The prior art usually adopts a bucket partitioning method. For sub-scalars of the same order in each scalar, each point vector with a sub-scalar of k is first accumulated, and then the accumulated result is multiplied with the corresponding sub-scalar and stored in each storage bucket; then the accumulated results in each storage bucket are accumulated to obtain a sub-accumulated result; then each sub-accumulated result is multiplied by cj times, where j is the partitioning stage of the multi-scalar multiplication; finally, each sub-accumulated result that has been multiplied by cj times is accumulated to obtain the calculation result of the multi-scalar multiplication.
[0008] In this process, when accumulating the vectors of each point with subscalar k, a depth of A memory of depth of When the accumulated results in each bucket are accumulated to obtain a sub-accumulated result, a memory with a depth of 1 is required. When each sub-accumulated result is multiplied by cj, a memory with a depth of In this calculation process, especially when the calculation results of each block are multiplied, a deep memory is required, which will cause huge storage resource overhead. Summary of the invention
[0009] The embodiments of the present application provide a data processing apparatus, method and related equipment, aiming to solve the problem in the prior art that a huge depth of memory is required when performing multi-scalar multiplication.
[0010] In a first aspect, an embodiment of the present application provides a data processing device, the device comprising: a scalar splitter, a computing module, and a bucket memory;
[0011] The scalar divider is used to receive the data to be processed, and divide the scalar of each point vector in the data to be processed from low to high bits into multiple sub-scalars according to a preset bit width;
[0012] The computing module is used for:
[0013] In each segmentation stage, based on the point vectors corresponding to the sub-scalars of each data size, the accumulation results of each point in each segmentation stage are accumulated respectively; the accumulation results of each point in each segmentation stage are written into the corresponding bucket of the bucket memory; wherein, the accumulation results of each point in the same segmentation stage are located in the same bucket, and the depth address of each point accumulation result in the bucket is the data size of the sub-scalar corresponding to each point accumulation result;
[0014] According to the order of the segmentation stages from high to low, the depth address corresponding to the accumulated result of each point in each segmentation stage is used as the weight of the accumulated result of each point, and the accumulated result of each point is multiplied by the corresponding weight, and the sub-accumulated result of each segmentation stage is accumulated;
[0015] According to the order of the segmentation stages from high to low, the iterative result of each segmentation stage is determined, and the iterative result of each segmentation stage is stored in the bucket corresponding to the bucket memory, and the iterative result of each segmentation stage is stored in the storage space of at least two different depth addresses in the corresponding bucket, and the sum of the at least two different depth addresses is , until the iterative result of the last segmentation stage is obtained, and the iterative result of the last segmentation stage is used as the calculation result of the multi-scalar multiplication of the data to be processed; c is the bit width of the sub-scalar;
[0016] Among them, the iterative result of each segmentation stage is obtained by multiplying the iterative result of the previous segmentation stage by its corresponding weight, and then accumulating it with the sub-accumulated result of the current segmentation stage; the weight of the iterative result of each segmentation stage is the depth address stored in the corresponding bucket, and the iterative result of the highest-order segmentation stage is its corresponding sub-accumulated result.
[0017] In the embodiment of the present application, the operation module includes a first state machine and a first multiplier-adder;
[0018] The scalar divider is used to send each sub-scalar and the corresponding point vector obtained in each division stage to the first state machine;
[0019] The first state machine is used to receive the sub-scalars and corresponding point vectors obtained in each segmentation stage, and in each segmentation stage, send the point vectors corresponding to the sub-scalars of each data size to the first multiplier-adder respectively;
[0020] The first multiplier-adder is used to receive point vectors corresponding to sub-scalars of each data size in each segmentation stage, and to accumulate all point vectors corresponding to sub-scalars of the same data size to obtain accumulation results of each point in each segmentation stage.
[0021] In an embodiment of the present application, the first multiplier is also used to: write the accumulated results of each point in each segmentation stage into the corresponding bucket of the bucket memory; wherein the accumulated results of each point in the same segmentation stage are located in the same bucket, and the depth address of each point accumulated result in the bucket is the data size of the sub-scalar corresponding to the accumulated result of each point.
[0022] In the embodiment of the present application, the operation module further includes a second state machine and a second multiplier-adder;
[0023] The second state machine is used to: read the accumulated results of each point in each segmentation stage and the corresponding depth address from the bucket corresponding to each segmentation stage in the order of the segmentation stages from high to low, and send the accumulated results of each point in each segmentation stage and the corresponding depth address to the second multiplier;
[0024] The second multiplier is used to: receive the accumulated results of each point in each segmentation stage and the corresponding depth address in the order of the segmentation stages from high to low; and use the depth address corresponding to the accumulated results of each point in each segmentation stage as the weight of the accumulated results of each point, multiply the accumulated results of each point with their corresponding weights, and accumulate to obtain the sub-accumulated results of each segmentation stage.
[0025] In the embodiment of the present application, the second multiplier-adder is further used to store the sub-accumulation result of each segmentation stage in the bucket corresponding to each segmentation stage.
[0026] In the embodiment of the present application, the operation module further includes a third state machine and a third multiplier-adder;
[0027] The third state machine is used to: read the iteration result and the corresponding depth address of the previous segmentation stage, and the sub-accumulation result of the current segmentation stage from the bucket corresponding to each segmentation stage in the order of the segmentation stages from high to low, and send the iteration result and the corresponding depth address of the previous segmentation stage, and the sub-accumulation result of the current segmentation stage to the third multiplier;
[0028] The third multiplier-adder is used to use the address depth of the iteration result of the previous segmentation stage as a weight, multiply it with the iteration result of the previous segmentation stage, and then accumulate it with the sub-accumulation result of the current segmentation stage to obtain the iteration result of the current segmentation stage according to the segmentation stage order from high to low;
[0029] The iteration result of the current segmentation stage is stored in the storage space of at least two different depth addresses of the bucket corresponding to each segmentation stage, and the sum of the at least two different depth addresses is , where c is the bit width of the sub-scalar, and the corresponding iteration result in the segmentation stage of the highest bit is its corresponding sub-accumulation result.
[0030] In a second aspect, an embodiment of the present application provides a data processing method, which is applied to the data processing device described in any of the above embodiments, and the method includes:
[0031] Receive data to be processed, and divide the scalar of each point vector in the data to be processed into multiple sub-scalars from low to high bits according to a preset bit width;
[0032] In each segmentation stage, the point vectors corresponding to the sub-scalars of each data size are accumulated to obtain the accumulation results of each point in each segmentation stage; the accumulation results of each point in each segmentation stage are written into the corresponding bucket of the bucket memory; wherein, the accumulation results of each point in the same segmentation stage are located in the same bucket, and the depth address of each point accumulation result in the bucket is the data size of the sub-scalar corresponding to each point accumulation result;
[0033] According to the order of the segmentation stages from high to low, the depth address corresponding to the accumulated result of each point in each segmentation stage is used as the weight of the accumulated result of each point, and the accumulated result of each point is multiplied by the corresponding weight, and the sub-accumulated result of each segmentation stage is accumulated;
[0034] According to the order of the segmentation stages from high to low, the iterative result of each segmentation stage is determined, and the iterative result of each segmentation stage is stored in the bucket corresponding to the bucket memory, and the iterative result of each segmentation stage is stored in the storage space of at least two different depth addresses of the corresponding bucket, and the sum of the at least two different depth addresses is , until the iterative result of the last segmentation stage is obtained, and the iterative result of the last segmentation stage is used as the calculation result of the multi-scalar multiplication of the data to be processed; c is the bit width of the sub-scalar;
[0035] Among them, the iterative result of each segmentation stage is obtained by multiplying the iterative result of the previous segmentation stage by its corresponding weight, and then accumulating it with the sub-accumulated result of the current segmentation stage; the weight of the iterative result of each segmentation stage is the depth address stored in the corresponding bucket, and the iterative result of the highest-order segmentation stage is its corresponding sub-accumulated result.
[0036] In the embodiment of the present application, after obtaining the sub-scalars in each segmentation stage, the data processing method further includes:
[0037] Sending each sub-scalar and the corresponding point vector obtained in each segmentation stage to the first state machine;
[0038] In each segmentation stage, point vectors corresponding to sub-scalars of each data size are sent to the first multiplier-adder based on the first state machine;
[0039] In each segmentation stage, all point vectors corresponding to sub-scalars of the same data size are accumulated based on the first multiplier to obtain accumulation results of each point in each segmentation stage.
[0040] In the embodiment of the present application, after obtaining the accumulation results of each point in each segmentation stage, the data processing method further includes:
[0041] The accumulated results of each point in each segmentation stage are written into the corresponding bucket of the bucket memory; wherein, the accumulated results of each point in the same segmentation stage are located in the same bucket, and the depth address of each accumulated result of each point in the bucket is the data size of the sub-scalar corresponding to the accumulated result of each point.
[0042] In the embodiment of the present application, after the point accumulation results of each segmentation stage are written into the corresponding bucket, the data processing method further includes:
[0043] Read the accumulated results of each point in each segmentation stage and the corresponding depth address from the bucket corresponding to each segmentation stage in the order of the segmentation stages from high to low, and send the accumulated results of each point in each segmentation stage and the corresponding depth address to the second multiplier-adder;
[0044] According to the order of the segmentation stages from high to low, based on the second multiplier, the depth address corresponding to the accumulation result of each point in each segmentation stage is used as the weight of the accumulation result of each point, and the accumulation result of each point is multiplied by the corresponding weight to obtain the sub-accumulation result of each segmentation stage.
[0045] In the embodiment of the present application, after obtaining the sub-accumulation result of each segmentation stage, the data processing method further includes:
[0046] Based on the second multiplier-adder, the sub-accumulation result of each segmentation stage is stored in the bucket corresponding to each segmentation stage.
[0047] In the embodiment of the present application, after storing the sub-accumulated results of each segmentation stage in the bucket corresponding to each segmentation stage, the data processing method further includes:
[0048] According to the order of the segmentation stages from high to low, read the iteration result and the corresponding depth address of the previous segmentation stage, and the sub-accumulation result of the current segmentation stage from the bucket corresponding to each segmentation stage, and send the iteration result and the corresponding depth address of the previous segmentation stage, and the sub-accumulation result of the current segmentation stage to the third multiplier-adder;
[0049] According to the order of the segmentation stages from high to low, the third multiplier-adder uses the address depth of the iteration result of the previous segmentation stage as a weight, multiplies it with the iteration result of the previous segmentation stage, and then accumulates it with the sub-accumulation result of the current segmentation stage to obtain the iteration result of the current segmentation stage;
[0050] The iteration result of the current segmentation stage is stored in the storage space of at least two different depth addresses of the bucket corresponding to each segmentation stage, and the sum of the at least two different depth addresses is , where c is the bit width of the sub-scalar, and the corresponding iteration result in the segmentation stage of the highest bit is its corresponding sub-accumulation result.
[0051] In a third aspect, an embodiment of the present application provides a computer-readable storage medium, which includes instructions, which, when executed on a computer, enables the computer to execute the data processing method described in the second aspect.
[0052] In a fourth aspect, an embodiment of the present application provides a chip, which includes the data processing device described in any one of the above embodiments.
[0053] In the embodiment of the present application, by setting a calculation module and a bucket memory, after calculating the sub-accumulated result of each segmentation stage, the iterative result of each segmentation stage is obtained one by one from high to low, and the iterative result of each segmentation stage is stored in the storage space of at least two different depth addresses of the corresponding bucket, and the sum of the at least two depth addresses is , then the depth of the bucket is no greater than -1, therefore, the iterative result of each segmentation stage can be stored in the bucket used to store the accumulated result of each segmentation stage point without referencing additional memory, which can avoid the waste of storage space and facilitate the reading and calculation of the next segmentation stage. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] By reading the detailed description of the embodiments of the present application with reference to the accompanying drawings, the objectives, features and advantages of the embodiments of the present application will become easy to understand. Among them:
[0055] Figure 1 A module diagram of a data processing device in an embodiment of the present application;
[0056] Figure 2 This is a module diagram of a data processing device in another embodiment of the present application;
[0057] Figure 3 A step diagram of a data processing method in an embodiment of the present application;
[0058] Figure 4 A schematic diagram of the structure of a computer-readable storage medium according to an embodiment of the present application.
[0059] In the drawings, the same or corresponding reference numerals represent the same or corresponding parts. DETAILED DESCRIPTION
[0060] The terms "first", "second", etc. in the specification and claims of the embodiments of the present application and the above-mentioned drawings are used to distinguish similar objects (for example, the first storage address and the second storage address are respectively represented as storage addresses of different depths, and the others are similar), and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device including a series of steps or modules is not necessarily limited to those steps or modules clearly listed, but may include other steps or modules that are not clearly listed or inherent to these processes, methods, products or devices. The division of modules that appear in the embodiments of the present application is only a logical division, and there may be other division methods when implemented in actual applications, for example, multiple modules can be combined into or integrated in another system, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces, indirect coupling between modules, and the communication connection may be electrical or other similar forms, which are not limited in the embodiments of the present application. Moreover, the modules or submodules described as separate components may or may not be physically separated, may or may not be physical modules, or may be distributed in multiple circuit modules, and some or all of the modules may be selected according to actual needs to achieve the purpose of the embodiments of the present application.
[0061] Exemplary Devices
[0062] Reference Figure 1 , Figure 1 A block diagram of a data processing device 100 provided in an embodiment of the present application, wherein the data processing device 100 includes: a scalar divider 110, a computing module 120 and a bucket memory 130;
[0063] The scalar divider 110 is used to receive the data to be processed, and divide the scalar of each point vector in the data to be processed from low to high bits into multiple sub-scalars according to a preset bit width;
[0064] The computing module 120 is used for:
[0065] In each segmentation stage, the point vectors corresponding to the sub-scalars of each data size are accumulated to obtain the accumulation results of each point in each segmentation stage; the accumulation results of each point in each segmentation stage are written into the corresponding bucket of the bucket memory 130; wherein, the accumulation results of each point in the same segmentation stage are located in the same bucket, and the depth address of each point accumulation result in the bucket is the data size of the sub-scalar corresponding to each point accumulation result;
[0066] According to the order of the segmentation stages from high to low, the depth address corresponding to the accumulated result of each point in each segmentation stage is used as the weight of the accumulated result of each point, and the accumulated result of each point is multiplied by the corresponding weight, and the sub-accumulated result of each segmentation stage is accumulated;
[0067] According to the order of the segmentation stages from high to low, the iteration result of each segmentation stage is determined, and the iteration result of each segmentation stage is stored in the bucket corresponding to the bucket memory 130, and the iteration result of each segmentation stage is stored in the storage space of at least two different depth addresses in the corresponding bucket, and the sum of the at least two different depth addresses is , until the iterative result of the last segmentation stage is obtained, and the iterative result of the last segmentation stage is used as the calculation result of the multi-scalar multiplication of the data to be processed; c is the bit width of the sub-scalar;
[0068] Among them, the iterative result of each segmentation stage is obtained by multiplying the iterative result of the previous segmentation stage by its corresponding weight, and then accumulating it with the sub-accumulated result of the current segmentation stage; the weight of the iterative result of each segmentation stage is the depth address stored in the corresponding bucket, and the iterative result of the highest-order segmentation stage is its corresponding sub-accumulated result.
[0069] In an embodiment of the present application, the data to be processed may be encrypted data involved in fields such as zero-knowledge proof or elliptic curve cryptography, which includes a point vector and a corresponding scalar. There are multiple point vectors to be processed, and the multiple point vectors are located at points on the same elliptic curve.
[0070] Assumptions is the point vector to be processed, is a point vector The corresponding scalar, i∈[1,n], represents that there are n point vectors for multi-scalar multiplication in the data to be processed, and has corresponding n scalars.
[0071] Among them, when the data to be processed is subjected to multi-scalar multiplication, it can be expressed by the following formula (1):
[0072]
[0073] in, represents the result at segmentation stage j, Represents the result after accumulating the results of each segmentation stage, that is, the final calculation result of multi-scalar multiplication of the data to be processed.
[0074] In the application embodiment, after receiving the data to be processed, the scalar divider 110 can divide each scalar from low bits to high bits into multiple sub-scalars according to the preset bit width.
[0075] In the embodiment of the present application, each scalar is binary coded, such as 1110001, 10001, 1010, etc. The above is only an example, and the bit width of the scalar actually involved in the calculation is not limited to this.
[0076] Usually, the bit width of the scalar involved in the calculation is large and the calculation process is relatively complicated. Therefore, in an embodiment of the present application, the scalars involved in the calculation can be divided into scalars with multiple data bits from low to high bits according to the preset bit width to obtain multiple sub-scalars with smaller bit widths.
[0077] For example, in the embodiment of the present application, it is assumed that the scalars involved in the calculation include , , ,in,
[0078] =11001110;
[0079] =110001;
[0080] =1001.
[0081] Assume that the preset bit width c=2, then It can be divided into sub-scalars as shown in Table 1 below:
[0082] Table 1
[0083]
[0084] Referring to Table 1, when performing segmentation, the scalar is segmented from low to high bits according to the preset bit width, that is, segmented from the right side of the scalar to the left side. j represents the segmentation stage. For example, the column with j=0 is the sub-scalar obtained by segmenting each scalar at segmentation stage 0, and the column with j=1 is the sub-scalar obtained by segmenting each scalar at segmentation stage 1. Other segmentation stages can be deduced in the same way, and will not be described in detail here.
[0085] use express The sub-scalar obtained in the j-division stage, then for the scalar In terms of , it is divided into four sub-scalars in the four segmentation stages and can be expressed as =10, 11, =00, .
[0086] Similarly, scalar is split into four sub-scalars and can be expressed as =10, 00, =10, ;
[0087] Scalar is split into two sub-scalars and can be expressed as =10, 10. Due to The bit width is small, so there is no subscalar in the split stage 2 and 3, or it can also be written as =00, 00.
[0088] In addition, in the embodiment of the present application, each scalar and each sub-scalar obtained by segmentation also have the following relationship:
[0089]
[0090] in, is a scalar, For The sub-scalar obtained by the j-th segmentation (segmentation stage j), c represents the preset bit width, , m is the value of the highest bit width in each scalar. For example, The maximum bit width is 8, so m is 8, r=4, j=0, 1, 2, 3.
[0091] In addition, in the embodiment of the present application, the sub-scalars in the same segmentation stage have the same position order in the original scalar. , , The number of digits in each original scalar is the first and second digits from the lowest to the highest, then , , is in the same segmentation stage. Similarly, , , For the same segmentation stage, that is, each sub-scalar Sub-scalars with the same j value are in the same segmentation stage.
[0092] like Figure 2 As shown, in the embodiment of the present application, the operation module 120 includes a first state machine 121 and a first multiplier 122;
[0093] The scalar divider 110 is used to send each sub-scalar and the corresponding point vector obtained in each division stage to the first state machine 121;
[0094] The first state machine 121 is used to receive the sub-scalars and corresponding point vectors obtained in each segmentation stage, and in each segmentation stage, send the point vectors corresponding to the sub-scalars of each data size to the first multiplier 122 respectively;
[0095] The first multiplier 122 is used to receive point vectors corresponding to sub-scalars of each data size in each segmentation stage, and to accumulate all point vectors corresponding to sub-scalars of the same data size to obtain accumulation results of each point in each segmentation stage.
[0096] For example, in the embodiment of the present application, the bit width of the sub-scalar is c, then the minimum sub-scalar is 0 and the maximum is -1, so after receiving each sub-scalar and the corresponding point vector in any segmentation stage, the first state machine 121 can =0 starts to traverse to -1, each point vector corresponding to the sub-scalar value of 0 is sent to the first multiplier 122, and each point vector corresponding to the sub-scalar value of 1 is sent to the first multiplier 122, until the sub-scalar value is -1 are sent to the first multiplier 122. After the first multiplier 122 obtains the point vectors corresponding to the sub-scalars of different data sizes, it accumulates the point vectors corresponding to the sub-scalars of the same data size to obtain the point accumulation results corresponding to the sub-scalars of each data size, that is, the number of point accumulation results in any segmentation stage is the same as the number of sub-scalars of different data sizes in the segmentation stage.
[0097] In the embodiment of the present application, the point accumulation result can also be calculated by referring to the following formula (3):
[0098]
[0099] Among them, k∈[0, ], is a scalar, is a scalar The subscalar in segmentation stage j, is the point vector to be processed, i∈[1,n], is in the segmentation stage j, The accumulated result of the points at that time.
[0100] In an embodiment of the present application, the first multiplier 122 is also used to write the accumulated results of each point in each segmentation stage into the corresponding bucket of the bucket memory 130; wherein the accumulated results of each point in the same segmentation stage are located in the same bucket, and the depth address of the storage space of each point accumulated result in the bucket is the data size of the sub-scalar corresponding to the accumulated result of each point.
[0101] The bucket storage 130 is pre-set with multiple buckets, each bucket corresponding to a segmentation stage. As shown in the above formula (3), according to Calculate the cumulative results of each point, then k can be used as the depth address of the cumulative results of each point in its corresponding bucket, k∈[0, -1], so the depth of the bucket used to store each segmentation stage is , including 0 -1Total For example, when k=0, Stored in the storage space with a depth address of 0 in the corresponding bucket, when k=1, Stored in the storage space with a depth address of 1 in the corresponding bucket, and so on. -1 hour, The depth address stored in the corresponding bucket is -1 storage space.
[0102] like Figure 2 As shown, in the embodiment of the present application, the operation module 120 further includes a second state machine 123 and a second multiplier 124;
[0103] The second state machine 123 is used to: read the accumulated results of each point in each segmentation stage and the corresponding depth address from the bucket corresponding to each segmentation stage in the order of the segmentation stages from high to low, and send the accumulated results of each point in each segmentation stage and the corresponding depth address to the second multiplier 124;
[0104] The second multiplier 124 is used to: receive the accumulated results of each point in each segmentation stage and the corresponding depth address in the order of the segmentation stages from high to low; and use the depth address corresponding to the accumulated results of each point in each segmentation stage as the weight of the accumulated results of each point, multiply the accumulated results of each point with their corresponding weights, and accumulate them to obtain the sub-accumulated results of each segmentation stage.
[0105] After calculating the point accumulation result in each segmentation stage, in the embodiment of the application, the sub-accumulation result of each segmentation stage can also be determined based on the second state machine 123 and the second multiplier 124 according to the segmentation stage order from high to low.
[0106] in, , m is the highest bit width in each scalar, c is the bit width of the preset sub-scalar, then r represents the maximum number of times that can be divided in each scalar, such as =11001110, is a scalar with the maximum bit width, so m=8, c=2, then r=4, that is, each scalar in the data to be processed can be divided at most four times, j starts from 0, so j∈[0, r-1] is j∈[0, 3].
[0107] In the same segmentation stage, after the accumulation results of each point are obtained, the sub-scalar data size corresponding to the accumulation results of each point is used as the depth address, and the accumulation results of each point are stored in the storage space of the corresponding depth address of the corresponding bucket. Therefore, in the embodiment of the present application, when calculating the sub-accumulation results of each segmentation stage, the accumulation results of each point in the current segmentation stage and the depth address stored in the bucket can be read based on the second state machine 123, and sent to the second multiplier 124. The second multiplier 124 uses the depth address of each point accumulation result as the respective weight, multiplies the accumulation results of each point with the respective corresponding weights, and then accumulates them to obtain the sub-accumulation results of the current segmentation stage.
[0108] In the embodiment of the present application, the sub-accumulation results of each segmentation stage can also be calculated by the following formulas (4) and (5):
[0109] (4)
[0110]
[0111] in, is the accumulated result of the points with depth address k in segmentation stage j, At segmentation stage j, the product of the accumulated result of the points in depth address k and the weight k, is the sub-accumulated result of segmentation stage j.
[0112] Based on the second state machine 123 and the second multiplier 124, the sub-accumulation result of each segmentation stage can be calculated by the above formulas (4) and (5). At this time, the calculation result of the multi-scalar multiplication of the data to be processed can be calculated based on the following formula (6):
[0113]
[0114] in, Represents the calculation result of multi-scalar multiplication of n point vectors and scalars, j represents the segmentation stage number, j∈[0,r-1], represents the sub-accumulated result of segmentation stage j, Indicates that the sub-accumulated result of segmentation stage j needs to be flipped times.
[0115] In the prior art, it is necessary to accumulate the sub-summation results of each segmentation stage After obtaining the sub-accumulated results of all the segmentation stages, store the sub-accumulated results of each segmentation stage. In this process, firstly, j memories are needed to store the sub-accumulated results of j segmentation stages; in addition, in order to facilitate calculation, the depth address of the sub-accumulated results of each segmentation stage is often determined by the multiples required to multiply them. For example, It needs to be multiplied by cj, so the depth address when it is stored is , so that the depth address can be Read from storage space After that, directly multiply the depth address as its corresponding weight, that is, Since j memories are needed to store the sub-accumulated results of j segmentation stages, As When storing a deep address, the required memory depth is at least , which will result in a larger memory depth and waste of storage resources.
[0116] In the embodiment of the present application, based on the above formula (6), it can be known that j∈[0, r-1], there are a total of r segmentation stages, and each segmentation stage has a sub-accumulation result, as shown below:
[0117] When j=0, the sub-accumulation result is , need to be multiplied by ;
[0118] When j=1, the sub-accumulation result is , need to be multiplied by ;
[0119] When j=2, the sub-accumulation result is , need to be multiplied by ;
[0120] And so on.
[0121] When j=r-3, the sub-accumulation result is , need to be multiplied by ;
[0122] When j=r-2, the sub-accumulation result is , need to be multiplied by ;
[0123] When j=r-1, the sub-accumulation result is , need to be multiplied by ;
[0124] Then, the above formula (6) can be transformed into:
[0125] = * + * + * +……+ * + *
[0126] =( * + )* + * +……+ * + *
[0127] =(( * + )* + )* +……+ * + *
[0128] =(((( * + )* + )* +……+ ) + )*
[0129] =((( * + )* + )* +……+ ) +
[0130] Based on the above derivation, we can know that:
[0131] is the sub-accumulated result at the split stage r-1, multiplied by Then add the sub-accumulation result of the segmentation stage r-2 have the same weight, that is, have a weight of (r-2)c, as shown in the above formula ( * + )* ;
[0132] Will * + Multiply After that, the sub-accumulation result of the segmentation stage r-3 is added have the same weight, that is, have a weight of (r-3)c, as shown in the above formula (( * + )* + )* .
[0133] Therefore, in an embodiment of the present application, iteration can be performed based on the sub-accumulated results of each segmentation stage. For example, in an embodiment of the present application, the sub-accumulated results of each segmentation stage are calculated in the order of the segmentation stages from high to low, that is, in the order of j=r-1, r-2, r-3, ..., j=2, j=1, j=0.
[0134] When j=r-1, the sub-accumulation result is , the initial segmentation stage, which is taken as the iterative result of the initial segmentation stage;
[0135] When j=r-2, the sub-accumulation result is , multiply the iterative result of the segmentation stage j=r-1 by After that, it is accumulated with the sub-accumulation result of this segmentation stage as the iterative result of this segmentation stage, that is, * + ;
[0136] When j=r-3, the sub-accumulation result is , multiply the iterative result of the segmentation stage j=r-2 by After that, it is accumulated with the sub-accumulation result of this segmentation stage as the iterative result of this segmentation stage, that is, ( * + )* + ;
[0137] By analogy, when j=0, the sub-accumulation result is , multiply the iterative result of the segmentation stage j=1 by After that, it is accumulated with the sub-accumulation result of this segmentation stage as the iterative result of this segmentation stage, that is, ((( * + )* + )* +……+ ) + , combined with the conversion process of the above formula (6), it can be seen that the iterative result of the segmentation stage j = 0 is the final calculation result, that is, the iterative result of the segmentation stage j = 0 is .
[0138] Therefore, in the embodiment of the present application, the second state machine 123 and the second multiplier 124 determine the point accumulation result corresponding to each segmentation stage according to the segmentation stage sequence from high to low.
[0139] In addition, in the segmentation stage, after the accumulation results of each point are obtained, the sub-scalar data size corresponding to the accumulation results of each point is used as the depth address, and the accumulation results of each point are stored in the corresponding bucket. , k is the depth address where the accumulated results of each point are stored in the bucket, k∈[0, -1], so the depth of the bucket used to store the accumulated results of each segmentation stage is , the maximum depth address is -1. When calculating the sub-accumulated results, it is necessary to read each point accumulated result from the corresponding bucket. Therefore, when the sub-accumulated result of each segmentation stage is obtained, the corresponding bucket for storing the point accumulated result of the segmentation stage is in an empty state. Therefore, in the embodiment of the present application, the second multiplier 124 is also used to store the sub-accumulated result of each segmentation stage in the bucket corresponding to each segmentation stage, thereby saving storage space.
[0140] like Figure 2As shown, in the embodiment of the present application, the operation module 120 further includes a third state machine 125 and a third multiplier 126;
[0141] The third state machine 125 is used to: read the iteration result and the corresponding depth address of the previous segmentation stage, and the sub-accumulation result of the current segmentation stage from the bucket corresponding to each segmentation stage in the order of the segmentation stages from high to low, and send the iteration result and the corresponding depth address of the previous segmentation stage, and the sub-accumulation result of the current segmentation stage to the third multiplier 126;
[0142] The third multiplier-adder 126 is used to use the address depth of the iteration result of the previous segmentation stage as a weight, multiply it with the iteration result of the previous segmentation stage, and then accumulate it with the sub-accumulation result of the current segmentation stage to obtain the iteration result of the current segmentation stage;
[0143] The iteration result of the current segmentation stage is stored in the storage space of at least two different depth addresses of the bucket corresponding to each segmentation stage, and the sum of the at least two different depth addresses is , where c is the bit width of the sub-scalar, and the corresponding iteration result in the segmentation stage of the highest bit is its corresponding sub-accumulation result.
[0144] Based on the conversion process of the above formula (6), it can be known that the sub-accumulated results of each segmentation stage can be calculated respectively according to the segmentation stage order from high to low, and the iterative result can be calculated. Among them, the iterative result in the highest-order segmentation stage (j=r-1) is the sub-accumulated result of the segmentation stage. Then, in the highest-order segmentation stage, the third state machine 125 only reads the sub-accumulated result in the bucket corresponding to the highest-order segmentation stage, and sends it to the third multiplier 126. The third multiplier 126 uses the sub-accumulated result of the highest-order segmentation stage as the iterative result of the highest-order segmentation stage and stores it.
[0145] In addition, based on the conversion process of the above formula (6), it can be seen that each segmentation stage needs to read the iterative result of the previous segmentation stage and multiply it by , in order to facilitate calculation, the third multiplier 126 can store the iteration result of each segmentation stage in the corresponding bucket with a depth address of Therefore, when the iteration result is read from the depth address, the address depth of the storage space can be directly used as the weight and directly multiplied by the read iteration result.
[0146] In addition, considering that the maximum depth address used to store the point accumulation results of each segmentation stage is -1, less than Therefore, in the embodiment of the present application, the third multiplier 126 can also store the iterative result of each segmentation stage in the storage space of at least two different depth addresses of the corresponding bucket, and the sum of the at least two different depth addresses is .
[0147] For example, for each point in the highest bit segmentation stage r-1, the cumulative result is The first multiplier 122 stores the data at depth addresses 0 to 1 in the corresponding buckets. -1 storage space, the sub-accumulation result of the segmentation stage r-1 is calculated When the second state machine 123 has accumulated the results of each point From the corresponding bucket depth address 0 to -1 is taken out, then the iterative result is obtained in the third multiplier 126 After that, it can be stored in the bucket corresponding to the segmentation stage with depth addresses 1 and -1, in the segmentation stage r-2, when the third state machine 125 reads the iteration result of the segmentation stage r-1, it reads from the storage address 1 and -1 for reading, including and the corresponding depth address 1, and And the corresponding depth address -1, and sends it to the third multiplier 126, which uses the depth address as a weight to multiply the read iteration result and then accumulates it.
[0148] Specifically, the multiplier-accumulator receives a depth address of 1 and , multiplied to get 1* ;
[0149] The depth address received by the multiplier-accumulator -1 and , multiply and we get ;
[0150] After accumulation, + , that is, it is still Thus, not only can the iterative result of each segmentation stage be stored in the bucket used to store the point accumulation result of the segmentation stage without the need for additional memory, but also, when the iterative result of the previous stage is read in the next segmentation stage, the depth address corresponding to the iterative result in the previous segmentation stage can be directly multiplied as its weight, which also facilitates calculation.
[0151] In addition, it should be noted that in addition to storing the iteration results in the storage space of the two depth addresses of the corresponding bucket, the iteration results can also be stored in the storage space of more depth addresses of the corresponding bucket, for example, the storage space of the corresponding bucket at depth addresses 1, 2, 3, The corresponding storage space is stored, as long as the sum of the depth addresses of the storage space used to store the iteration results is That's it.
[0152] In addition, the third state is also used to obtain the sub-accumulation result of the current segmentation stage and send it to the third multiplier 126. The third multiplier 126 multiplies the iterative result of the previous stage and the corresponding depth address, and adds it with the point accumulation result of the current segmentation stage to obtain the iterative result of the current segmentation stage, and stores it in the corresponding bucket of the current segmentation stage. The above-mentioned method of storing in the storage space of at least two depth addresses is still used during storage, which will not be elaborated here.
[0153] After multiple iterations, the iterative results of each segmentation stage can be gradually obtained until the iterative result of the last segmentation stage (lowest bit) is obtained. At this time, the iterative result of the last segmentation stage is the calculation result of the multi-scalar multiplication of the data to be processed.
[0154] In the embodiment of the present application, by setting the operation module 120 and the bucket memory 130, after calculating the sub-accumulated result of each segmentation stage, the iterative result of each segmentation stage is obtained one by one from high to low, and the iterative result of each segmentation stage is stored in the storage space of at least two different depth addresses of the corresponding bucket, and the sum of the at least two depth addresses is , then the depth of the bucket is no greater than -1, therefore, the iterative result of each segmentation stage can be stored in the bucket used to store the accumulated result of each segmentation stage point without referencing additional memory, which can avoid the waste of storage space and facilitate the reading and calculation of the next segmentation stage.
[0155] Exemplary Methods
[0156] A data processing device in an embodiment of the present application is described above. Next, a data processing method in an exemplary embodiment of the present application is introduced.
[0157] Reference Figure 3 In an embodiment of the present application, the data processing method is applied to the data processing device described in any of the above embodiments, and the data processing method includes the following steps S100-S400:
[0158] Step S100: receiving data to be processed, and dividing the scalar of each point vector in the data to be processed into a plurality of sub-scalars from low to high bits according to a preset bit width;
[0159] Step S200: In each segmentation stage, based on the point vectors corresponding to the sub-scalars of each data size, the accumulation results of each point in each segmentation stage are accumulated respectively; and the accumulation results of each point in each segmentation stage are written into the corresponding bucket of the bucket memory; wherein, the accumulation results of each point in the same segmentation stage are located in the same bucket, and the depth address of each point accumulation result in the bucket is the data size of the sub-scalar corresponding to each point accumulation result;
[0160] Step S300: According to the order of the segmentation stages from high to low, the depth address corresponding to the accumulated result of each point in each segmentation stage is used as the weight of the accumulated result of each point, and the accumulated result of each point is multiplied by the corresponding weight to obtain the sub-accumulated result of each segmentation stage;
[0161] Step S400: Determine the iterative result of each segmentation stage according to the segmentation stage sequence from high to low, and store the iterative result of each segmentation stage in the bucket corresponding to the bucket memory, and the iterative result of each segmentation stage is stored in the storage space of at least two different depth addresses in the corresponding bucket, and the sum of the at least two different depth addresses is , until the iterative result of the last segmentation stage is obtained, and the iterative result of the last segmentation stage is used as the calculation result of the multi-scalar multiplication of the data to be processed; c is the bit width of the sub-scalar;
[0162] Among them, the iterative result of each segmentation stage is obtained by multiplying the iterative result of the previous segmentation stage by its corresponding weight, and then accumulating it with the sub-accumulated result of the current segmentation stage; the weight of the iterative result of each segmentation stage is the depth address stored in the corresponding bucket, and the iterative result of the highest-order segmentation stage is its corresponding sub-accumulated result.
[0163] The specific implementation method of each step refers to the various embodiments of the above exemplary device, which will not be described in detail here.
[0164] In the embodiment of the present application, after obtaining the sub-scalars in each segmentation stage, the data processing method further includes:
[0165] Sending each sub-scalar and the corresponding point vector obtained in each segmentation stage to the first state machine;
[0166] In each segmentation stage, point vectors corresponding to sub-scalars of each data size are sent to the first multiplier-adder based on the first state machine;
[0167] In each segmentation stage, all point vectors corresponding to sub-scalars of the same data size are accumulated based on the first multiplier to obtain accumulation results of each point in each segmentation stage.
[0168] In the embodiment of the present application, after obtaining the accumulation results of each point in each segmentation stage, the data processing method further includes:
[0169] The accumulated results of each point in each segmentation stage are written into the corresponding bucket of the bucket memory; wherein, the accumulated results of each point in the same segmentation stage are located in the same bucket, and the depth address of each accumulated result of each point in the bucket is the data size of the sub-scalar corresponding to the accumulated result of each point.
[0170] In the embodiment of the present application, after the point accumulation results of each segmentation stage are written into the corresponding bucket, the data processing method further includes:
[0171] Read the accumulated results of each point in each segmentation stage and the corresponding depth address from the bucket corresponding to each segmentation stage in the order of the segmentation stages from high to low, and send the accumulated results of each point in each segmentation stage and the corresponding depth address to the second multiplier-adder;
[0172] According to the order of the segmentation stages from high to low, based on the second multiplier, the depth address corresponding to the accumulation result of each point in each segmentation stage is used as the weight of the accumulation result of each point, and the accumulation result of each point is multiplied by the corresponding weight to obtain the sub-accumulation result of each segmentation stage.
[0173] In the embodiment of the present application, after obtaining the sub-accumulation result of each segmentation stage, the data processing method further includes:
[0174] Based on the second multiplier-adder, the sub-accumulation result of each segmentation stage is stored in the bucket corresponding to each segmentation stage.
[0175] In the embodiment of the present application, after storing the sub-accumulated results of each segmentation stage in the bucket corresponding to each segmentation stage, the data processing method further includes:
[0176] According to the order of the segmentation stages from high to low, read the iteration result and the corresponding depth address of the previous segmentation stage, and the sub-accumulation result of the current segmentation stage from the bucket corresponding to each segmentation stage, and send the iteration result and the corresponding depth address of the previous segmentation stage, and the sub-accumulation result of the current segmentation stage to the third multiplier-adder;
[0177] According to the order of the segmentation stages from high to low, the third multiplier-adder uses the address depth of the iteration result of the previous segmentation stage as a weight, multiplies it with the iteration result of the previous segmentation stage, and then accumulates it with the sub-accumulation result of the current segmentation stage to obtain the iteration result of the current segmentation stage;
[0178] The iteration result of the current segmentation stage is stored in the storage space of at least two different depth addresses of the bucket corresponding to each segmentation stage, and the sum of the at least two different depth addresses is , where c is the bit width of the sub-scalar, and the corresponding iteration result in the segmentation stage of the highest bit is its corresponding sub-accumulation result.
[0179] The data processing method in the embodiment of the present application, after calculating the sub-accumulated result of each segmentation stage, uses the method of obtaining the iterative result of each segmentation stage one by one from high to low, and stores the iterative result of each segmentation stage in the storage space of at least two different depth addresses of the corresponding bucket, and the sum of the at least two depth addresses is , then the depth of the bucket is no greater than -1, therefore, the iterative result of each segmentation stage can be stored in the bucket used to store the accumulated result of each segmentation stage point without referencing additional memory, which can avoid the waste of storage space and facilitate the reading and calculation of the next segmentation stage.
[0180] Exemplary Media
[0181] After introducing the method, medium and system of the exemplary embodiments of the present application, next, reference is made to Figure 4 For a description of the computer-readable storage medium of the exemplary embodiment of the present application, please refer to Figure 4 , the computer-readable storage medium shown is a CD 70, on which a computer program (i.e., a program product) is stored. When the computer program is executed by the processor, the steps described in the above method implementation are implemented, for example, receiving data to be processed, dividing the scalar of each point vector in the data to be processed from low to high bits into a plurality of sub-scalars according to a preset bit width; in each segmentation stage, the point vectors corresponding to the sub-scalars based on the size of each data are respectively accumulated to obtain the accumulation results of each point in each segmentation stage; and writing the accumulation results of each point in each segmentation stage into the corresponding bucket of the bucket memory. In the segmentation stage sequence from high to low, the depth address corresponding to the accumulated result of each point in each segmentation stage is used as the weight of the accumulated result of each point, and the accumulated result of each point is multiplied by the corresponding weight, and the sub-accumulated result of each segmentation stage is accumulated; according to the segmentation stage sequence from high to low, the iterative result of each segmentation stage is determined, and the iterative result of each segmentation stage is stored in the bucket corresponding to the bucket memory, and the iterative result of each segmentation stage is stored in the storage space of at least two different depth addresses of the corresponding bucket, and the sum of the at least two different depth addresses is , until the iterative result of the last segmentation stage is obtained, and the iterative result of the last segmentation stage is used as the calculation result of the multi-scalar multiplication of the data to be processed; c is the bit width of the sub-scalar. The specific implementation method of each step is not repeated here. It should be noted that examples of the computer-readable storage medium may also include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other optical or magnetic storage media, which are not repeated here one by one.
[0182] Exemplary Chips
[0183] After introducing the method, apparatus, and medium of the exemplary embodiment of the present application, next, the chip of the exemplary embodiment of the present application is introduced.
[0184] The present application proposes a chip, including the data processing device described in any of the above embodiments, and thus has all the beneficial effects of the data processing device described in the above embodiments, which will not be described one by one here.
[0185] Finally, it should be noted that the above-described embodiments are only specific implementation methods of the present application, which are used to illustrate the technical solutions of the present application, rather than to limit them. The protection scope of the present application is not limited thereto. Although the present application is described in detail with reference to the above-described embodiments, ordinary technicians in the field should understand that any technician familiar with the technical field can still modify the technical solutions recorded in the above-described embodiments within the technical scope disclosed in the present application, or can easily think of changes, or make equivalent replacements for some of the technical features therein; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be included in the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.
[0186] In addition, although the operations of the method of the present application are described in a specific order in the drawings, this does not require or imply that the operations must be performed in this specific order, or that all the operations shown must be performed to achieve the desired results. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step, and / or one step may be decomposed into multiple steps.
Claims
1. A data processing device, characterized in that: The device comprises: a scalar divider, an operation module and a bucket memory; The scalar divider is used to receive the data to be processed, and divide the scalar of each point vector in the data to be processed from low to high bits into multiple sub-scalars according to a preset bit width; The computing module is used for: In each segmentation stage, based on the point vectors corresponding to the sub-scalars of each data size, the accumulation results of each point in each segmentation stage are accumulated respectively; the accumulation results of each point in each segmentation stage are written into the corresponding bucket of the bucket memory; wherein, the accumulation results of each point in the same segmentation stage are located in the same bucket, and the depth address of each point accumulation result in the bucket is the data size of the sub-scalar corresponding to each point accumulation result; According to the order of the segmentation stages from high to low, the depth address corresponding to the accumulated result of each point in each segmentation stage is used as the weight of the accumulated result of each point, and the accumulated result of each point is multiplied by the corresponding weight, and the sub-accumulated result of each segmentation stage is accumulated; According to the order of the segmentation stages from high to low, the iterative result of each segmentation stage is determined, and the iterative result of each segmentation stage is stored in the bucket corresponding to the bucket memory, and the iterative result of each segmentation stage is stored in the storage space of at least two different depth addresses in the corresponding bucket, and the sum of the at least two different depth addresses is , until the iterative result of the last segmentation stage is obtained, and the iterative result of the last segmentation stage is used as the calculation result of the multi-scalar multiplication of the data to be processed; c is the bit width of the sub-scalar; Among them, the iterative result of each segmentation stage is obtained by multiplying the iterative result of the previous segmentation stage by its corresponding weight, and then accumulating it with the sub-accumulated result of the current segmentation stage; the weight of the iterative result of each segmentation stage is the depth address stored in the corresponding bucket, and the iterative result of the highest-order segmentation stage is its corresponding sub-accumulated result.
2. The data processing device according to claim 1, characterized in that The operation module includes a first state machine and a first multiplier-adder; The scalar divider is used to send each sub-scalar and the corresponding point vector obtained in each division stage to the first state machine; The first state machine is used to receive the sub-scalars and corresponding point vectors obtained in each segmentation stage, and in each segmentation stage, send the point vectors corresponding to the sub-scalars of each data size to the first multiplier-adder respectively; The first multiplier-adder is used to receive point vectors corresponding to sub-scalars of each data size in each segmentation stage, and to accumulate all point vectors corresponding to sub-scalars of the same data size to obtain accumulation results of each point in each segmentation stage.
3. The data processing device according to claim 2, characterized in that The first multiplier is also used to: write the accumulated results of each point in each segmentation stage into the corresponding bucket of the bucket memory; wherein the accumulated results of each point in the same segmentation stage are located in the same bucket, and the depth address of each point accumulated result in the bucket is the data size of the sub-scalar corresponding to the accumulated result of each point.
4. The data processing device according to claim 3, characterized in that The operation module also includes a second state machine and a second multiplier-adder; The second state machine is used to: read the accumulated results of each point in each segmentation stage and the corresponding depth address from the bucket corresponding to each segmentation stage in the order of the segmentation stages from high to low, and send the accumulated results of each point in each segmentation stage and the corresponding depth address to the second multiplier; The second multiplier is used to: receive the accumulated results of each point in each segmentation stage and the corresponding depth address in the order of the segmentation stages from high to low; and use the depth address corresponding to the accumulated results of each point in each segmentation stage as the weight of the accumulated results of each point, multiply the accumulated results of each point with their corresponding weights, and accumulate to obtain the sub-accumulated results of each segmentation stage.
5. The data processing device according to claim 4, characterized in that The second multiplier-adder is further used to store the sub-accumulation result of each segmentation stage in the bucket corresponding to each segmentation stage.
6. The data processing device according to claim 5, characterized in that The operation module also includes a third state machine and a third multiplier-adder; The third state machine is used to: read the iteration result and the corresponding depth address of the previous segmentation stage, and the sub-accumulation result of the current segmentation stage from the bucket corresponding to each segmentation stage in the order of the segmentation stages from high to low, and send the iteration result and the corresponding depth address of the previous segmentation stage, and the sub-accumulation result of the current segmentation stage to the third multiplier; The third multiplier-adder is used to use the address depth of the iteration result of the previous segmentation stage as a weight, multiply it with the iteration result of the previous segmentation stage, and then accumulate it with the sub-accumulation result of the current segmentation stage to obtain the iteration result of the current segmentation stage according to the segmentation stage order from high to low; The iteration result of the current segmentation stage is stored in the storage space of at least two different depth addresses of the bucket corresponding to each segmentation stage, and the sum of the at least two different depth addresses is , where c is the bit width of the sub-scalar, and the corresponding iteration result in the segmentation stage of the highest bit is its corresponding sub-accumulation result.
7. A data processing method, applied to the data processing device according to any one of claims 1 to 6, the data processing method comprising: Receive data to be processed, and divide the scalar of each point vector in the data to be processed into multiple sub-scalars from low to high bits according to a preset bit width; In each segmentation stage, the point vectors corresponding to the sub-scalars of each data size are accumulated to obtain the accumulation results of each point in each segmentation stage; the accumulation results of each point in each segmentation stage are written into the corresponding bucket of the bucket memory; wherein, the accumulation results of each point in the same segmentation stage are located in the same bucket, and the depth address of each point accumulation result in the bucket is the data size of the sub-scalar corresponding to each point accumulation result; According to the order of the segmentation stages from high to low, the depth address corresponding to the accumulated result of each point in each segmentation stage is used as the weight of the accumulated result of each point, and the accumulated result of each point is multiplied by the corresponding weight, and the sub-accumulated result of each segmentation stage is accumulated; According to the order of the segmentation stages from high to low, the iterative result of each segmentation stage is determined, and the iterative result of each segmentation stage is stored in the bucket corresponding to the bucket memory, and the iterative result of each segmentation stage is stored in the storage space of at least two different depth addresses of the corresponding bucket, and the sum of the at least two different depth addresses is , until the iterative result of the last segmentation stage is obtained, and the iterative result of the last segmentation stage is used as the calculation result of the multi-scalar multiplication of the data to be processed; c is the bit width of the sub-scalar; Among them, the iterative result of each segmentation stage is obtained by multiplying the iterative result of the previous segmentation stage by its corresponding weight, and then accumulating it with the sub-accumulated result of the current segmentation stage; the weight of the iterative result of each segmentation stage is the depth address stored in the corresponding bucket, and the iterative result of the highest-order segmentation stage is its corresponding sub-accumulated result.
8. The data processing method according to claim 7, after obtaining the sub-scalars in each segmentation stage, the data processing method further comprises: Send each sub-scalar and the corresponding point vector obtained in each segmentation stage to the first state machine; In each segmentation stage, point vectors corresponding to sub-scalars of each data size are respectively sent to the first multiplier-adder based on the first state machine; In each segmentation stage, all point vectors corresponding to sub-scalars of the same data size are accumulated based on the first multiplier to obtain accumulation results of each point in each segmentation stage.
9. The data processing method according to claim 8, characterized in that: After obtaining the accumulated results of each point in each segmentation stage, the data processing method further includes: The accumulated results of each point in each segmentation stage are written into the corresponding bucket of the bucket memory; wherein, the accumulated results of each point in the same segmentation stage are located in the same bucket, and the depth address of each accumulated result of each point in the bucket is the data size of the sub-scalar corresponding to the accumulated result of each point.
10. The data processing method according to claim 9, after writing the point accumulation results of each segmentation stage into the corresponding bucket, the data processing method further comprises: According to the order of the segmentation stages from high to low, the accumulated results of each point in each segmentation stage and the corresponding depth address are read from the bucket corresponding to each segmentation stage, and the accumulated results of each point in each segmentation stage and the corresponding depth address are sent to the second multiplier-adder; According to the order of the segmentation stages from high to low, based on the second multiplier, the depth address corresponding to the accumulation result of each point in each segmentation stage is used as the weight of the accumulation result of each point, and the accumulation result of each point is multiplied by the corresponding weight to obtain the sub-accumulation result of each segmentation stage.
11. The data processing method according to claim 10, after obtaining the sub-accumulation result of each segmentation stage, the data processing method further comprises: Based on the second multiplier-adder, the sub-accumulation result of each segmentation stage is stored in the bucket corresponding to each segmentation stage.
12. The data processing method according to claim 11, after storing the sub-accumulation results of each segmentation stage in the bucket corresponding to each segmentation stage, the data processing method further comprises: According to the order of the segmentation stages from high to low, the iteration result and the corresponding depth address of the previous segmentation stage, as well as the sub-accumulation result of the current segmentation stage are read from the bucket corresponding to each segmentation stage, and the iteration result and the corresponding depth address of the previous segmentation stage, as well as the sub-accumulation result of the current segmentation stage are sent to the third multiplier-adder; According to the order of the segmentation stages from high to low, the third multiplier-adder uses the address depth of the iteration result of the previous segmentation stage as a weight, multiplies it with the iteration result of the previous segmentation stage, and then accumulates it with the sub-accumulation result of the current segmentation stage to obtain the iteration result of the current segmentation stage; The iteration result of the current segmentation stage is stored in the storage space of at least two different depth addresses of the bucket corresponding to each segmentation stage, and the sum of the at least two different depth addresses is , where c is the bit width of the sub-scalar, and the corresponding iteration result in the segmentation stage of the highest bit is its corresponding sub-accumulation result.
13. A computer-readable storage medium, characterized in that: It comprises instructions which, when executed on a computer, cause the computer to perform the method as claimed in any one of claims 7 to 12.
14. A chip, characterized in that: The chip includes the data processing device as described in any one of claims 1-6.
Citation Information
Patent Citations
Multi-scalar multiplier and acceleration method
CN116954559A
Data processing device and method, medium and computing equipment
CN118349213A