Method and device for multiple computing nodes to collaboratively perform exponential normalization computing
By synchronizing the local maximum value and local index summing results among multiple computing nodes, the problem of frequent data synchronization and reading in parallel exponential normalization calculations of multiple computing nodes is solved, which improves computing efficiency and reduces latency and power consumption.
Patent Information
- Application Number
- CN202510058604.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-15
- Publication Date
- 2025-05-13
AI Technical Summary
When multiple computing nodes perform exponential normalization calculations in parallel, data synchronization and reading are frequent, resulting in high delays, large power consumption, and possible data transmission conflicts, reducing computing efficiency.
Each computing node calculates the local maximum value and local index summation results, and synchronizes these results among multiple computing nodes through a dedicated parameter sharing bus to avoid global data synchronization and multiple reads.
The calculation efficiency of multiple computing nodes in collaborative exponential normalization calculations is improved, data delay and data transmission conflicts are reduced, and power consumption is saved.
Smart Images

Figure CN119988009A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computing technology, and in particular to a method and device for multiple computing nodes to collaboratively perform exponential normalization computing. Background Art
[0002] In recent years, various large models based on neural networks have shown strong performance in various applications. In neural networks, the exponential normalization operator (softmax) is a very important and common nonlinear operator. When training large models, exponential normalization calculation is one of the key factors affecting training efficiency. Summary of the invention
[0003] In large models based on neural networks, the input data is large in scale, and it is usually impossible to complete all calculations for all input data in the same computing node, but parallel calculations need to be performed in different computing nodes. In this case, data synchronization between computing nodes will take up a lot of computing resources. In addition, data synchronization between multiple computing nodes often uses existing data paths, which may cause data transmission conflicts and reduce transmission efficiency.
[0004] When performing exponential normalization (softmax) calculations in large models, exponential normalization calculations are usually performed for different parts of the input data in multiple computing nodes. However, exponential normalization calculations do not only involve the respective input data (local input data) of each computing node, but require the global maximum and global exponential summation results of all input data (global input data) of all computing nodes. In the related art, when multiple computing nodes perform exponential normalization calculations in parallel, not only is it necessary to synchronize the input data of each computing node to other nodes so that each node obtains the global input data, but also it is necessary to read the global input data multiple times during the calculation process. Such large-scale data synchronization and data reading will cause higher delays and larger power consumption overhead. In addition, data synchronization between multiple computing nodes often uses existing data paths, which may cause data transmission conflicts and reduce computing efficiency.
[0005] In view of this, the present application proposes a method, apparatus, system, equipment, storage medium and program product for multiple computing nodes to collaboratively perform exponential normalization calculations, so as to improve the computing efficiency of multiple computing nodes collaboratively performing exponential normalization calculations, reduce data delays and data transmission conflicts, and save power consumption.
[0006] Specifically, the present application is implemented through the following technical solutions:
[0007] In a first aspect, the present application provides a method for multiple computing nodes to collaboratively perform exponential normalization calculations, the method comprising: at each computing node among the multiple computing nodes: calculating the local maximum and local exponential summation result of the input data of the computing node; sending the local maximum and local exponential summation result of the computing node to other computing nodes among the multiple computing nodes, and receiving the local maximum and local exponential summation results of the other computing nodes; calculating the global maximum and global exponential summation result based on the local maximum and local exponential summation result of the computing node and the local maximum and local exponential summation results of the other computing nodes; and performing exponential normalization on the input data of the computing node based on the global maximum and global exponential summation result.
[0008] In a second aspect, the present application provides an apparatus for multiple computing nodes to collaboratively perform exponential normalization calculations, each of the multiple computing nodes comprising the apparatus, and the apparatus comprising: a computing control module, configured to parse instructions sent by a controller included in the apparatus, and control the operation of each module in the apparatus according to the instructions; an exponential calculation module, configured to perform exponential calculation on the input data of the computing node, and to perform exponential normalization on the input data of the computing node based on the global maximum value and the global exponential summation result; a preprocessing module, configured to calculate the local maximum value of the input data of the computing node, and to sum the results of the exponential calculation to obtain a local exponential summation result, and to calculate the global maximum value and the global exponential summation result based on the local maximum value and the local exponential summation result of the computing node and the local maximum value and the local exponential summation result of other computing nodes in the multiple computing nodes; and a parameter sharing bus, configured to send the local maximum value and the local exponential summation result of the computing node to the other computing nodes, and to receive the local maximum value and the local exponential summation result of the other computing nodes.
[0009] In the third aspect, the present application also provides a computing system composed of multiple computing nodes, which collaboratively perform exponential normalization calculations. At each of the multiple computing nodes: the local maximum and local exponential summation results of the input data of this computing node are calculated; the local maximum and local exponential summation results of this computing node are sent to other computing nodes in the multiple computing nodes, and the local maximum and local exponential summation results of the other computing nodes are received; based on the local maximum and local exponential summation results of this computing node and the local maximum and local exponential summation results of the other computing nodes, the global maximum and global exponential summation results are calculated; and based on the global maximum and global exponential summation results, the input data of this computing node is exponentially normalized.
[0010] In a fourth aspect, the present application also provides an electronic device for multiple computing nodes to collaboratively perform exponential normalization calculations, the electronic device comprising: a processor; and a memory, including a computer program code stored thereon, wherein, when the computer program code is executed by the processor, the processor implements a method for multiple computing nodes to collaboratively perform exponential normalization calculations according to an embodiment of the present application.
[0011] In a fifth aspect, the present application also provides a non-transitory computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the processor implements a method for multiple computing nodes to collaboratively perform exponential normalization calculations according to an embodiment of the present application.
[0012] In a sixth aspect, the present application also provides a computer program product, including computer instructions, which, when executed by a processor, enable the processor to implement a method for multiple computing nodes to collaboratively perform exponential normalization calculations according to an embodiment of the present application.
[0013] Compared with the method in which multiple computing nodes perform exponential normalization calculations in parallel in the related art, the method provided in the present application for multiple computing nodes to collaboratively perform exponential normalization calculations does not need to synchronize the input data of each computing node to other nodes, nor does it need to read the global input data multiple times during the calculation process. Instead, it makes full use of the characteristics of exponential calculations to calculate the local maximum value and the local exponential summation result for the local input data of each computing node, and only needs to synchronize these two local result values among multiple computing nodes, thereby improving the computational efficiency of multiple computing nodes collaboratively performing exponential normalization calculations, reducing data delays and data transmission conflicts, and saving power consumption overhead. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 A flowchart of a method for multiple computing nodes to collaboratively perform exponential normalization calculations according to an embodiment of the present application;
[0015] Figure 2 It is a structural schematic diagram of a device for multiple computing nodes to collaboratively perform exponential normalization calculation according to an embodiment of the present application;
[0016] Figure 3 It is a structural schematic diagram of a preprocessing module of an apparatus for multiple computing nodes to collaboratively perform exponential normalization calculation according to an embodiment of the present application;
[0017] Figure 4 It is a structural schematic diagram of an index calculation module of an apparatus for multiple computing nodes to collaboratively perform index normalization calculation according to an embodiment of the present application;
[0018] Figure 5It is a structural schematic diagram of a parameter sharing bus of an apparatus for multiple computing nodes to collaboratively perform exponential normalization calculation according to an embodiment of the present application;
[0019] Fig. 6A , Figure 6B , Figure 6C and Fig.6D It is a structural diagram of an example of a device for multiple computing nodes to collaboratively perform exponential normalization calculation and its main parts according to an embodiment of the present application;
[0020] Figure 7 Schematic diagram of the interconnection structure of an apparatus for multiple computing nodes to collaboratively perform exponential normalization calculations according to an embodiment of the present application. DETAILED DESCRIPTION
[0021] The technical solution of the present application is described clearly and completely below. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in the field without creative work are within the scope of protection of the present application.
[0022] Unless otherwise defined, the technical terms or scientific terms used in this application should be understood by people with ordinary skills in the field to which this application belongs. The "first", "second" and similar words used in this application do not indicate any order, quantity or importance, but are only used to distinguish different components. Similarly, similar words such as "one" or "one" do not indicate quantity restrictions, but indicate the existence of at least one. Similar words such as "connect" or "connected" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative positional relationships. When the absolute position of the described object changes, the relative positional relationship also changes accordingly.
[0023] Due to the large scale and numerous parameters of large models, it is usually necessary to train the same large model in parallel on multiple computing nodes. For example, the input data (global input data) of the large model can be split into multiple groups with the same number of computing nodes, and the input data of each group (local input data) is input to each computing node. Each computing node calculates its own local input data, thereby realizing the calculation of the global input data.
[0024] In large-scale neural networks, the exponential normalization operator (softmax) is a very important and common nonlinear operator. For global input data x0, x1, ... x N The exponential normalization calculation formula of is as follows, where x iis the global input data x0, x1, ...x N One of the input data (scalar), x m The maximum value of the global input data is the global maximum value.
[0025]
[0026] When multiple computing nodes perform exponential normalization calculations in parallel, although each computing node only needs to perform exponential normalization calculations on the input data of the node (local input data), according to the above formula for exponential normalization calculation, the global maximum value x needs to be used when performing exponential normalization calculations on local input data. m And the global exponential summation result
[0027] The exponential normalization calculation process in the related art is as follows. At each computing node, first read the input data x0, x1, ... x n And synchronize these input data to other nodes, and receive other input data x from other nodes n+1 、x n+2 ,……x N , get the global input data x0, x1, ... x N When performing exponential normalization calculation, the global input data x0, x1, ... x is read for the first time. N And calculate the global maximum x m , and then read the global input data x0, x1, ... x for the second time N And calculate the global exponential summation result Next, read the input data x0, x1, ... x n And calculate the exponential normalized value of each input data
[0028] Such an exponential normalization calculation process not only requires synchronizing the input data of each computing node to other nodes so that each node can obtain global input data, but also requires reading global input data multiple times during the calculation process. Such large-scale data synchronization and data reading will cause high latency and high power consumption. In addition, data synchronization between multiple computing nodes often uses existing data paths, which may cause data transmission conflicts and reduce computing efficiency.
[0029] In view of this, the present application proposes a method for multiple computing nodes to collaboratively perform exponential normalization calculations, which fully utilizes the characteristics of exponential calculations, calculates the local maximum value and the local exponential summation result for the local input data of each computing node, and synchronizes these two local result values among multiple computing nodes only through a dedicated parameter sharing bus, thereby improving the computational efficiency of multiple computing nodes collaboratively performing exponential normalization calculations, reducing data delays and data transmission conflicts, and saving power consumption overhead.
[0030] It should be understood that the method provided in the present application for multiple computing nodes to collaboratively perform exponential normalization calculations can be applied to a computing system including multiple computing nodes (sometimes referred to as "nodes" below), each computing node including a device for multiple computing nodes to collaboratively perform exponential normalization calculations (sometimes also referred to as "exponential normalization calculation acceleration system" below).
[0031] Next, the method, device, computing system, electronic device, storage medium and program product for multiple computing nodes to collaboratively perform exponential normalization calculation according to an embodiment of the present application will be described in detail with reference to the accompanying drawings.
[0032] Figure 1 A flowchart of a method 100 for multiple computing nodes to collaboratively perform exponential normalization calculations according to an embodiment of the present application is shown. Figure 1 In the embodiment, each step of method 100 may be executed on each computing node of the plurality of computing nodes.
[0033] like Figure 1 As shown, at step S102, the local maximum value and the local exponential summation result of the input data (local input data) of the computing node are calculated at each computing node among the multiple computing nodes.
[0034] Assume that the input data of each computing node is x0, x1, ... x n , where n≥1. First, read the input data x0~x n , calculate the input data x0~x n The maximum value of is the local maximum x m0 , and then read the input data x0~x for this computing node for the second time n , calculate the local exponential summation result E0 according to the following formula:
[0035]
[0036] Next, go to step S104.
[0037] In step S104, the local maximum value and the local exponential summation result of the computing node are sent to other computing nodes in the plurality of computing nodes, and the local maximum value and the local exponential summation result of other computing nodes are received.
[0038] For example, each computing node among the multiple computing nodes can send the local maximum value and local exponent summation results calculated by the computing node in step S102 to other computing nodes through the parameter sharing bus described later, and receive the local maximum value and local exponent summation results from other computing nodes through the parameter sharing bus, so that each computing node can obtain the local maximum value and local exponent summation results of all computing nodes.
[0039] Next, go to step S106.
[0040] In step S106, a global maximum value and a global index summation result are calculated based on the local maximum value and the local index summation result of the current computing node and the local maximum value and the local index summation result of the other computing nodes.
[0041] For example, each computing node can calculate the local maximum value x m0 As the initial global maximum x m , and use the local exponential summation result E0 of this computing node as the initial global exponential summation result E. That is, at the beginning of the stage of calculating the global maximum value and the global exponential summation result, let the global maximum value x m Equal to the local maximum value x of this computing node m0 , let the global exponential summation result E be equal to the local exponential summation result E0 of this computing node.
[0042] After receiving the local maximum value and local index summation result of one or more other computing nodes, the local maximum value of the one or more other computing nodes is added to the current global maximum value x m The largest one among them is taken as the updated global maximum value x m The number of one or more other computing nodes is not limited. For example, each time a local maximum value of another computing node is received (for example, set to x m1 ) and the sum of the local exponentials (e.g. set to E1), then the local maximum value x m1 With the current global maximum x m Compare. If x m1 >x m , then x m =x m1 , if x m1 ≤x m , then keep x m unchanged (can be understood as the updated global maximum xm With the current global maximum x m Or, for example, each time the local maximum of two other computing nodes is received (for example, set to x m1 and x m2 ) and the local exponential summation result (for example, set to E1 and E2), the two local maximum values x m1 and x m2 With the current global maximum x m Compare and take the largest of the three as the updated global maximum value x m Alternatively, the above operation may be performed after receiving the local maximum values and local exponential summation results of more other computing nodes, or after receiving the local maximum values and local exponential summation results of all other nodes.
[0043] Next, according to the updated global maximum x m , the local exponential summation results of the one or more other computing nodes and the current global exponential summation result E are not based on the global maximum value x m The calculated value is calibrated. Whether it is the local exponential summation result of other computing nodes or the current global exponential summation result, it can be expressed as Among them, for the local exponential summation results of other computing nodes, x0~x n is the input data of the node, x mk is the local maximum value of the node; for the current global exponential summation result, x0~x n is the input data of this node and the accumulated input data of other computing nodes, x mk is the current global maximum value (the global maximum value before the above update). When the updated global maximum value is equal to the current global maximum value, it can be considered that the current global exponential summation result E calculated based on the current global maximum value is based on the updated global maximum value x m Calculated, so no calibration is required. When the updated global maximum is greater than the current global maximum, the current global exponential summation result E needs to be calibrated. Similarly, if the local maximum of other computing nodes is equal to the updated global maximum, there is no need to calibrate the local exponential summation result of the node. If the local maximum of other computing nodes is less than the updated global maximum, the local exponential summation result of the node needs to be calibrated. Since the updated global maximum is always equal to at least one of the local maximum of other computing nodes and the current global maximum, the corresponding at least one exponential summation result does not need to be calibrated but is already based on the updated global maximum x m Calculated.
[0044] The calibration can be performed, for example, in the following manner.
[0045] The computing node will receive the local exponential summation results of one or more other computing nodes and the current global exponential summation result E that is not based on the updated global maximum value x m The calculated value is set to E k , and determine the value E k is based on the maximum value x mk Calculated, that is As mentioned above, for the local exponential summation results of other computing nodes, x mk is the local maximum value of the node. For the current global exponential summation result, x mk is the current global maximum value (the global maximum value before the above update). Then calculate E according to the following formula k The calibrated value E k ': Expanding this formula can be written as: It can be seen that after the global maximum value is updated as described above, by replacing the global maximum value x that is not based on the update m The calculated exponential summation result E k Multiply by the calibration factor Taking full advantage of the fact that exponential multiplication is equivalent to exponential addition in exponential calculation, we can get the updated global maximum value x m Calculated exponential sum result Thus, by replacing the global maximum x with the one not based on the above update m The calculated value is calibrated to the global maximum x based on this update m The calculation form can unify the maximum value of the minuend in the exponential part of all exponential summation results (exponential summation results of this calculation node and other calculation nodes) as accumulation objects, so that all exponential summation results can be directly accumulated regardless of their source.
[0046] After calibration, the local index summation results of the one or more other computing nodes are accumulated to the current global index summation result E to obtain an updated global index summation result E. In the case where the local index summation results of other computing nodes are calibrated, the calibrated local index summation results of the other computing nodes are accumulated to the current global index summation result E. In the case where the current global index summation result E is calibrated, the local index summation results of the other computing nodes are accumulated to the calibrated current global index summation result E.
[0047] After performing the above operations on the local maximum values and local exponential summation results of all other computing nodes, the global exponential summation result E that accumulates the local exponential summation results of all multiple computing nodes is used as the final global exponential summation result E, and the global maximum value x at this time is m As the final global maximum x m .
[0048] Next, go to step S108.
[0049] In step S108, based on the global maximum value and the global exponential summation result, the input data of the current computing node is exponentially normalized.
[0050] At each computing node, the input data x0, x1, ... x for this computing node is read for the third time. n , based on the global maximum value x calculated in step S106 m The global exponential summation result E is used to normalize the input data of this computing node according to the following formula: Expanding this formula can be written as: It can be seen that the exponential normalization calculation performed in the above manner is the same as the calculation formula of the aforementioned exponential normalization operator.
[0051] According to the method for multiple computing nodes to collaboratively perform exponential normalization calculations according to an embodiment of the present application, there is no need to synchronize all input data between all computing nodes, nor is there any need to read all input data multiple times. Instead, the local maximum value and the local exponential summation result are calculated for the local input data of each computing node, and only these two local result values need to be synchronized between multiple computing nodes. The characteristics of exponential calculations are fully utilized to calibrate the local exponential summation results to obtain the global exponential summation result, thereby improving the computational efficiency of multiple computing nodes collaboratively performing exponential normalization calculations, reducing data delays and data transmission conflicts, and saving power consumption overhead.
[0052] The method for multiple computing nodes to collaboratively perform exponential normalization calculations is described in detail below with a specific example.
[0053] In this example, the computing system includes three computing nodes, namely node 0, node 1 and node 2, and the global input data is [x0, x1, x2...x8], that is, N = 8. Among them, the input data of node 0 is [x0, x1, x2], that is, n = 2. In order to avoid confusion, the input data of node 1 and node 2 are not represented by the aforementioned local input data x0, x1, ...x nInstead of being represented by the index of the global input data, that is, the input data of node 1 is [x3, x4, x5], and the input data of node 2 is [x6, x7, x8]. However, it should be understood that when the local input data x0, x1, ... x n To represent, for node 0, its local input data x0, x1, ... x n is [x0, x1, x2]. For node 1, its local input data is x0, x1, ... x n is [x3, x4, x5]. For node 2, its local input data is x0, x1, ... x n is [x6, x7, x8].
[0054] In this example, for node 0, first, node 0 reads the input data [x0, x1, x2] of this node for the first time and calculates the local maximum value x m0 , and then read the input data of this node [x0, x1, x2] for the second time, according to the local maximum value x m0 Perform exponential calculation and sum to obtain the local exponential summation result E0. Specifically, when calculating E0, the following relationship is satisfied: Similarly, node 1 reads the input data [x3, x4, x5] of this node for the first time and calculates the local maximum value x m1 , and then read the input data of this node [x3, x4, x5] for the second time, according to the local maximum value x m1 Perform exponential calculation and sum to obtain the local exponential summation result E1. Specifically, when calculating E1, the following relationship is satisfied: Node 2 reads the input data [x6, x7, x8] of this node for the first time and calculates the local maximum value x m2 , and then read the input data of this node [x6, x7, x8] for the second time, according to the local maximum value x m2 Perform exponential calculation and sum to obtain the local exponential summation result E2. Specifically, when calculating E2, the following relationship is satisfied:
[0055] Furthermore, parameter synchronization is performed between the nodes. Specifically, node 0 sets its local maximum value x m0 The summation result E0 is sent to nodes 1 and 2, and the local maximum value x is received from node 1. m1 The local exponential summation result E1 receives the local maximum value x from node 2 m2 Node 1 and node 2 perform similar operations, so that each node obtains the local maximum value x of all nodes. m0 、x m1 、x m2And the local exponential summation results E0, E1, E2.
[0056] In one example, it is assumed that in each node, the accumulation is performed only after receiving the local maximum and local exponential summation results of all other nodes. Therefore, according to the previous description, in each node, the local maximum value x m0 、x m1 、x m2 The largest one among them is taken as the global maximum value x m Then, based on the global maximum x m Perform calibration processing.
[0057] Take node 0 as an example. First, the local maximum value x of this node is m0 Set to the initial global maximum x m , and set the local exponential summation result E0 of this node as the initial global exponential summation result E. After receiving the local maximum value x of node 1 and node 2 m1 、x m2 After the local exponential summation results E1 and E2, first compare the current global maximum value x m and x m1 、x m2 Assume that the local maximum x of node 1 is m1 Maximum, then let x m =x m1 , and get the updated x m Since the local index summation results of nodes 0 and 2 are not based on the updated x m That is x m1 If the calculation is done, the sum of the two indices needs to be calibrated. The calibration process satisfies the following relationship.
[0058]
[0059] Similarly,
[0060]
[0061] After calibration, the exponential summation results of nodes 1 and 2 can be added to the current global exponential summation result E to obtain an updated global exponential summation result E:
[0062]
[0063] At this time, since the local exponential summation results of all nodes have been accumulated, the global exponential summation result E at this time is the final global exponential summation result E. The global maximum value x at this time is m That is the final global maximum value x m .
[0064] Similarly, nodes 1 and 2 also obtain the global exponential summation result E and the global maximum value x in the above manner. m .
[0065] At each node, the global maximum value x is calculated m After summing the global exponential result E, each node reads the input data of this node for the third time and performs exponential normalization calculation. For example, node 0 reads the input data of this node [x0, x1, x2] and performs exponential normalization calculation as follows:
[0066]
[0067] Similarly, nodes 1 and 2 are also based on the global maximum x m The input data of this node is exponentially normalized with the global exponential summation result E.
[0068] In another example, assume that each node accumulates the local maximum and local exponential sum of each other node. Therefore, node 0 accumulates the local maximum x of its own node. m0 Set to the initial global maximum x m , and sets the local exponential summation result E0 of this node as the initial global exponential summation result E. When receiving the local maximum value x of node 1, m1 After the local exponential summation result E1, the current global maximum value x m The local maximum x at node 1 m1 Assume that the current global maximum value x m If x is larger, there is no need to update the global maximum value x m (It can also be considered that the updated global maximum value x m With the current global maximum x m In this case, the local exponential summation result E1 of node 1 is not based on the global maximum value x m , but based on the local maximum x of node 1 m1 Therefore, the local exponential summation result E1 of node 1 needs to be calibrated:
[0069]
[0070] The calibrated local exponential summation result E1' of node 1 is added to the current global exponential summation result E to obtain an updated global exponential summation result E:
[0071]
[0072] Next, after receiving the local maximum x at node 2 m2After the local exponential summation result E2 is calculated, the current global maximum value x m The local maximum x at node 2 m2 Assume that the local maximum x at node 2 is m2 is larger, then the global maximum value x m Update to x m2 , get the updated global maximum x m (Its value is equal to x m2 ). In this case, the current global exponential summation result E is based on the global maximum value x before the update. m Therefore, the current global exponential summation result E needs to be calibrated:
[0073]
[0074] The calibrated global exponential summation result E' is added to the local exponential summation result E2 of node 2 to obtain an updated global exponential summation result E:
[0075]
[0076] At this time, since the local exponential summation results of all nodes have been accumulated, the global exponential summation result E at this time is the final global exponential summation result E. The global maximum value x at this time is m That is the final global maximum value x m Similarly, nodes 1 and 2 also obtain the global exponential summation result E and the global maximum value x in the above manner. m The subsequent exponential normalization calculation is the same as in the previous example and will not be repeated.
[0077] It is worth noting that the above examples are only simplified examples for the sake of convenience and do not constitute a limitation on the present application. That is, the number of computing nodes is not limited to 3, but can be any number, and the number of input data can also be any number, that is, the global input data maximum index N and the local input data maximum index n can be any value greater than 1. In practical applications, there may be a large number of computing nodes and a large amount of input data.
[0078] Figure 2 The schematic diagram of the structure of the exponential normalization calculation unit 200 for multiple computing nodes to collaboratively perform exponential normalization calculation according to an embodiment of the present application is shown. The exponential normalization calculation unit 200 can also be understood as a part of the device for multiple computing nodes to collaboratively perform exponential normalization calculation, which is included in each of the multiple computing nodes of the computing system.
[0079] According to the exponential normalization calculation unit 200 (eg Fig. 6AThe exponential normalization calculation unit in the embodiment of the present invention) and the device for performing exponential normalization calculation in collaboration with multiple computing nodes include a memory (e.g. Fig. 6A system memory in the Fig. 6A System Controller in the Figure 2 As shown, the exponential normalization calculation unit 200 includes: a calculation control module 202 (eg Fig. 6A , a calculation control unit in the preprocessing module 204 (e.g. Fig. 6A preprocessing unit in), index calculation module 206 (e.g. Fig. 6A exponential calculation unit in ) and parameter sharing bus 208 (e.g. Fig. 6A parameter sharing bus in the .
[0080] The computing control module 202 is configured to parse instructions sent by a controller included in the device, and control the operation of each module in the device according to the instructions.
[0081] The preprocessing module 204 is configured to calculate the local maximum value x of the input data of the computing node m0 , and the result of the index calculation by the index calculation module 206 The local exponential summation result E0 is obtained by summing up, as well as the local maximum value and local exponential summation result of this calculation node and the local maximum value and local exponential summation result of other calculation nodes, and the global maximum value x is calculated. m And the global exponential summation result E.
[0082] The index calculation module 206 is configured to perform index calculation on the input data of the computing node. And based on the global maximum x m And the global exponential summation result E is used to exponentially normalize the input data of this computing node.
[0083] The parameter sharing bus 208 is configured to transmit the local maximum value x m0 The local exponential summation result E0 is synchronized to other computing nodes, and the local maximum value and local exponential summation results of other computing nodes are received.
[0084] For a specific example of a device for multiple computing nodes to collaboratively perform exponential normalization calculations, see Fig. 6A .like Fig. 6A As shown, in addition to the above modules, the device for multiple computing nodes to collaboratively perform exponential normalization calculation according to the embodiment of the present application may also include: a memory ( Fig. 6A The system controller in the computing node is used to store the input data of the computing node; the controller ( Fig. 6AThe system memory in the device is used to send control instructions to the exponential normalization calculation unit; the data reading unit is used to read the input data of the computing node from the memory of the device; the data writing unit is used to write the exponential normalization calculation result to the memory of the device. Fig. 6A In order to make the drawings clear, the connections between some modules / units are not shown, but it should be understood that the modules described in the specification that have information / energy transfer such as control and being controlled, data sending and receiving, etc. actually have a connection relationship.
[0085] Figure 3 FIG. 2 is a schematic diagram showing the structure of the preprocessing module 204 of the apparatus for multiple computing nodes to collaboratively perform exponential normalization calculation according to an embodiment of the present application. Figure 3 As shown, the pre-processing module 204 includes:
[0086] The control submodule includes: a local cumulative counter configured to count input data of the computing node; and a remote cumulative counter configured to count local exponential summation results of the other computing nodes;
[0087] The maximum value calculation submodule is configured to calculate the input data x0~x n The maximum value of is the local maximum x m0 , and compare the local maximum of the other computing nodes with the current global maximum x m Compare to get the updated global maximum x m , where the local maximum value x of this computation node is m0 As the initial global maximum x m ;
[0088] The calibration submodule is configured to update the global maximum x m , the local index summation results of the other computing nodes and the current global index summation result E are not based on the global maximum value x m The calculated value is calibrated, wherein the local exponential summation result E0 of the current computing node is used as the initial global exponential summation result E;
[0089] The accumulator is configured to calculate the input data x0 to x1 of the current computing node calculated by the exponential computing module. n Index Accumulate to obtain the local exponential summation result E0 of the computing node, and after the calibration submodule performs calibration, add the local exponential summation results of the other computing nodes to the current global exponential summation result E to obtain the final global exponential summation result E;
[0090] The selector is configured to select the input data of the accumulator as the input data x0~x n Index Or the local exponential summation result of the other computing nodes;
[0091] The remote wait flag register is configured as the local cumulative counter in the control submodule to count the input data x0~x of the computing node. n After the number n+1, the selector changes the input data of the accumulator from the input data x0 to x n Index Switch to the local exponential summation result of the other computing nodes.
[0092] For a specific example of a preprocessing module, see Figure 6B The control submodule of the preprocessing module can be, for example, Figure 6B The computing control unit in the embodiment includes a local cumulative counter for counting the input data of the computing node; and a remote cumulative counter for counting the local exponential summation results of other computing nodes. When calculating the exponential summation result of the present node, the local cumulative counter counts the input data of the present node, and when the number of input data is counted, it is considered that the calculation of the input data of the present node has been completed. When the local exponential summation results of other computing nodes are added to the local exponential summation result of the present node, the remote cumulative counter counts the local exponential summation results of other nodes, and when the number of other nodes is counted (the number of all computing nodes minus one), it is considered that the calculation of the data of other nodes has been completed.
[0093] The maximum value calculation submodule can be, for example, Figure 6B The maximum value calculation unit in is used to compare the size of the input data of this node to obtain the local maximum value of this node, and compare the size of the local maximum value of this node with the local maximum values of other nodes to obtain the global maximum value. Figure 6B For the sake of clarity, the input of the node data to the maximum value calculation unit is not shown, but it should be understood that when the maximum value calculation unit compares the input data of the node, the input data of the node actually needs to be input to the maximum value calculation unit. In addition, a register can be set in the maximum value calculation unit to temporarily store the current local maximum value or the current global maximum value during the comparison process.
[0094] The calibration submodule may be, for example, Figure 6B The calibration unit in is used to calibrate the local exponential summation results of other computing nodes or the current global exponential summation results as described above. The accumulator, selector and remote wait flag register can be, for example, Figure 6BThe accumulator performs all the addition operations described in the embodiments of the present application, and the selector can select whether the object of addition is the data of this node or other nodes according to the flag in the remote wait flag register. For example, the remote wait flag register can change the flag bit when the local accumulator counter counts the number of input data of this node, thereby controlling the switching of the selector. Figure 6B The connections between some modules / units are not shown, but it should be understood that the modules described in the specification that have information / energy transmission such as control and being controlled, data sending and receiving, etc. actually have a connection relationship.
[0095] Figure 4 FIG. 2 shows a schematic diagram of the structure of the index calculation module 206 of the apparatus for multiple computing nodes to collaboratively perform index normalization calculation according to an embodiment of the present application. Figure 4 As shown, the index calculation module 206 includes:
[0096] The subtraction calculation submodule is configured to calculate the input data x0 to x1 of the calculation node to which the exponential normalization calculation unit 200 belongs. n With the local maximum x m0 The difference x i -x m0 , and input data x0~x n With the global maximum x m The difference x i -x m ;
[0097] The exponential calculation submodule is configured to calculate the input data x0~x n Perform index calculations And index calculation
[0098] The multiplication submodule is configured to convert the exponential calculation result Multiply by the inverse of the global exponential summation result E to get the exponential normalization result
[0099] For a specific example of the index calculation module, see Figure 6C . Figure 6C The subtraction calculator, exponent calculator, and multiplication calculator correspond to the subtraction calculation submodule, the exponent calculation submodule, and the multiplication calculation submodule respectively.
[0100] Figure 5 FIG. 2 shows a schematic diagram of a parameter sharing bus 208 of an apparatus for collaboratively performing exponential normalization calculations on multiple computing nodes according to an embodiment of the present application. The parameter sharing bus 208 may include, for example, a parameter sharing bus interface. Figure 5 As shown, the parameter sharing bus interface may include:
[0101] a parameter sharing bus application unit, configured to send a parameter sharing bus occupation application to a parameter sharing bus arbitrator outside the device;
[0102] The parameter distribution unit is configured to notify the parameter sharing bus application unit to send the application when receiving the local maximum value and the local index summation result of the computing node from the preprocessing module, and temporarily store the local maximum value and the local index summation result of the computing node;
[0103] The parameter receiving unit is configured to receive the local maximum value and the local exponential summation result of the other computing nodes through the parameter sharing bus.
[0104] For a specific example of the parameter sharing bus interface, see Fig.6D . Fig.6D The parameter sharing bus application unit, parameter distribution unit, and parameter receiving unit in the parameter distribution unit correspond to the above-mentioned units with the same name. For example, a maximum value register and an accumulated sum register can be set in the parameter distribution unit to temporarily store the local maximum value and the local exponential summation result of the computing node respectively.
[0105] Figure 7 The interconnection structure of an apparatus for multiple computing nodes to collaboratively perform exponential normalization calculations according to an embodiment of the present application is shown.
[0106] like Figure 7 As shown, when computing nodes 0 to N respectively include exponential normalization computing units 0 to N of a device for collaborative exponential normalization computing of multiple computing nodes, the parameter sharing buses of all exponential normalization computing units are connected to a parameter sharing bus arbiter through a parameter sharing bus interface. After the parameter sharing bus arbiter receives the parameter sharing bus occupation application from the exponential normalization computing unit, it will approve an exponential normalization computing unit to occupy the parameter sharing bus according to the priority, so that the exponential normalization computing unit broadcasts the data to be synchronized (local maximum value and local exponential summation result) to other exponential normalization computing units through the parameter sharing bus. The parameter sharing buses in other exponential normalization computing units will receive the sent data and send the data to the preprocessing module for subsequent processing. In order to adapt to the different requirements of different computing systems for parameter sharing efficiency and circuit overhead, the bit width of the parameter sharing bus is variable, and can be designed to transmit two data at a time or only one data at a time according to the requirements.
[0107] It is worth noting that the exponential normalization calculation unit for multiple computing nodes to collaboratively perform exponential normalization calculations according to the present application can implement the various embodiments of the above-mentioned method for multiple computing nodes to collaboratively perform exponential normalization calculations, and can achieve the same beneficial effects, which will not be repeated here.
[0108] The present application also provides an exponential normalization calculation device for multiple computing nodes to collaboratively perform exponential normalization calculations, the exponential normalization calculation device comprising: a processor; and a memory, including a computer program code stored thereon, wherein, when the computer program code is executed by the processor, the processor implements the exponential normalization calculation method described in the embodiment of the present application.
[0109] It is worth noting that the exponential normalization calculation device can implement various embodiments of the above-mentioned exponential normalization method and can achieve the same beneficial effects, which will not be described in detail here.
[0110] The present application also provides a non-transitory computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the processor implements the exponential normalization calculation method described in the embodiment of the present application.
[0111] The present application also provides a computer program product, including computer instructions. When the computer instructions are executed by a processor, the processor implements the exponential normalization calculation method described in the embodiment of the present application.
[0112] According to the method for multiple computing nodes to collaboratively perform exponential normalization calculations according to an embodiment of the present application, there is no need to synchronize all input data between all computing nodes, nor is there any need to read all input data multiple times. Instead, the local maximum value and the local exponential summation result are calculated for the local input data of each computing node, and only these two local result values need to be synchronized between multiple computing nodes. The characteristics of exponential calculations are fully utilized to calibrate the local exponential summation results to obtain the global exponential summation result, thereby improving the computational efficiency of multiple computing nodes collaboratively performing exponential normalization calculations, reducing data delays and data transmission conflicts, and saving power consumption overhead.
[0113] The preferred specific embodiments of the present application are described in detail above. It should be understood that a person skilled in the art can make many modifications and changes based on the concept of the present application without creative work. Therefore, any technical solution that can be obtained by a person skilled in the art based on the concept of the present application through logical analysis, reasoning or limited experiments on the basis of the prior art should be within the scope of protection determined by the claims.
Claims
1. A method for multiple computing nodes to collaboratively perform exponential normalization calculations, characterized in that: The method comprises: At each computing node of the plurality of computing nodes: Calculate the local maximum value and local exponential summation result of the input data of this computing node; Sending the local maximum value and the local exponential summation result of the computing node to other computing nodes in the plurality of computing nodes, and receiving the local maximum value and the local exponential summation result of the other computing nodes; Calculate the global maximum value and the global index summation result based on the local maximum value and the local index summation result of the current computing node and the local maximum value and the local index summation result of the other computing nodes; and Based on the global maximum value and the global exponential summation results, the input data of this calculation node is exponentially normalized.
2. The method according to claim 1, characterized in that The calculation of the local maximum value and the local exponential summation result of the input data of the computing node includes: Read the input data x0~x for this computing node n , calculate the input data x0~x n The maximum value of is the local maximum x m0 , where n ≥ 1; and The second time to read the input data x0~x for this computing node n , calculate the local exponential summation result E0 according to the following formula:
3. The method according to claim 2, characterized in that The calculating of the global maximum value and the global exponential summation result based on the local maximum value and the local exponential summation result of the computing node and the local maximum value and the local exponential summation result of the other computing nodes includes: The local maximum value x of this computation node m0 As the initial global maximum x m , and use the local exponential summation result E0 of this computing node as the initial global exponential summation result E; After receiving the local maximum value and the local exponential summation result of one or more other computing nodes, the local maximum value of the one or more other computing nodes is added to the current global maximum value x m The largest one among them is taken as the updated global maximum value x m ; According to the updated global maximum x m , the local index summation results of the one or more other computing nodes and the current global index summation result E are not based on the global maximum value x m The calculated values are calibrated; After the calibration is performed, the local exponential summation results of the one or more other computing nodes are accumulated to the current global exponential summation result E to obtain an updated global exponential summation result E; The global exponential summation result E obtained by accumulating the local exponential summation results of all the multiple computing nodes is used as the final global exponential summation result E, and the global maximum value x at this time is m As the final global maximum x m .
4. The method according to claim 3, characterized in that The global maximum value x is updated according to m , the local index summation results of the one or more other computing nodes and the current global index summation result E are not based on the global maximum value x m The calculated values are calibrated to include: The local index summation results of the one or more other computing nodes and the current global index summation result E that are not based on the global maximum value x m The calculated value is set to E k , and determine the value E k is based on the maximum value x mk calculated; The value E is calculated according to the following formula k The calibrated value E k ':
5. The method according to claim 3 or 4, characterized in that: The exponential normalization of the input data of the computing node based on the global maximum value and the global exponential summation result includes: The third time, read the input data x0~x for this computing node n , based on the global maximum x m The global exponential summation result E is used to perform exponential normalization on the input data of this computing node according to the following formula:
6. A device for multiple computing nodes to collaboratively perform exponential normalization calculations, characterized in that: Each computing node of the plurality of computing nodes comprises the device, wherein the device comprises: A computing control module, configured to parse instructions sent by a controller included in the device, and control the operation of each module in the device according to the instructions; An exponential calculation module, configured to perform exponential calculation on the input data of the computing node, and to perform exponential normalization on the input data of the computing node based on a global maximum value and a global exponential summation result; a preprocessing module configured to calculate a local maximum value of input data of the computing node, and to sum the results of the exponential calculation to obtain a local exponential summation result, and to calculate the global maximum value and the global exponential summation result based on the local maximum value and the local exponential summation result of the computing node and the local maximum values and the local exponential summation results of other computing nodes in the plurality of computing nodes; and The parameter sharing bus is configured to send the local maximum value and the local exponential summation result of the computing node to the other computing nodes, and receive the local maximum value and the local exponential summation result of the other computing nodes.
7. The device according to claim 6, characterized in that The preprocessing module comprises: The control submodule includes: a local cumulative counter configured to count input data of the computing node; and a remote cumulative counter configured to count local exponential summation results of the other computing nodes; The maximum value calculation submodule is configured to calculate the input data x0~x n The maximum value of is the local maximum x m0 , and compare the local maximum of the other computing nodes with the current global maximum x m Compare to get the updated global maximum x m , where the local maximum value x of this computation node is m0 As the initial global maximum x m ; The calibration submodule is configured to update the global maximum x m , the local index summation results of the other computing nodes and the current global index summation result E are not based on the global maximum value x m The calculated value is calibrated, wherein the local exponential summation result E0 of the current computing node is used as the initial global exponential summation result E; The accumulator is configured to calculate the input data x0 to x1 of the current computing node calculated by the exponential computing module. n Index Accumulate to obtain the local exponential summation result E0 of the computing node, and after the calibration submodule performs calibration, add the local exponential summation results of the other computing nodes to the current global exponential summation result E to obtain the final global exponential summation result E; The selector is configured to select the input data of the accumulator as the input data x0~x n Index Or the local exponential summation result of the other computing nodes; The remote wait flag register is configured as the local cumulative counter in the control submodule to count the input data x0~x of the computing node. n After the number n+1, the selector changes the input data of the accumulator from the input data x0 to x n Index Switch to the local exponential summation result of the other computing nodes.
8. The device according to claim 6 or 7, characterized in that The parameter sharing bus includes a parameter sharing bus interface, and the parameter sharing bus interface includes: a parameter sharing bus application unit, configured to send a parameter sharing bus occupation application to a parameter sharing bus arbitrator outside the device; The parameter distribution unit is configured to notify the parameter sharing bus application unit to send the application when receiving the local maximum value and the local index summation result of the computing node from the preprocessing module, and temporarily store the local maximum value and the local index summation result of the computing node; The parameter receiving unit is configured to receive the local maximum value and the local exponential summation result of the other computing nodes through the parameter sharing bus.
9. A computing system composed of a plurality of computing nodes, wherein the plurality of computing nodes cooperate to perform exponential normalization calculations, characterized in that: At each computing node of the plurality of computing nodes: Calculate the local maximum value and local exponential summation result of the input data of this computing node; Sending the local maximum value and the local exponential summation result of the computing node to other computing nodes in the plurality of computing nodes, and receiving the local maximum value and the local exponential summation result of the other computing nodes; Calculate the global maximum value and the global index summation result based on the local maximum value and the local index summation result of the current computing node and the local maximum value and the local index summation result of the other computing nodes; as well as Based on the global maximum value and the global exponential summation results, the input data of this calculation node is exponentially normalized.
10. An electronic device for multiple computing nodes to collaboratively perform exponential normalization calculations, characterized in that: The electronic device comprises: Processor; and a memory including computer program code stored thereon, When the computer program code is executed by the processor, the processor is caused to implement the method according to any one of claims 1 to 5.
11. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the processor is caused to implement the method according to any one of claims 1 to 5.
Citation Information
Cited By
Operator calculation method, electronic equipment and storage medium
CN121052300A