Distributed computing error analysis method, distributed computing system, computer equipment and storage medium

By acquiring multi-order tensors in a distributed computing system and correcting error factors in real time, the problems of computational inaccuracy and system instability caused by accumulated errors are solved, achieving efficient and accurate error analysis and correction while reducing costs.

CN121144652APending Publication Date: 2025-12-16SHANGHAI SMARTLOGIC TECHNOLOGY LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511221135.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-28
Publication Date
2025-12-16

AI Technical Summary

Technical Problem

In a distributed computing environment, accumulated errors lead to inaccurate calculation results and system instability. Existing error analysis methods lack accuracy and universality, and require additional storage and computing costs.

Method used

By acquiring multi-order tensors in a distributed computing system and distributing them to various computing nodes, the computing nodes determine the product results and the correlation matrix, the correction nodes accumulate the results to determine the error factor coefficients, and correct the calculation results in real time, thus avoiding downtime and the creation of additional datasets.

Benefits of technology

It enables real-time error analysis and correction during distributed computing, improving the accuracy and universality of calculation results while reducing downtime costs and storage requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121144652A_ABST
    Figure CN121144652A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a distributed computing error analysis method, a distributed computing system, computer equipment and a storage medium, and the method comprises the steps: obtaining a multi-order tensor in a target computing instance, and distributing all elements of the multi-order tensor to all computing nodes; for each computational node, determining a first product result of each tensor element in the current computational node and a corresponding weight tensor, a second product result of a difference of each tensor element and a corresponding weight tensor, a third product result of each tensor element and a corresponding difference and weight tensor, and an action incidence matrix corresponding to each tensor element; the correction node accumulates the multi-order weight tensors according to the same dimension to obtain an accumulation result; determining an error factor coefficient corresponding to each computing node according to the accumulation result, the first product result, the second product result, the third product result and the action incidence matrix, and positioning the computing node with data deviation by performing error analysis in real time in the distributed computing process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of distributed computing or cluster computing technology, specifically to a distributed computing error analysis method, a distributed computing system, a computer device, and a storage medium. Background Technology

[0002] In a distributed computing environment, data computation is often distributed across multiple nodes within the computing system. Each node performs its assigned computational processing on the data and then transmits the data to a predetermined or calculated target node for further iterative computation, targeting itself or other nodes. During this process, errors from each node accumulate and negatively impact the system's computational accuracy and stability. Sources of error generally include: accumulated errors from computation (such as fixed-point to floating-point conversion, convolution kernel design, numerical integration, and bit-width limitations of computation and storage components) and accumulated errors from transmission (such as transmission bandwidth limitations). In this case, accumulated errors can lead to system computational errors and system downtime. Furthermore, system downtime and offline error troubleshooting result in performance degradation and downtime costs. Therefore, error analysis and control of computational results during the iterative process are crucial for improving computational accuracy, system stability, and computational performance and efficiency.

[0003] In related technologies, error analysis of computational results is typically achieved using two methods: First, establishing offline datasets. This involves creating one or more offline datasets for a single or a specific type of computational instance using methods such as generation or measurement, and periodically comparing the computational results with these offline datasets, or when errors occur. Second, training machine learning models. This involves predicting the generation and error nodes of errors for a single or a specific type of computational instance by either building a custom model or using an existing one. This involves creating one or more offline datasets for a single or a specific type of computational instance using methods such as generation or measurement, using these datasets to train a machine learning model, and then using the trained model to predict the generation and error nodes of errors for that single or a specific type of computational instance.

[0004] However, while the above two methods achieve error analysis of the computational results to some extent, they still have the following limitations: For establishing offline datasets, methods such as generation and measurement are required. The accuracy and coverage of these datasets directly affect the subsequent comparison results with real data. Furthermore, offline datasets are only for a single or specific type of computational instance; multiple datasets need to be established for different types of computational instances (common computational instance categories include molecular dynamics, quantum chemistry, and computational fluid dynamics). Additionally, offline datasets require additional storage media and equipment, incurring additional costs. Finally, comparing real data with offline datasets requires shutting down the computing system and exporting the real data for comparison, resulting in performance loss and downtime costs. For training machine learning models, it is necessary to determine a model suitable for a single or specific type of computational instance by either building a custom model or using an existing one. The model's fit and accuracy directly affect the prediction results. Moreover, machine-trained models struggle to optimize prediction results for different types of computational instances, limiting their universality. Additionally, the datasets required for training machine learning models still suffer from the limitations of offline datasets. Finally, building, calling, and training machine learning models yourself will incur additional costs. Summary of the Invention

[0005] This application provides a distributed computing error analysis method, a distributed computing system, a computer device, and a storage medium.

[0006] A first aspect of this application provides a distributed computing error analysis method applied to a distributed computing system, the distributed computing system including multiple computing nodes and at least one correction node, the method comprising:

[0007] Obtain the multi-order tensor involved in the computation in the target computation instance, and distribute all elements of the multi-order tensor to all computation nodes respectively;

[0008] For each computing node, determine the first product of each tensor element and its corresponding weight tensor, the second product of the difference of each tensor element and its corresponding weight tensor, the third product of each tensor element and its corresponding difference and weight tensor, and the action correlation matrix corresponding to each tensor element. Then send the first product, second product, third product and action correlation matrix to the correction node.

[0009] The calibration node accumulates the multi-order weight tensors along the same dimension to obtain the accumulated result;

[0010] The error factor coefficients for each computation node are determined based on the accumulated results, the first product result, the second product result, the third product result, and the effect correlation matrix.

[0011] In an optional embodiment of this application, the method further includes:

[0012] The calibration node sends the error factor coefficients corresponding to each computing node to the corresponding computing node;

[0013] Each computing node determines the quantization deviation value corresponding to each tensor element in the current computing node based on the error factor coefficient;

[0014] The calculation results of each calculation node are corrected based on the quantization deviation value corresponding to each tensor element to obtain the corrected result.

[0015] In an optional embodiment of this application, the quantization deviation value corresponding to each tensor element in the current computing node is determined based on the error factor coefficient using the following expression:

[0016]

[0017] Wherein, Δg i (t) represents the quantization deviation value corresponding to the i-th tensor element, kΔg(t) is the error factor coefficient, and g i (t) represents the i-th tensor element. This is the result of the first product. This is the cumulative result.

[0018] In an optional embodiment of this application, the calculation result of each computation node is corrected based on the quantization deviation value corresponding to each tensor element using the following expression to obtain the corrected result:

[0019]

[0020] in, For the corrected result of the i-th computation node, Δg l (t) represents the quantization bias value corresponding to the l-th tensor vector. Let l be the weight tensor vector in the i-th computation node. Let l be the l-th tensor vector in the i-th computation node.

[0021] In an optional embodiment of this application, the error factor coefficient corresponding to each computing node is determined using the following expression, based on the accumulated result, the first product result, the second product result, the third product result, and the effect correlation matrix:

[0022]

[0023] Where, k Δg (t) represents the error factor coefficient corresponding to the i-th computation node. Let l be the weight tensor vector in the i-th computation node. Let l be the l-th tensor vector in the i-th computation node. This is the result of the first product. For the cumulative result, For the result of the second product, I total_inv It is the inverse of the correlation matrix.

[0024] In an optional embodiment of this application, the following expressions are used to determine the first product of each tensor element and its corresponding weight tensor in the current computing node, the second product of the difference between each tensor element and its corresponding weight tensor, the third product of each tensor element and its corresponding difference and weight tensor, and the action correlation matrix corresponding to each tensor element:

[0025]

[0026]

[0027]

[0028] I[p][q]=g i (t)[p]*g i (t)[q]*m i

[0029] in, This is the result of the first product. This is the result of the second product. This is the result of the third product. Let l be the weight tensor vector in the i-th computation node. Let l be the l-th tensor vector in the i-th computation node. Let I[p][q] be the l-th tensor difference vector in the i-th computation node, and let g be a tensor element. i (t) corresponds to the element of the correlation matrix, g i (t)[p] is a tensor element g i (t) The value of the p-th dimension, g i (t)[q] is a tensor element g i (t) The value of the q-th dimension, m i Let i be the i-th weight tensor element.

[0030] In an optional embodiment of this application, the method further includes:

[0031] According to the preset correction step size parameter k, an error analysis is performed every k rounds of distributed computing in the distributed computing system.

[0032] A second aspect of this application provides a distributed computing system, which includes multiple computing nodes, at least one calibration node, and an allocation module.

[0033] The allocation module is used to obtain the multi-order tensor participating in the computation in the target computing instance, and allocate all elements of the multi-order tensor to all computing nodes respectively;

[0034] Each computing node is used to determine the first product result of each tensor element and its corresponding weight tensor, the second product result of the difference between each tensor element and its corresponding weight tensor, the third product result of each tensor element and its corresponding difference and weight tensor, and the action correlation matrix corresponding to each tensor element, and then send the first product result, the second product result, the third product result and the action correlation matrix to the correction node.

[0035] The calibration node is used to accumulate the multi-order weight tensors along the same dimension to obtain the accumulation result. The error factor coefficient corresponding to each calculation node is determined based on the accumulation result, the first product result, the second product result, the third product result, and the action correlation matrix.

[0036] A third aspect of this application provides a computer device, including a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the steps of any of the above distributed computing error analysis methods.

[0037] A fourth aspect of the embodiments of this application provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of the distributed computing error analysis method as described above.

[0038] Compared with the prior art, the technical solutions provided in this application have at least some or all of the following advantages:

[0039] The distributed computing error analysis method described in this application obtains the multi-order tensors involved in the calculation of the target computing instance, and distributes all elements of the multi-order tensors to all computing nodes. For each computing node, the method determines the first product result of each tensor element and its corresponding weight tensor, the second product result of the difference between each tensor element and its corresponding weight tensor, the third product result of each tensor element and its corresponding difference and weight tensor, and the action correlation matrix corresponding to each tensor element. The method then sends the first product result, the second product result, the third product result, and the action correlation matrix to the correction node. The correction node accumulates the multi-order weight tensors along the same dimension to obtain the accumulation result. Based on the accumulation result, the first product result, the second product result, the third product result, and the action correlation matrix, the method determines the error factor coefficient corresponding to each computing node. Without stopping the calculation, the method performs error analysis and correction on the calculation results in real time during the distributed computing process, which can locate computing nodes with data deviations. Attached Figure Description

[0040] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:

[0041] Figure 1 A flowchart illustrating a distributed computing error analysis method provided in one embodiment of this application;

[0042] Figure 2 A flowchart illustrating a distributed computing error correction method provided in one embodiment of this application;

[0043] Figure 3 This is a schematic diagram of a distributed computing system architecture provided in one embodiment of this application;

[0044] Figure 4 This is a schematic diagram of a computer device structure provided in one embodiment of this application. Detailed Implementation

[0045] In the process of developing this application, the inventors discovered that current methods for generating distributed computing error analysis are poor in both accuracy and universality.

[0046] To address the aforementioned issues, this application provides a distributed computing error analysis method, a distributed computing system, a computer device, and a storage medium to improve the accuracy and versatility of distributed computing error analysis.

[0047] The solutions in this application embodiment can be implemented using various computer languages, such as the object-oriented programming language Java and the interpreted scripting language JavaScript.

[0048] To make the technical solutions and advantages of the embodiments of this application clearer, the exemplary embodiments of this application will be described in further detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not an exhaustive list of all embodiments. It should be noted that, unless otherwise specified, the embodiments and features in the embodiments of this application can be combined with each other.

[0049] Please see Figure 1 The distributed computing error analysis method provided in this application is applied to a distributed computing system, which includes multiple computing nodes and at least one correction node. The method includes the following steps S100 to S400:

[0050] S100: Obtain the multi-order tensor involved in the computation in the target computation instance, and distribute all elements of the multi-order tensor to all computation nodes respectively;

[0051] S200, for each computing node, determine the first product result of each tensor element and its corresponding weight tensor, the second product result of the difference between each tensor element and its corresponding weight tensor, the third product result of each tensor element and its corresponding difference and weight tensor, and the action correlation matrix corresponding to each tensor element, and send the first product result, the second product result, the third product result and the action correlation matrix to the correction node;

[0052] S300, the calibration node accumulates the multi-order weight tensors along the same dimension to obtain the accumulation result;

[0053] S400, determine the error factor coefficient corresponding to each calculation node based on the accumulated result, the first product result, the second product result, the third product result, and the effect correlation matrix.

[0054] In an optional embodiment of this application, if the distributed computing system is a 64-bit computing system with m cores, each core having n arithmetic logic units (ALUs), and the cores are physically interconnected, then each ALU can be numbered as ALU. 00 ALU 01 ALU 02 ...ALU mn In ALU 00 ALU 01 ALU 02 ...ALU mn In this context, at least one arithmetic logic unit (ALU) can be arbitrarily designated as a correction node to execute the error analysis and correction process. When the system's software flow is not in the "error analysis and correction process," the correction node performs the assigned computational tasks in the same way as other ALUs. 00 ALU01 ALU 02 ...ALU mn As a computing node. Alternatively, the calibration node can be specified in any computing node of the computing system, or selected from any device or board attached to or supplemented to the computing system that has computing capabilities, meets computing performance requirements, and is capable of implementing the method of this embodiment.

[0055] In an optional embodiment of this application, if it is known that the current computation instance contains one j-th order tensor participating in the computation... Wherein, the tensor vector is denoted as t represents the mapping relationship in the computation instance; its corresponding j-th order weight tensor is denoted as The weight tensor vector is denoted as Determine the ALU for each compute node 00 ALU 01 ALU 02 ...ALU mn The results of the first product of each tensor element and its corresponding weight tensor, the second product of the difference between each tensor element and its corresponding weight tensor, the third product of each tensor element, its corresponding difference, and its weight tensor, and the action correlation matrix corresponding to each tensor element are all included. The action correlation matrix consists of the j-order tensors distributed across the computation nodes in the current computation round. Inter-element interactions, ALUs of each computing node 00 ALU 01 ALU 02 …ALU mn For the tensor vectors in it Construct a j*j correlation matrix I j×j The order of the constructed I matrix and its tensor Order-related, for example, If the order is 3, then the constructed I matrix is ​​also of order 3. (3rd order tensor) I 3×3 The specific method for constructing the matrix is ​​as follows, and this method can be generalized to j-order tensors. I j×j Matrix construction:

[0056] I[1][1]=g i (t)[1]*g i (t)[1]*m i

[0057] I[1][2]=g i (t)[1]]*g i (t)[2]*m i

[0058] I[1][3]=gi (t)[1]*g i (t)[3]*m i

[0059] I[2][1]=g i (t)[2]*g i (t)[1]*m i

[0060] I[2][2]=g i (t)[2]]*g i (t)[2]*m i

[0061] I[2][3]=g i (t)[2]*g i (t)[3]*m i

[0062] I[3][1]=g i (t)[3]*g i (t)[1]*m i

[0063] I[3][2]=g i (t)[3]]*g i (t)[2]*m i

[0064] I[3][3]=g i (t)[3]*g i (t)[3]*m i

[0065] Where I[1][1] is I 3×3 The element in the first row and first column of the matrix, g i (t)[1] is a 3rd order tensor medium element g i The first dimension element of (t), g i (t)[2] is a third-order tensor medium element g i The second dimension element of (t), g i (t)[3] is a third-order tensor medium element g i The third dimension element of (t), m i For tensor element g i The weight tensor element corresponding to (t) in the 3rd order tensor Including the x-axis, y-axis, and z-axis directions, the elements of the first dimension are the elements of the x-axis, the elements of the second dimension are the elements of the y-axis, and the elements of the third dimension are the elements of the z-axis.

[0066] The distributed computing error analysis method of this application does not require stopping the computation or establishing an offline dataset. By continuously or periodically applying this method, error analysis and correction of the computation results are performed in real time during the computation iteration process. It can also locate the computation nodes with data deviations. At the same time, this method directly processes the computation data of each computation node. This process is not limited to a certain type of computation instance and has a certain degree of universality.

[0067] In an optional embodiment of this application, see [link to relevant documentation]. Figure 2 The method further includes:

[0068] S210, the calibration node sends the error factor coefficients corresponding to each computing node to the corresponding computing node;

[0069] S220, each computing node determines the quantization deviation value corresponding to each tensor element in the current computing node according to the error factor coefficient;

[0070] S230 corrects the calculation results of each calculation node based on the quantization deviation value corresponding to each tensor element, and obtains the corrected result.

[0071] The distributed computing error analysis method of this application can correct each tensor element in each computing node by using the error factor coefficient corresponding to each computing node. This can achieve real-time correction in the distributed computing process and obtain accurate calculation results simply and quickly.

[0072] In an optional embodiment of this application, in step S220, the quantization deviation value corresponding to each tensor element in the current computing node is determined according to the error factor coefficient using the following expression:

[0073]

[0074] Wherein, Δg i (t) represents the quantization bias value corresponding to the i-th tensor element, k Δg (t) is the error factor coefficient, g i (t) represents the i-th tensor element. This is the result of the first product. This is the cumulative result.

[0075] The distributed computing error analysis method of this application calculates the quantization deviation value of each tensor element directly by using the first product result, error factor coefficient, accumulation result and tensor element corresponding to each tensor element in the current computing node when each computing node receives the error factor coefficient. This allows the quantization deviation value to be used to correct each tensor element in the distributed computing process before the distributed computing is performed, thus ensuring that the distributed computing result obtained is the corrected computing result.

[0076] In an optional embodiment of this application, in step S230, the calculation result of each computation node is corrected according to the quantization deviation value corresponding to each tensor element using the following expression to obtain the corrected result:

[0077]

[0078] in, For the corrected result of the i-th computation node, Δg l (t) represents the quantization bias value corresponding to the l-th tensor vector. Let l be the weight tensor vector in the i-th computation node. Let l be the l-th tensor vector in the i-th computation node.

[0079] The distributed computing error analysis method of this application corrects each tensor element using quantization deviation values ​​while performing distributed computing, thus avoiding problems such as stopping the computing due to correction after obtaining the distributed computing results.

[0080] In an optional embodiment of this application, in step S400, the error factor coefficient corresponding to each computing node is determined based on the accumulated result, the first product result, the second product result, the third product result, and the effect correlation matrix using the following expression:

[0081]

[0082] Where, k Δg (t) represents the error factor coefficient corresponding to the i-th computation node. Let l be the weight tensor vector in the i-th computation node. Let l be the l-th tensor vector in the i-th computation node. This is the result of the first product. For the cumulative result, For the result of the second product, I total_inv It is the inverse of the correlation matrix.

[0083] The distributed computing error analysis method of this application uses the error coefficient factor as a quantitative value of the impact of the deviation of the computing results of each computing node on the entire computing instance. The larger the value of the error coefficient factor, the greater the impact of the deviation of the computing node's computing results on the entire computing instance. By analyzing the error coefficient factor, the effect of quantifying the computing deviation and analyzing the error can be achieved. In addition, since the calculation data of the error coefficient factor comes from the data sent by all computing nodes received by the correction node, the correction node can reverse the calculation to locate which computing node the data deviation comes from, thereby realizing the location of the source node of the error.

[0084] In an optional embodiment of this application, in step S200, the following expressions are used to determine the first product result of each tensor element in the current computing node with its corresponding weight tensor, the second product result of the difference between each tensor element and its corresponding weight tensor, the third product result of each tensor element with its corresponding difference and weight tensor, and the action correlation matrix corresponding to each tensor element:

[0085]

[0086]

[0087] I[p][q]=g i (t)[p]*g i (t)[q]*m i

[0088] in, This is the result of the first product. This is the result of the second product. This is the result of the third product. Let l be the weight tensor vector in the i-th computation node. Let l be the l-th tensor vector in the i-th computation node. Let I[p][q] be the l-th tensor difference vector in the i-th computation node, and let g be a tensor element. i (t) corresponds to the element of the correlation matrix, g i (t)[p] is a tensor element g i (t) The value of the p-th dimension, g i (t)[q] is a tensor element g i (t) The value of the q-th dimension, m i For tensor element g i (t) corresponds to the weight tensor element, p∈[1, ..., t) i ],q∈[1, i ] From tensor element g i (t) corresponds to the elements of the correlation matrix that form the tensor element g. i The I matrix corresponding to (t) is obtained by summing the I matrices corresponding to all tensor elements according to their dimensions. total , as the correlation matrix.

[0089] The distributed computing error analysis method of this application determines, for each computing node, the first product result of each tensor element and its corresponding weight tensor, the second product result of the difference of each tensor element and its corresponding weight tensor, the third product result of each tensor element and its corresponding difference and weight tensor, and the action correlation matrix corresponding to each tensor element as the calculation parameters of the error coefficient factor, thus obtaining an accurate error coefficient factor.

[0090] In an optional embodiment of this application, the method further includes:

[0091] According to the preset correction step size parameter k, an error analysis is performed every k rounds of distributed computing in the distributed computing system. Parameter k is used to balance computational efficiency and accuracy, and its value can be determined based on specific computing instances or experience. Furthermore, for computing instances or systems with high stability, a longer step size can be used to reduce computational costs and balance computational efficiency; conversely, for computing instances or systems with sophisticated computational characteristics, a shorter step size can be used.

[0092] The distributed computing error analysis method of this application pre-sets corresponding correction step size parameters for different computing instances, which can ensure the computing efficiency of distributed computing while taking into account the computing accuracy of distributed computing.

[0093] In an optional embodiment of this application, the distributed computing error analysis method of this application is applied to the determination method of electrostatic interaction force in obtaining the short-range electrostatic interaction of each atom and obtaining the target electrostatic interaction force. Multi-order tensors are constructed based on atomic charge and atomic position, respectively, to analyze and correct the error in the electrostatic interaction calculation in the distributed nodes. The method for determining the electrostatic interaction force includes:

[0094] The target molecular dynamics simulation system is divided into multiple three-dimensional grids, where each three-dimensional grid is treated as a node. Each node contains a computational unit and a storage unit, and each node contains at least one atom. Different nodes communicate with each other through point-to-point direct connection channels, and edge nodes are periodically connected to form a closed loop.

[0095] For each atom in the target molecular dynamics simulation system, the current atom is taken as the first atom and traversed with the specified second atom to determine the electrostatic force between the first atom and the second atom. The electrostatic force between the non-electrostatic pairs of atoms in the first atom and the second atom is canceled to obtain the target electrostatic force, which is taken as the long-range electrostatic force of the current atom. The second atom diffuses around the first atom and the diffusion radius is within a preset value. All atoms involved in the rigid atomic group are stored in the same storage unit.

[0096] The long-range electrostatic interaction of each atom is corrected based on the long-range electrostatic interaction of all atoms in the target molecular dynamics simulation system, and the corrected long-range electrostatic interaction is obtained.

[0097] Obtain the short-range electrostatic interaction of each atom, and sum the corrected long-range electrostatic interaction of each atom with the short-range electrostatic interaction. The sum is taken as the electrostatic interaction force of each atom.

[0098] In an optional embodiment of this application, the designated second atom is obtained through the following steps:

[0099] For each atom other than the first atom, atoms whose distance from the first atom is greater than a preset value are considered as atoms to be eliminated.

[0100] The atoms in the target molecular dynamics simulation system other than the first atom and the atom to be eliminated are designated as the second atom.

[0101] In an optional embodiment of this application, the long-range electrostatic interactions of all atoms in the target molecular dynamics simulation system are obtained through the following steps:

[0102] The target molecular dynamics simulation system is divided into multiple regions by multiple three-dimensional grids. The cumulative value of long-range electrostatic interactions within each region is collected. The number of regions is less than the number of three-dimensional grids.

[0103] The summation of the long-range electrostatic interactions within each region in all regions yields the total long-range electrostatic interaction of the target molecular dynamics simulation system, which is then used as the long-range electrostatic interaction of all atoms in the target molecular dynamics simulation system.

[0104] In an optional embodiment of this application, collecting the cumulative value of long-range electrostatic interactions within each region includes:

[0105] The receiving function of each computing unit within the configuration area is interrupted;

[0106] The broadcast function between computing units in the configuration area is interrupted;

[0107] For each region, each computing unit within the current region sends long-range electrostatic discharge (ESD) signals to the same designated computing unit. The designated computing unit accumulates the received ESD signals. When the number of ESD signals received by the designated computing unit reaches the first specified number in the receive function interruption configuration, the designated computing unit stops receiving ESD signals, obtaining the cumulative value of long-range ESD signals for the current region, and clearing the interruption configuration of the receive function for each computing unit within the region.

[0108] The summation of the cumulative long-range electrostatic effects within each region of all regions includes:

[0109] Each designated computing unit broadcasts the cumulative value of long-range electrostatic discharge (ESD) within the region to the fixed address of designated computing units in other regions. The designated computing unit then accumulates the broadcast ESD. When the number of ESD received by the designated computing unit reaches the second specified number in the broadcast function interruption configuration, the cumulative value of long-range ESD between regions is obtained, and the broadcast function interruption configuration of each computing unit in the region is cleared.

[0110] In an optional embodiment of this application, the step of correcting the long-range electrostatic interaction of each atom based on the long-range electrostatic interactions of all atoms in the target molecular dynamics simulation system to obtain the corrected long-range electrostatic interaction includes:

[0111] The long-range correction force for each atom is determined based on the long-range electrostatic interactions of all atoms in the target molecular dynamics simulation system using the following expression:

[0112]

[0113] Among them, F cor The long-range correction force for each atom, F tot Let N represent the long-range electrostatic interactions of all atoms in the target molecular dynamics simulation system, and N be the number of atoms in the target molecular dynamics simulation system.

[0114] The long-range correction force is broadcast regionally to each computational unit. The long-range electrostatic interaction of each atom is corrected using the following expression, resulting in the corrected long-range electrostatic interaction:

[0115] F v,long =F long -F cor

[0116] Among them, F v,long For the corrected long-range electrostatic interaction of each atom, F long For the long-range electrostatic interaction before correction for each atom, F cor The long-range correction force for each atom.

[0117] In an optional embodiment of this application, the following expression is used to determine the electrostatic force between the first atom and the second atom by traversing the current atom as the first atom and a specified second atom, and to cancel the electrostatic force between the non-electrostatic pairs of atoms in the first and second atoms, thereby obtaining the target electrostatic force, including:

[0118]

[0119] Among them, F long Let f be the target electrostatic force, α be the electrostatic conversion factor, and q be the Gaussian cloud factor.i q j The charges carried by atoms i and j are respectively, r ij Let N be the distance between atoms i and j, N be the total number of atoms, and k be the distance between atoms i and j. x k y k z L is the number of lattice points for three-way diffusion. x L y L z To simulate the side length of the box.

[0120] In an optional embodiment of this application, the short-range electrostatic interaction of each atom is obtained through the following expression, including:

[0121]

[0122] Among them, F short Let f be the short-range electrostatic interaction for each atom, f be the electrostatic conversion factor, α be the Gaussian cloud factor, and q be the short-range electrostatic interaction for each atom. i q j The charges carried by atoms i and j are respectively, r ij Let N be the distance between atoms i and j, N be the total number of atoms, and k be the distance between atoms i and j. x k y k z L is the number of lattice points for three-way diffusion. x L y L z To simulate the side length of the box, erfc(x) is the error complementation function.

[0123] In an optional embodiment of this application, the distributed computing error analysis method of this application is applied to the three-dimensional FFT calculation method. The three-dimensional FFT points are used as elements of a third-order tensor to construct a third-order tensor, and the errors in the FFT calculation of the distributed nodes are analyzed and corrected. The FFT calculation method includes:

[0124] The target 3D data is preprocessed to make each dimension of the target 3D data m*2. n Multiples of, where the array in the microcode memory is 2 n ×2 n ×2 n m is the number of microcode storage devices. The preprocessed three-dimensional data includes first-dimensional data, second-dimensional data, and third-dimensional data.

[0125] The data in the first dimension are divided and FFT calculated sequentially according to radix-5, radix-3, and radix-2 to obtain the FFT calculation result of the first dimension.

[0126] The FFT calculation result of the first dimension is used as the data of the second dimension. The data of the second dimension is divided and FFT calculated in turn according to the radix-5, radix-3 and radix-2 to obtain the FFT calculation result of the second dimension.

[0127] The second-dimensional FFT calculation result is used as the third-dimensional data. The third-dimensional data is then divided and FFT calculated sequentially according to radix 5, radix 3, and radix 2 to obtain the third-dimensional FFT calculation result, which is used as the three-dimensional FFT calculation result.

[0128] In an optional embodiment of this application, the step of sequentially partitioning and FFT-computing the first-dimensional data according to radix-5, radix-3, and radix-2 to obtain the first-dimensional FFT calculation result includes:

[0129] Determine if the number of points in the first dimension is divisible by 5;

[0130] If the number of points in the first dimension data is divisible by 5, the first dimension data is divided according to radix 5 to obtain the t1th batch of 5 radix 5 subsequences. The t1th batch of 5 radix 5 subsequences is transformed into the frequency domain using the radix 5 FFT algorithm, which is used as the result of the t1th batch of 5 radix 5 subsequences, where t1 is the number of times the data is divided according to radix 5.

[0131] If the number of data points in the first dimension is not divisible by 5, or if the results of the 5 groups of radix-5 subsequences in the t1th batch are obtained, determine whether the number of points in the first dimension data or the results of each group of radix-5 subsequences in the t1th batch is divisible by 3.

[0132] If the number of points in the first dimension data or the result of each radix-5 subsequence in the t1th batch is divisible by 3, the first dimension data or the result of each radix-5 subsequence in the t1th batch is divided according to radix-3 to obtain the 3 radix-3 subsequences in the t2th batch. The 3 radix-3 subsequences in the t2th batch are transformed into the frequency domain using the radix-3 FFT algorithm, which is used as the result of the 3 radix-3 subsequences in the t2th batch, where t2 is the number of times the data is divided according to radix-3.

[0133] If the number of data points in the first dimension is not divisible by 3, or if the results of the 3 groups of radix-3 subsequences in the t2th batch are obtained, determine whether the number of points in the first dimension data or the results of each group of radix-3 subsequences in the t2th batch is divisible by 2.

[0134] If the number of points in the first dimension data or each radix-3 subsequence result is divisible by 2, the first dimension data or each radix-3 subsequence result is divided according to radix-2 to obtain the t3th batch of 2 radix-2 subsequences. The t3th batch of 2 radix-2 subsequences is transformed into the frequency domain using the radix-2 FFT algorithm and used as the first dimension FFT calculation result, where t3 is the number of times the data is divided according to radix-2.

[0135] In an optional embodiment of this application, the step of dividing the first dimension data according to radix-5 to obtain the t1th batch of 5-group radix-5 subsequences, and using a radix-5 FFT algorithm to transform the t1th batch of 5-group radix-5 subsequences to the frequency domain as the result of the t1th batch of 5-group radix-5 subsequences, includes:

[0136] Step 1: Divide the data in the first dimension according to radix-5 to obtain 5 radix-5 subsequences;

[0137] Step 2: Use the radix-5 FFT algorithm to transform the 5 radix-5 subsequences into the frequency domain to obtain the first batch of 5 radix-5 subsequence results;

[0138] Step 3: Determine whether the number of points in each of the first batch of base-5 subsequence results is divisible by 5;

[0139] Step 4: If the number of points in each group of radix-5 subsequences in the first batch is divisible by 5, divide each group of radix-5 subsequences in the first batch according to radix-5 to obtain the first, second, third, fourth and fifth subsequence segments of each group of radix-5 subsequences in the first batch. Concatenate the first subsequence segments corresponding to the first to fifth groups of radix-5 subsequences in the first batch to obtain the first group of radix-5 subsequences in the second batch. Concatenate the second subsequence segments corresponding to the first to fifth groups of radix-5 subsequences in the first batch to obtain the second group of radix-5 subsequences in the second batch, and so on, until the second batch of 5 groups of radix-5 subsequences is obtained.

[0140] Step 5: Use the radix-5 FFT algorithm to transform the second batch of 5 radix-5 subsequences to the frequency domain to obtain the second batch of 5 radix-5 subsequence results. Return to step 3 until the number of points in each radix-5 subsequence result is no longer divisible by 5, and obtain the t1th batch of 5 radix-5 subsequence results.

[0141] In an optional embodiment of this application, the first batch of 5 radix-5 subsequences is obtained by transforming the 5 radix-5 subsequences to the frequency domain using the radix-5 FFT algorithm through the following expression:

[0142]

[0143] Where, X(k1), and The results for the first batch of 5 radix-5 subsequences are as follows: A1, B1, C1, D1, and E1 are the first batch of 5 radix-5 subsequences. and As a base 5 butterfly factor, and radix is ​​the base-5 adjustment factor, where radix is ​​the base number.

[0144] In an optional embodiment of this application, the step of dividing the first dimension data or the radix-5 subsequence results of each group in the t1 batch according to radix-3 to obtain the 3 groups of radix-3 subsequences in the t2 batch includes:

[0145] Step 1: Divide the data of the first dimension into radix 3 or the radix 5 subsequence results of each group in the t1 batch according to radix 3, to obtain 3 groups of radix 3 subsequences;

[0146] Step 2: Use the radix-3 FFT algorithm to transform the three radix-3 subsequences into the frequency domain to obtain the first batch of three radix-3 subsequence results;

[0147] Step 3: Determine whether the number of points in each of the first batch of base-3 subsequence results is divisible by 3;

[0148] Step 4: If the number of points in each group of radix-3 subsequences in the first batch is divisible by 3, divide each group of radix-3 subsequences in the first batch according to radix-3 to obtain the first, second, and third subsequence segments after the division of each group of radix-3 subsequences in the first batch. Concatenate the first subsequence segments corresponding to the first to third groups of radix-3 subsequences in the first batch to obtain the first group of radix-3 subsequences in the second batch. Concatenate the second subsequence segments corresponding to the first to third groups of radix-3 subsequences in the first batch to obtain the second group of radix-3 subsequences in the second batch, and so on, until the second batch of 3 groups of radix-3 subsequences is obtained.

[0149] Step 5: Use the radix-3 FFT algorithm to transform the second batch of 3 groups of radix-3 subsequences into the frequency domain to obtain the results of the second batch of 3 groups of radix-3 subsequences. Return to step 3 and continue until the number of points in each group of radix-3 subsequences is no longer divisible by 3, to obtain the results of the t2th batch of 3 groups of radix-3 subsequences.

[0150] In an optional embodiment of this application, the three sets of radix-3 subsequences are transformed to the frequency domain using the radix-3 FFT algorithm through the following expression to obtain the first batch of three sets of radix-3 subsequence results:

[0151]

[0152] Where X(k2), and The results for the first batch of three radix-3 subsequences are shown. A2, B2, and C2 are the first batch of three radix-3 subsequences. and As a base 3 butterfly factor, and The base is 3, which is the adjustment factor.

[0153] In an optional embodiment of this application, the step of dividing the first-dimensional data or each group of radix-3 subsequences according to radix-2 to obtain the t3th batch of 2 groups of radix-2 subsequences, and using the radix-2 FFT algorithm to transform the t3th batch of 2 groups of radix-2 subsequences into the frequency domain as the first-dimensional FFT calculation result includes:

[0154] Step 1: Divide each group of base 3 subsequences into two groups of base 2 subsequences according to base 2.

[0155] Step 2: Use the radix-2 FFT algorithm to transform the two sets of radix-2 subsequences into the frequency domain to obtain the first batch of two sets of radix-2 subsequence results;

[0156] Step 3: Determine whether the number of points in each base-2 subsequence result of the first batch is divisible by 2;

[0157] Step 4: If the number of points in each group of radix-2 subsequences in the first batch is divisible by 2, divide each group of radix-2 subsequences according to radix-2 to obtain the first and second subsequence segments of each group of radix-2 subsequences in the first batch. Concatenate the first subsequence segments corresponding to the first and second groups of radix-2 subsequences in the first batch to obtain the first group of radix-2 subsequences in the second batch. Concatenate the second subsequence segments of the first and second groups of radix-2 subsequences in the first batch to obtain the second group of radix-2 subsequences in the second batch, thus obtaining the second batch of 2 groups of radix-2 subsequences.

[0158] Step 5: Use the radix-2 FFT algorithm to transform the second batch of two groups of radix-2 subsequences to the frequency domain, obtaining the results of the second batch of two groups of radix-2 subsequences. Return to step 3, and repeat until the number of points in each group of radix-2 subsequences is no longer divisible by 2, to obtain the results of the t3th batch of two groups of radix-2 subsequences.

[0159] The following expression is used to transform the two sets of radix-2 subsequences to the frequency domain using the radix-2 FFT algorithm, resulting in the first batch of two sets of radix-2 subsequences:

[0160]

[0161] Where X(k3) and These are the results of the first two sets of radix-2 subsequences. A3 and B3 are the first two sets of radix-2 subsequences. The base is a butterfly factor.

[0162] It should be understood that although the steps in the flowchart are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order constraint on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the diagram may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.

[0163] Please see Figure 3 One embodiment of this application provides a distributed computing system 300, which includes multiple computing nodes, at least one calibration node, and an allocation module.

[0164] The allocation module 310 is used to obtain the multi-order tensor participating in the calculation in the target computing instance and allocate all elements of the multi-order tensor to all computing nodes respectively;

[0165] Each computing node 320 is used to determine the first product result of each tensor element and its corresponding weight tensor, the second product result of the difference between each tensor element and its corresponding weight tensor, the third product result of each tensor element and its corresponding difference and weight tensor, and the action correlation matrix corresponding to each tensor element, and sends the first product result, the second product result, the third product result and the action correlation matrix to the correction node.

[0166] The correction node 330 is used to accumulate the multi-order weight tensors along the same dimension to obtain the accumulation result. Based on the accumulation result, the first product result, the second product result, the third product result, and the action correlation matrix, the error factor coefficient corresponding to each calculation node is determined.

[0167] For specific limitations regarding the aforementioned system 300, please refer to the limitations on the distributed computing error analysis method described above, which will not be repeated here. Each module in the aforementioned system 300 can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in the computer device in hardware form, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to each module.

[0168] In one embodiment, a computer device is provided, the internal structure of which can be as follows: Figure 4As shown. The computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system, computer programs, and the database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The database stores data. The network interface communicates with external terminals via a network connection. When the computer program is executed by the processor, it implements the distributed computing error analysis method described above. It includes: memory and a processor; the memory stores the computer program; and the processor executes the computer program to implement any step of the distributed computing error analysis method described above.

[0169] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, can perform any of the steps in the distributed computing error analysis method described above.

[0170] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0171] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A system that specifies functions in one or more boxes.

[0172] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including an instruction set implemented in a process. Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0173] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0174] Although preferred embodiments of this application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of this application.

[0175] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.

Claims

1. A distributed computing error analysis method, characterized in that, Applied to a distributed computing system, the distributed computing system including multiple computing nodes and at least one calibration node, the method includes: Obtain the multi-order tensor involved in the computation in the target computation instance, and distribute all elements of the multi-order tensor to all computation nodes respectively; For each computing node, determine the first product of each tensor element and its corresponding weight tensor, the second product of the difference of each tensor element and its corresponding weight tensor, the third product of each tensor element and its corresponding difference and weight tensor, and the action correlation matrix corresponding to each tensor element. Then send the first product, second product, third product and action correlation matrix to the correction node. The calibration node accumulates the multi-order weight tensors along the same dimension to obtain the accumulated result; The error factor coefficients for each computation node are determined based on the accumulated results, the first product result, the second product result, the third product result, and the effect correlation matrix.

2. The method according to claim 1, characterized in that, The method further includes: The calibration node sends the error factor coefficients corresponding to each computing node to the corresponding computing node; Each computing node determines the quantization deviation value corresponding to each tensor element in the current computing node based on the error factor coefficient; The calculation results of each calculation node are corrected based on the quantization deviation value corresponding to each tensor element to obtain the corrected result.

3. The method according to claim 2, characterized in that, The following expression is used to determine the quantization deviation value corresponding to each tensor element in the current computing node based on the error factor coefficient: Wherein, Δg i (t) represents the quantization bias value corresponding to the i-th tensor element, k Δg(t) For the error factor coefficient, g i (t) represents the i-th tensor element. This is the result of the first product. This is the cumulative result.

4. The method according to claim 2, characterized in that, The calculation results of each computation node are corrected based on the quantization deviation values ​​corresponding to each tensor element using the following expression, resulting in the corrected result: in, For the corrected result of the i-th computation node, Δg l (t) represents the quantization bias value corresponding to the l-th tensor vector. Let l be the weight tensor vector in the i-th computation node. Let l be the l-th tensor vector in the i-th computation node.

5. The method according to claim 1, characterized in that, The error factor coefficients for each computation node are determined using the following expression, based on the accumulated results, the first product result, the second product result, the third product result, and the effect correlation matrix: Where, k Δg(t) The error factor coefficients corresponding to the i-th computation node are... Let l be the weight tensor vector in the i-th computation node. Let l be the l-th tensor vector in the i-th computation node. This is the result of the first product. For the cumulative result, For the result of the second product, I total_inv It is the inverse of the correlation matrix.

6. The method according to claim 1, characterized in that, The following expressions determine the first product of each tensor element and its corresponding weight tensor in the current computation node, the second product of the difference between each tensor element and its corresponding weight tensor, the third product of each tensor element and its corresponding difference and weight tensor, and the action correlation matrix corresponding to each tensor element: in, This is the result of the first product. This is the result of the second product. This is the result of the third product. Let l be the weight tensor vector in the i-th computation node. Let l be the l-th tensor vector in the i-th computation node. Let I[p][q] be the l-th tensor difference vector in the i-th computation node, and let g be a tensor element. i (t) corresponds to the element of the correlation matrix, g i (t)[p] is a tensor element g i (t) The value of the p-th dimension, g i (t)[q] is a tensor element g i (t) The value of the q-th dimension, m i Let i be the i-th weight tensor element.

7. The method according to claim 1, characterized in that, The method further includes: According to the preset correction step size parameter k, an error analysis is performed every k rounds of distributed computing in the distributed computing system.

8. A distributed computing system, characterized in that, This distributed computing system includes multiple computing nodes, at least one calibration node, and an allocation module. The allocation module is used to obtain the multi-order tensor participating in the computation in the target computing instance, and allocate all elements of the multi-order tensor to all computing nodes respectively; Each computing node is used to determine the first product result of each tensor element and its corresponding weight tensor, the second product result of the difference between each tensor element and its corresponding weight tensor, the third product result of each tensor element and its corresponding difference and weight tensor, and the action correlation matrix corresponding to each tensor element, and then send the first product result, the second product result, the third product result and the action correlation matrix to the correction node. The calibration node is used to accumulate the multi-order weight tensors along the same dimension to obtain the accumulation result. The error factor coefficient corresponding to each calculation node is determined based on the accumulation result, the first product result, the second product result, the third product result, and the action correlation matrix.

9. A computer device, comprising: The system includes a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the distributed computing error analysis method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the distributed computing error analysis method according to any one of claims 1 to 7.