Distributed matrices multiplication method rooted in network structure of distributed computing system

The distributed matrix multiplication operation method addresses the challenges of accuracy and efficiency in high-dimensional matrix operations by using group algebra-based coded computing, which improves speed, accuracy, and load handling even in noisy wireless environments.

WO2025105928A1PCT designated stage expired Publication Date: 2025-05-22SEOUL NATIONAL UNIVERSITY R&DB FOUNDATION
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/096558
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-09-05
Filing Date
2024-11-14
Publication Date
2025-05-22

AI Technical Summary

Technical Problem

Existing distributed computing methods for high-dimensional matrix multiplication face challenges in accuracy and efficiency due to communication constraints, noise in wireless environments, and the assumption of error-free communication links.

Method used

A distributed matrix multiplication operation method utilizing group algebra-based coded computing, which involves a master node and worker nodes, where the master node generates partitioning functions and broadcasts them through a wireless network, enabling worker nodes to perform encoded matrix multiplication operations and restore the original results.

Benefits of technology

This method improves the speed, accuracy, and load handling of high-dimensional matrix multiplication operations, even in noisy wireless communication environments, by minimizing calculation errors and alleviating complexity compared to conventional studies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024096558_22052025_PF_FP_ABST
    Figure KR2024096558_22052025_PF_FP_ABST
Patent Text Reader

Abstract

A distributed matrices multiplication method, of a distributed computing system having a master node and a plurality of worker nodes, according to an embodiment of the present invention, includes: a step in which the master node generates a first partition function and second partition function based on group algebra, according to a partition order (pm) of a first input matrix and a partition order (pn) of a second input matrix, and broadcasts the partition functions through a network having a multi-access channel structure; a step in which at least pmn worker nodes each encode different first sub-matrices and second sub-matrices which are a part of the first input matrix and second input matrix, respectively, on the basis of the first partition function and second partition function received through the network, to generate first encoded matrices and second encoded matrices; a step in which each worker node performs multiplication of encoded matrices by using the first encoded matrix and second encoded matrix generated by the worker node; and a step in which the master node aggregates the results of the multiplication from each of the worker nodes to reconstruct the result of the multiplication of the first input matrix and second input matrix.
Need to check novelty before this filing date? Find Prior Art

Description

Distributed matrix multiplication operation method based on the network structure of a distributed computing system

[0001] The present invention relates to a distributed matrix multiplication operation method and a distributed computing system for performing the same, and more particularly, to a method and system for improving the accuracy of a distributed matrix multiplication operation due to the structural characteristics of a network.

[0002] Recently, services that involve programs or applications called high-intelligence, such as big data analysis, augmented reality, and tactile communication, are being introduced, and high-dimensional matrix operations are essential for their actual commercialization.

[0003] Because it takes a long time to perform such a large load of calculations on a single device, research is being conducted on distributed computing that distributes communication across multiple edge devices.

[0004] Among them, encoded computing has emerged to solve the problem of congestion in accuracy, speed, and load of operations, and it is a method that enables restoration of the original matrix multiplication result using the encoded matrix multiplication result.

[0005] However, although communication channels, noise, and network structures directly affect the computational accuracy of coded computing, existing coded computing methods do not take these factors into account, and the network structures also simply utilize P2P communication.

[0006] Specifically, due to the computing power or communication limitations of the edge device, there may be a significant delay in the time it takes for the parallel processing operation to be transmitted to the main device. Even if the operation results are all received from the remaining edge devices, if the results are not transmitted from one or more edge devices, the result of the entire matrix multiplication operation cannot be restored.

[0007] That is, in response to a situation in which results are inefficiently obtained due to straggler edge devices, techniques have been proposed that can restore the operation values ​​of the entire matrix product based on the encoding of the sub-matrix even if the operation results are not transmitted from at least some edge devices.

[0008] However, most existing studies have been conducted under the assumption that the communication link between the main device and the edge device is error-free, which has made it difficult to extend distributed computing techniques to wireless communication environments because it does not reflect the phenomenon of fading or noise in communication.

[0009] In this regard, prior art literature includes the paper Q. Yu, MA Maddah-Ali, and AS Avestimehr, "Straggler mitigation in distributed matrix multiplication: Fundamental limits and optimal coding," IEEE Transactions on Information Theory, vol. 66, no. 3, pp. 1920-1933, 2020.

[0010] The present invention is intended to solve the problems of the prior art described above, and aims to provide a system and method for performing distributed multiplication operations by utilizing structural characteristics in a wireless network environment based on a group algebra-based coding computing technique.

[0011] However, the technical tasks that this embodiment seeks to accomplish are not limited to the technical tasks described above, and other technical tasks may exist.

[0012] A distributed computing system for performing a distributed matrix multiplication operation method according to one embodiment of the present invention includes a master node for processing a multiplication operation of a first input matrix and a second input matrix, and a plurality of worker nodes configured to process in parallel a portion of the multiplication operation of the first input matrix and the second input matrix distributed by the master node and transmit it to the master node.

[0013] According to one embodiment of the present invention, the distributed matrix multiplication operation method comprises: a step in which the master node generates a first partitioning function based on a number of units according to a partitioning order (pm) of the first input matrix and a second partitioning function based on a number of units according to a partitioning order (pn) of the second input matrix, and broadcasts the first partitioning function through a network having a multi-access channel structure; a step in which at least pmn worker nodes generate a first encoding matrix and a second encoding matrix by encoding different first sub-matrices and second sub-matrices forming a part of the first input matrix and the second input matrix, respectively, based on the first partitioning function and the second partitioning function received through the network; a step in which each worker node performs an encoding matrix multiplication operation using the first encoding matrix and the second encoding matrix it has generated; And the master node includes a step of collecting the results of the operation from each worker node and restoring the results of the product operation for the first input matrix and the second input matrix.

[0014] According to one embodiment of the present invention, the speed, accuracy and load of high-dimensional matrix multiplication operations are improved.

[0015] According to one embodiment of the present invention, the complexity of high-dimensional matrix operations is reduced compared to conventional research.

[0016] According to one embodiment of the present invention, it can be applied even in a noisy wireless communication environment and calculation errors can be minimized.

[0017] The effects of the present invention are not limited to those mentioned above, and other effects not mentioned will be clearly understood by those skilled in the art from the description below.

[0018] FIG. 1 is a structural diagram of a distributed computing system according to one embodiment of the present invention.

[0019] FIG. 2 is a block diagram of a configuration of a master node according to one embodiment of the present invention.

[0020] Figure 3 is a conceptual diagram for explaining a distributed matrix multiplication operation method that is the basis of the present invention.

[0021] Figure 4 is a conceptual diagram for explaining a first embodiment of a distributed matrix multiplication operation method of the present invention.

[0022] Fig. 5 is a conceptual diagram for explaining a second embodiment of a distributed matrix multiplication operation method of the present invention.

[0023] FIG. 6 is a flowchart of a distributed matrix multiplication operation method according to one embodiment of the present invention.

[0024] Figures 7a and 7b are drawings for explaining effects derived from an embodiment of the present invention.

[0025] Below, with reference to the attached drawings, embodiments of the present invention are described in detail so that those skilled in the art can easily implement them. However, the present invention may be implemented in various different forms and is not limited to the embodiments described herein. In the drawings, irrelevant parts have been omitted for clarity of description, and similar reference numerals have been used throughout the specification to indicate similar elements.

[0026] Throughout the specification, when a part is said to be "connected" to another part, this includes not only the cases where the parts are "directly connected" but also the cases where the parts are "electrically connected" with other elements intervening. Furthermore, when a part is said to "include" a component, this does not exclude other components, but rather includes other components, unless otherwise stated.

[0027] In the following description, terms such as "first" and "second" may be used to describe various components; however, these components should not be limited by these terms, and these terms are used solely to distinguish one component from another. Furthermore, in the following description, singular expressions may include plural expressions, unless the context clearly indicates otherwise.

[0028] The "user terminal" mentioned below may be implemented as a computer or portable terminal that can access a server or other terminal via a network. Here, the computer may include, for example, a notebook, desktop, or laptop equipped with a web browser. The portable terminal may include, for example, a wireless communication device that ensures portability and mobility, such as a smart phone, tablet PC, wearable device, various VR HMD devices, as well as various devices equipped with communication modules such as Bluetooth (BLE, Bluetooth Low Energy), NFC, RFID, ultrasonic, infrared, WiFi, LiFi, etc. In addition, the "network" refers to a connection structure that enables information exchange between each node, such as terminals and servers, and includes a local area network (LAN), a wide area network (WAN), the Internet (WWW: World Wide Web), a wired and wireless data communication network, a telephone network, a wired and wireless television communication network, etc. Examples of wireless data communication networks include, but are not limited to, 3G, 4G, 5G, 3GPP (3rd Generation Partnership Project), LTE (Long Term Evolution), WIMAX (World Interoperability for Microwave Access), Wi-Fi, Bluetooth, infrared, ultrasonic, visible light communication (VLC), and LiFi.

[0029] Hereinafter, an embodiment of the present invention will be described in detail using the attached drawings.

[0030] FIG. 1 is a structural diagram of a distributed computing system according to one embodiment of the present invention. Referring to FIG. 1, the distributed computing system (10) includes a master node (100) and a plurality of worker nodes (200) connected to the master node (100) via a network.

[0031] A master node (100) according to one embodiment refers to a main device or server within a distributed computing system (10), and can process high-intelligence tasks requested from a user terminal (not shown) and provide the results thereof. Here, the high-intelligence tasks are, for example, big data analysis, tactile communication, augmented reality, mixed reality, virtual reality, metaverse, artificial intelligence models, etc. implemented as programs or applications, and are not necessarily limited to the examples, and any task requiring high-dimensional matrix operations should be interpreted as being included in the present invention.

[0032] For example, the master node (100) can process a product operation of the first input matrix and the second input matrix to derive a desired result in a task requested by a user terminal. This is merely an exemplary expression for explanation, and the number of matrices that are actually the target of the product operation is not limited.

[0033] FIG. 2 is a block diagram of the configuration of a master node (100) according to one embodiment of the present invention. Referring to FIG. 2, the master node (100) may include a processor (110), a memory (120), and a communication module (130).

[0034] The memory (120) can store commands for the master node (100) to perform a product operation of the first input matrix and the second input matrix. The memory (120) can store an executable program that generates and executes one or more commands that implement the operation.

[0035] The memory (120) may include built-in memory and / or external memory, and may include volatile memory such as DRAM, SRAM, or SDRAM, nonvolatile memory such as OTPROM (one time programmable ROM), PROM, EPROM, EEPROM, mask ROM, flash ROM, NAND flash memory, or NOR flash memory, flash drive such as SSD, CF (compact flash) card, SD card, Micro-SD card, Mini-SD card, Xd card, or memory stick, or storage device such as HDD. The memory (120) may include, but is not limited to, magnetic storage media or flash storage media.

[0036] The processor (110) can execute its own process included in the distributed matrix multiplication operation method according to the following embodiment based on the program and instructions stored in the memory (120). The processor (110) can include any type of device capable of processing operations on data.

[0037] A processor (110) may refer to a data processing device built into hardware, for example, having a physically structured circuit to perform a function expressed by a code or command included in a program. Examples of such a data processing device built into hardware include, but are not limited to, processing devices such as a microprocessor, a central processing unit (CPU), a processor core, a multiprocessor, an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), and a graphics processing unit (GPU).

[0038] The communication module (130) may provide a communication interface that provides packet data as transmission and reception signals to and from external devices, including user terminals and worker nodes (200), using wireless communication technology. In addition, the communication module (130) may be a device including the hardware and software necessary to transmit and receive control signals or data signals, etc., with other network devices via wired or wireless connections.

[0039] According to one embodiment, a plurality of worker nodes (200) are included in a distributed computing system (10), and an environment in which a total of W worker nodes (WN1) to W worker nodes (WNW) are configured is exemplified below.

[0040] The worker node (200) refers to a type of edge device or server that processes in parallel a portion of the product operation of the first input matrix and the second input matrix distributed by the master node (100) for the distributed matrix product operation of the first input matrix and the second input matrix and transmits the result to the master node (100).

[0041] For example, the master node (100) may receive a task processing request signal from a user terminal. The master node (100) may request distributed processing of the task received from the user terminal to at least one worker node (200). Each worker node (200) may process different parts of the product operation of the first input matrix and the second input matrix in parallel and transmit the result to the master node (100) via wireless communication, and the master node (100) may collect the distributed processing results to generate a task processing result signal and transmit the result to the user terminal.

[0042] Hereinafter, the distributed matrix multiplication operation process according to the present invention will be specifically described using FIGS. 3 to 5.

[0043] First, let's define the terms used in the formula. and represent real and complex numbers, respectively. is defined as a finite field with size N. is defined as the set of integers that are the modulo N operations of the integer N. is defined as the remainder of R divided by 1. am.<V, W> is defined as the inner product of vector v and vector w. is defined as the expectation of the random variable X. |S| is defined as the size (cardinality) of the set of S. (g) X is defined as the multiplicative group (cyclic group) of a finite field g. (M) ζη represents the ζ-th row and η-th column of matrix M. (M) + denotes the Moore-Penrose inverse of matrix M. (M) H means the conjugate transpose of matrix M. ∥M∥ F means the Frobeniusnorm of matrix M. M N and M N represents the Kronecker product and the Hadamard product of vectors M and N, respectively.

[0044] In addition, the input matrix and its product operation performed according to the user's request are defined as follows. However, this is an example for explanation, and the scope of the present invention is not limited thereto.

[0045] In this example, two matrices and Consider a distributed computing system (10) for distributed multiplication operations. The master node (100) has a first input matrix (A T ) and the product of the second input matrix (B) , and to this end, partial operations are requested in parallel to at least some of the first worker node (WN1) to the Wth worker node (WNW).

[0046] At this time, the first input matrix (A T ) and the second input matrix (B) can be expressed by dividing into a first submatrix according to the first partition order (pm) and a second submatrix according to the second partition order (pn), as in Equation 1 below.

[0047] [Formula 1]

[0048]

[0049] At least pmn worker nodes (200) among W have a first encoding matrix ( ) and the second encoding matrix ( ) can be used to perform encoding matrix multiplication operations for each part. (w is in the range of 1 to W)

[0050] Based on the above assumptions, we will first explain an embodiment that serves as the basis for a distributed matrix multiplication operation method based on encoding using Fig. 3. Fig. 3 is a conceptual diagram for explaining a distributed matrix multiplication operation method that serves as the basis for the present invention.

[0051] The master node (100) encodes the first sub-matrix and the second sub-matrix as in Equation 2 below for distribution to the w-th worker node (Wnw) to create the first encoding matrix ( ) and the second encoding matrix ( ) is created.

[0052] [Formula 2]

[0053]

[0054] Here, and are each an encoding parameter, and are set differently for each worker node (200) by the master node (100) so that the original result C can be restored.

[0055] The master node (100) sets the preset transmission power P based on Equation 3 below. A and P B Wow, transmission time According to , the first encoding matrix ( ) and the second encoding matrix ( ) respectively and quantizes the vector.

[0056] [Formula 3]

[0057]

[0058] The master node (100) is During X w,A and X w,B is transmitted to the wth worker node (Wnw). Accordingly, the signal received by the worker node (Wnw) is set as in Equation 4 below.

[0059] [Formula 4]

[0060]

[0061] Here, and are respectively, X w,A and X w,Bis the channel coefficient of the wireless network from the master node (100) to the wth worker node (Wnw). In addition, and n w,B respectively represent the channel noise of the network. Meanwhile, represents a complex Gaussian noise process, whose components are complex Gaussian random variables with mean 0 and variance 1.

[0062] After receiving the encoded signal as above, the wth worker node (Wnw) receives Y w,A and Y w,B Each of them and By decoding, the encoding matrix multiplication operation is performed as in Equation 5 below.

[0063] [Formula 5]

[0064]

[0065] After that, the wth worker node (Wnw) Transmission time During transmission power P w According to quantizes the vector. Accordingly, the signal received by the master node (100) from the worker node (Wnw) is given as in Equation 6 below.

[0066] [Formula 6]

[0067]

[0068] Here, is Y w,C is the channel coefficient of the network from the wth worker node (Wnw) to the master node (100), and n w,C represents channel noise.

[0069] Master node (100) is Y w,C The first input matrix (A) T ) and the product operation of the second input matrix (B) Restore to C. That is, C and can be expressed as having mn submatrices, and C ij and can be defined as the (i, j)th partitioned submatrix.

[0070] As described above, the embodiment of the present invention as illustrated in FIG. 3 is a method in which the master node (100) individually communicates with each worker node (200) to transmit an encoding matrix for distribution of operations, and each worker node (200) also communicates in parallel with the master node (100) to transmit the result of the encoding matrix multiplication operation. In the present invention, this is defined as OMA-OMA (orthogonal multiple access-orthogonal multiple access).

[0071] Hereinafter, the first and second embodiments of the present invention, which are designed based on OMA-OMA, will be described. Fig. 4 is a conceptual diagram illustrating the first embodiment of the distributed matrix multiplication operation method of the present invention. Fig. 5 is a conceptual diagram illustrating the second embodiment of the distributed matrix multiplication operation method of the present invention.

[0072] First, the present invention proposes a method that is more advanced than OMA, taking advantage of the structural characteristics of wireless network environments. Specifically, the present invention proposes a transmission technique that minimizes the probability of computational errors by leveraging the characteristics of broadcasting (BC) and multiple access channels (MAC). This technique is defined herein as Broadcast-and-Compute (BnC).

[0073] The development of BnC relies on two established theories. The first theory, based on group algebra, explains matrix multiplication within a group theory framework by mapping each submatrix of the main matrix to elements of various types of groups. (H. Cohn and C. Umans, "A group-theoretic approach to fast matrix multiplication," in Proc. 44th Annu. IEEE Symp. Found. Comput. Sci., Cambridge, MA, USA, 2003, pp. 438-449.)

[0074] Another second theory, which seeks to restore linear combinations of given messages using MAC structures in a distributed computing environment, proposes a lattice-based encoder and decoder to restore messages at improved transmission rates. This framework is defined as Compute-and-Forward (CF). ("Compute-and-Forward: Harnessing Interference through Structured Codes," IEEE Trans. Inf. Theory, vol. 57, no. 10, pp. 6463-6486, Oct. 2011.)

[0075] In addition to the second theory, if w = 1, · · · ,W is indexed, k is transmitted to the dth relay for each transmitter. w When there is a message vector of length , which is a finite field is defined independently and uniformly. The goal of the rth relay is to stably restore the linear combination of messages, which can be expressed by Equation 7 below.

[0076] [Formula 7]

[0077]

[0078] The second theory introduces CF to restore the equation of linear combination at a desired transmission speed, and one embodiment of the present invention derives the following algorithm based on this.

[0079] Based on the second theory, all channel vectors and coefficient vector There may exist a nested lattice code Λ ⊂ Λ1 ⊂ · · · ⊂ΛW with rates r1, · · · , rW . In this case, the rth relay is given by Equation 8 below ( For some parts of the grid, the grid equation (grid point t) is as follows: Equation 9 w ∈ L w ) is the average error probability can be decoded as

[0080] [Formula 8]

[0081]

[0082] [Formula 9]

[0083]

[0084] Furthermore, based on the second theory, α that minimizes the upper noise bound of Equation 8 r can be calculated using the following formula 10.

[0085] [Formula 10]

[0086]

[0087] As a result, the operation speed can be derived from Equation 11 below.

[0088] [Formula 11]

[0089]

[0090] Based on these two theories, BnC requests each worker node (200) to perform an encoding matrix multiplication operation by broadcasting the same information over a wireless network. Furthermore, a method may be additionally introduced in which the results of the encoding matrix multiplication operation performed by each worker node (200) are linearly combined based on a MAC and collectively transmitted to the master node (100).

[0091] Referring to FIG. 4, the first embodiment of the present invention is a method of applying BnC when the master node (100) requests distribution of matrix multiplication operations to each worker node, which can be defined as BnC-OMA. Briefly, unlike OMA, where the master node (100) encodes each sub-matrix and sends it to the worker node (200), the first embodiment is a technique whereby when the master node (100) broadcasts the same information based on the number of troops, each worker node (200) utilizes this to encode each sub-matrix differently.

[0092] Specifically, the first embodiment proceeds as follows. First, the master node (100) inputs the first input matrix (A) based on the number of troops. T ) according to the partition order (pm) of the first partition function (A) and the second partition function (B) according to the partition order (pn) of the second input matrix (B) are generated as in Equation 12 below.

[0093] [Formula 12]

[0094]

[0095] That is, the first partition function (A) and the second partition function (B) are the first input matrix (A T ) and the first encoding matrix (B) corresponding to the different first sub-matrix and second sub-matrix of the second input matrix (B), each worker node (200) is responsible for the encoding matrix multiplication operation. ) and the second encoding matrix ( ) can be understood as a commonly applied frame for generating.

[0096] A more specific example of generating the first partition function (A) and the second partition function (B) is as follows. For the sake of the core explanation, group theory and representation theory, which are the basis of the first theory, are already well known and are therefore omitted. Instead, the triple product property applied to encoding matrix multiplication as a group number is introduced.

[0097] The triple product rule states that for a given group g, the right quotient set of a subset s ⊂g is denoted by Q(S). (i.e., Q(S) = {s1s2 -1 If for three subsets S, T, U ⊂g, each element g1∈ Q(S), g2∈ Q(T), g3∈ Q(U) obeys Equation 13 below, then S, T, and U can be seen to satisfy the triple product rule.

[0098] [Formula 13]

[0099] g1g2g3= 1, ifg1=g2=g3= 1

[0100] Here, the method of coding by partitioning the matrix product based on the military number considers S, T, U ⊂g that satisfy the triple product rule. In addition, two complex matrices, A with the size of |S| × |T| and B with the size of |T| × |U|, are considered. Accordingly, the military number-based encoding for A and B can be defined as in Equation 14 below.

[0101] [Formula 14]

[0102]

[0103] According to formula 14, Group element s -1From the coefficients of u, we can obtain the result of the operation of the AB matrix multiplication. Intuitively, this can be viewed as a generalization of the concept of cyclic convolution of the DFT (discrete Fourier transform) to an arbitrary group.

[0104] Therefore, based on the aforementioned military number concept, when p, m, and n corresponding to the partition order are coprime, distribution of matrix multiplication operation using cyclic group can be introduced. That is, the master node (100) can introduce the first input matrix (A T ) and the partitioning order of the second input matrix (B), the cyclic group g = {1, g, · · ·, g} in which the elements of the group are linearly combined and broadcasted in correspondence to each submatrix apmn-1} 2 , where α is an arbitrary integer, which is relatively very small compared to p, m, and n according to the well-known military number theory.

[0105] Meanwhile, the cyclic group g is considered to have three subsets as shown in Equation 15 below, and these satisfy the triple product rule.

[0106] [Formula 15]

[0107]

[0108] Therefore, based on the military number theory, the master node (100) is based on the cyclic group g and the first input matrix (A T ) and the second input matrix (B) can generate the first partition function (A) and the second partition function (B) by mapping each of the partitioned submatrices to an arbitrary matrix structure. That is, the master node (100) embeds the partitioned submatrices by utilizing the element g of g, which can be expressed by the above equation 12.

[0109] In this way, A and B are expressed as frames based on embedded military numbers, but since the encoding matrices processed by each worker node (200) are derived from the same military numbers A and B, A and B can be understood as a type of division function. Accordingly, the master node (100) can distribute the matrix multiplication operation to each worker node (200) by generating a first division function (A) and a second division function (B) and broadcasting them through a MAC structured network.

[0110] OMA is the first input matrix (A) at the master node (100) level. T ) and each submatrix of the second input matrix (B) is encoded to create the first encoding matrix ( ) and the second encoding matrix ( ) and distributes it to each worker node (200) through individual communication, or BnC transmits the first partition function (A) and the second partition function (B) in batches, and encodes each sub-matrix at the worker node (200) level to create the first encoding matrix ( ) and the second encoding matrix ( ) has a fundamental difference in generating the transmission time T. Accordingly, the difference in transmission method between BnC and OMA is the preset transmission time T. local In , the master node (100) is the first encoding matrix ( ) and the second encoding matrix ( ) can be interpreted as a quantization unit (Quantization Step Size) for transmitting the first partition function (A) and the second partition function (B).

[0111] Specifically, in BnC, the first partition function (A) and the second partition function (B) are military numbers and Since it is an element of , the first partition function (A) and the second partition function (B) are each and It can be expressed as a linear combination of elements (g) of a group cyclic group (g) with coefficients in the form of a complex matrix.

[0112] Based on this, the master node (100) transmits the total transmission time T of the first partition function (A) and the second partition function (B). loca The matrix multiplication operation can be distributed to each worker node (200) by dividing it into pmn sections with equal intervals and transmitting different coefficients for each section. That is, the master node (100) can distribute the matrix multiplication operation to the i-th section by transmitting the i-th group element of the first partition function (A) g i The coefficients can be transmitted over the network.

[0113] Here, since the coefficients of each group are complex matrices with dimensions α / m × β / p, the master node (100) is for the first partition function (A). It can be quantized into bits. The second partition function (B) is also quantized in the same way.

[0114] Therefore, the quantization unit for the master node (100) to broadcast the first division function (A) and the second division function (B) can be given by Equation 16 below.

[0115] [Formula 16]

[0116]

[0117] That is, the quantization step size for transmitting each element of the cyclic group in BnC is the transmission time and channel capacity of the first partition function (A) and the second partition function (B), and the first input matrix (A T ) and can be determined based on the partitioning order and dimension of the second input matrix (B).

[0118] When broadcasting by the master node (100) is in progress, at least pmn (W>pmn) worker nodes (200) generate a first input matrix (A) based on the first partition function (A) and the second partition function (B) received through the network. T ) and the different first sub-matrix and second sub-matrix forming part of the second input matrix (B) are respectively referred to as the first encoding matrix ( ) and the second encoding matrix ( ) is created.

[0119] As an example, in BnC, each worker node (200) generates a first encoding matrix ( ) and the second encoding matrix ( ) is derived by the first partition function (A) and the second partition function (B) according to Equation 12, and Equation 2 in OMA can be equivalently modeled as Equation 17 below.

[0120] [Formula 17]

[0121]

[0122] To explain this, each worker node (200) has a first encoding matrix ( ) and the second encoding matrix ( ) must perform matrix multiplication operation, so at this time, the coefficients of the cyclic group element g corresponding to the first submatrix and the second submatrix in the first partition function (A) and the second partition function (B) are complex values ​​( ) can be converted into. Based on the first theory mentioned above, the cyclic group (g = {1, g, ..., g apmn-1} 2 ) has a total of apmn irreducible representations as functions, can be expressed as in Equation 18 below.

[0123] [Formula 18]

[0124]

[0125] Here, am.

[0126] Also, in the above formula 18 According to, is established, is a cyclic group of any integer ( ) is a set of remainders divided by apmn. That is, as in Equation 19 below, the remainder of dividing the degree applied to the coefficient converted to a complex value by pmn is different for each worker node (200), and accordingly, each worker node (200) has a different first encoding matrix ( ) and the second encoding matrix ( ) can be derived.

[0127] [Formula 19]

[0128]

[0129] The first encoding matrix ( ) and the second encoding matrix ( ) is generated, each worker node (200) uses it to perform an encoding matrix multiplication operation.

[0130] As an example, the transmission signal in BnC can be equivalently modeled as Equation 20 below, with respect to Equations 4 and 6 described in OMA, by eliminating the effects of the encoder and decoder on the network.

[0131] [Formula 20]

[0132]

[0133] Here, N′ w,a (a = A, B) is [ ] is quantization noise that is uniformly distributed over the interval.

[0134] Accordingly, each worker node (200) receives the result of the operation , performs quantization, and transmits it to the master node (100).

[0135] As an example, each worker node (200) can quantize the results of the encoding matrix multiplication operations it performs for each sequentially set transmission section and transmit them to the master node (100). Here, the transmission section can mean the total transmission time taken to transmit the results of all encoding matrix multiplication operations divided into pmn equal intervals.

[0136] Specifically, the wth worker node (Wnw) in the wth transmission interval With an ergodic channel capacity of up to It can transmit bits. Also, Since the size of is α / m × γ / n, each element is can be quantized into . Therefore, The quantization step size is can be expressed as

[0137] And, after removing the influence of the encoder and decoder of the network channel, the received signal of the master node (100) can be expressed as . Here, N w,c Is [ ] means quantization noise that follows a uniform distribution in the interval (a=C).

[0138] Afterwards, the master node (100) collects the results of the encoding matrix multiplication operation from each worker node (200) and creates the first input matrix (A T ) and restore the product operation of the second input matrix (B).

[0139] As a specific example, the master node (100) receives the When you collect, you get what you originally wanted. can be restored in the following way. First, the result of the encoding matrix multiplication operation ( ) of the worker node (200) that delivered the index set. If defined as , the result of the encoding matrix multiplication operation ( ) can be expressed by the following equation 21, which connects the ζ-th row and the η-th column.

[0140] [Formula 21]

[0141]

[0142] Additionally, the heat vector can be defined as in Equation 22 below.

[0143] [Formula 22]

[0144]

[0145] Is Since it can be expressed as a linear combination of , also can be expressed as a linear combination of , and accordingly Z ζη and D ζη The linear equation of can be expressed as Equation 23 below.

[0146] [Formula 23]

[0147]

[0148] Here, P is a value preset to the master node (100), so the master node (100) By restoring, can be derived.

[0149] In the first embodiment described above, as shown in FIG. 4, BnC is applied in the process of distributing matrix multiplication from the master node (100) to the worker node (200), and the process of transmitting the result of the encoded matrix multiplication operation from the worker node (200) to the master node (100) follows OMA, and is defined as BnC-OMA in this specification.

[0150] Next, referring to FIG. 5, the second embodiment of the present invention is a method in which the master node (100) utilizes Compute-and-Forward (CF) according to the second theory in the process of collecting the encoded matrix multiplication operations of each worker node (200). Accordingly, in this specification, the second embodiment is defined as BnC-CF, and the process up to the front end where each worker node (200) performs the encoded matrix multiplication operation overlaps with the first embodiment, and will be replaced with the above.

[0151] According to the second embodiment, the result of each encoding matrix multiplication operation performed by each worker node (200) ) are linearly combined during transmission over the network according to the characteristics of the MAC structure. ) can be reached to the master node (100).

[0152] The difference between OMA and CF is the result of the distributed matrix multiplication operation performed by the wth worker node (Wnw). ) can be viewed as an encoding and decoding process. That is, by utilizing the grid-based CF encoder and decoder introduced in the second theory. A new speed for transmitting to the master node (100) can be derived.

[0153] To this end, a function for the encoding matrix multiplication operation of the worker node (200) is identified, and the result of the operation ( ) must be calculated to obtain the computational speed that can be transmitted based on CF.

[0154] Specifically, the first encoding matrix ( ) and the second encoding matrix ( ), according to Equation 17 which defines the first encoding matrix ( ) and the second encoding matrix ( ) The product can be derived using the following equation 24.

[0155] [Formula 24]

[0156]

[0157] According to Equation 24, if we calculate w=1, From the linear combination of the operation results, the (i, j)th submatrix of the matrix (C) to be obtained can be expressed as Equation 25 below.

[0158] [Formula 25]

[0159]

[0160] Here, i =1, · · ·, m, and j =1, · · ·, n.

[0161] Accordingly, based on Equation 7 of CF, the function is set as Equation 26 below, and in each worker node (200) When transmitted as , the master node (100) transmits the (i, j)th submatrix (C) of the matrix (C) through Equation 25. ij ) can be restored.

[0162] [Formula 26]

[0163]

[0164] In particular, according to Equation 11 based on the second theory, Equation 27 below can be derived, and the master node (100) can decode Equation 26 at an operation speed that satisfies it, thereby restoring each submatrix of the matrix (C) in a direction that minimizes the upper limit of noise.

[0165] [Formula 27]

[0166]

[0167] That is, the result of each encoding matrix multiplication operation ( ) is the first input matrix (A T) and the second input matrix (B) can be viewed as being linearly combined into an equation (Equation 26) in which lattice-based complex values ​​according to the partitioning order are applied as coefficients. Accordingly, the master node (100) can decode the equation (Equation 26) at an operation speed that satisfies a preset equation (Equation 27) based on noise on the network, and restore the result of the product operation for each of the first submatrix and the second submatrix.

[0168] Meanwhile, the result of the encoding matrix multiplication operation of each worker node (200) ) In transmission, the CF decoder transmits a without any distortion. ij It is assumed that the field size of CF is sufficiently large to obtain . Based on this, each worker node (200) can obtain a transmission speed (T master ) and operation speed (R ij ) as the result of each encoding matrix multiplication operation ( ) can be quantized and transmitted.

[0169] That is, each worker node (200) has a grid equation ( ) each element of R ij It can be quantized uniformly into bits, and the corresponding quantization unit (Quantization Step Size) can be expressed as Equation 28 below.

[0170] [Formula 28]

[0171]

[0172] Therefore, in the first embodiment, the reception signal according to OMA ( ) can be equivalently modeled as in Equation 29 below.

[0173] [Formula 29]

[0174]

[0175] Here, N ij is the section [ ] is a quantization noise with a uniform distribution over the whole area. ( )

[0176] Regarding the first and second embodiments of the present invention described above, effects supported by FIG. 7 are presented. FIG. 7a and FIG. 7b are drawings for explaining effects derived from the embodiments of the present invention.

[0177] Specifically, Figs. 7a and 7b are experimental data showing the average computation error according to the signal-to-noise ratio (SNR). p=6, m=n=1 were set, and Fig. 7a shows a graph showing the results of an experiment with 12 worker nodes (200), and Fig. 7b shows a graph showing the results of an experiment with 90 worker nodes. Here, UN stands for uncoded scheme, which means that encoding for distributed matrix multiplication operation is not performed. GA means that encoding is performed based on the first theory.

[0178] First, referring to FIGS. 7a and 7b, it can be confirmed that the GA, which performs distributed matrix multiplication based on the first theory (number of groups), shows better performance than the UN, which does not. In addition, it can be confirmed that the performance is superior when transmitting using BnC (first embodiment) compared to when using the existing OMA. In addition, as shown in FIG. 7b, when the number of worker nodes (200) is large, the utilization of CF (second embodiment) is more effective than when using the existing OMA, and thus it can be confirmed that it is effective for introduction to high-intelligence tasks.

[0179] Hereinafter, the main contents of the present invention will be summarized using Fig. 6. Fig. 6 is an operational flowchart of a distributed matrix multiplication operation method according to one embodiment of the present invention, and overlapping contents are replaced with the above.

[0180] First, the master node (100) inputs the first input matrix (A T ) and the product of the second input matrix (B) , and to this end, partial operations are requested in parallel to at least some of the first worker node (WN1) to the Wth worker node (WNW).

[0181] In step S610, the master node (100) inputs the first input matrix (A T ) and the first partition function (A) and the second partition function (B) based on the number of military units according to the partition order (pm) of the second input matrix (B) and the partition order (pn).

[0182] As an example, the master node (100) has a first input matrix (A T ) and the second input matrix (B) are divided into a cyclic group (g) composed of elements (g) corresponding to each of the divided submatrices, and based on this, the first partition function (A) and the second partition function (B) can be generated by linearly combining the submatrices by mapping them to an arbitrary matrix structure.

[0183] In step S620, the master node (100) broadcasts the first partition function (A) and the second partition function (B) through a network with a multi-access channel structure.

[0184] As an example, the master node (100) divides the preset transmission time of the first partition function (A) and the second partition function (B) into pmn sections with equal intervals, and transmits coefficients set differently corresponding to the first sub-matrix and the second sub-matrix for each of the divided sections.

[0185] As an example, the master node (100) can quantize the first partition function (A) and the second partition function (B) with a preset quantization unit (Quantization Step Size). Here, the quantization unit ( ) is the transmission time and channel capacity of the first partition function (A) and the second partition function (B), and the first input matrix (A T ) and can be set based on the partitioning order and dimension of the second input matrix (B).

[0186] In step S630, at least pmn worker nodes (200) generate a first input matrix (A) based on the first partition function (A) and the second partition function (B) received through the network. T ) and the different first sub-matrix and second sub-matrix forming part of the second input matrix (B) are respectively referred to as the first encoding matrix ( ) and the second encoding matrix ( ) is created.

[0187] As an example, each worker node (200) performs a complex value (coefficient of the first sub-matrix and the second sub-matrix, which are each responsible for the distributed multiplication operation in the first partition function (A) and the second partition function (B). ) can be converted into. Here, the remainder obtained by dividing the degree applied to the complex value by the above pmn is different for each worker node (200).

[0188] In step S640, each worker node (200) generates its own first encoding matrix ( ) and the second encoding matrix ( ) using the encoded matrix multiplication operation ( ) is performed.

[0189] As an optional first embodiment, each worker node (200) sequentially performs the result of the encoding matrix multiplication operation performed for each transmission section set ) can be quantized and transmitted to the master node (100). Here, the transmission section may be the total transmission time taken to transmit the results of all encoding matrix multiplication operations divided into pmn equal intervals.

[0190] As an optional second embodiment, the result of each encoding matrix multiplication operation performed by each worker node (200) ) are linearly combined while being transmitted over the network. ) can be reached to the master node (100). For example, the result of each encoding matrix multiplication operation ( ) is the first input matrix (AT ) and the equation in which the grid-based complex values ​​according to the partitioning order of the second input matrix (B) are applied as coefficients ( ) may be linearly combined.

[0191] In step S650, the master node (100) receives the result of the operation from each worker node (200). ) are combined to form the first input matrix (A T ) and restore the result (C) of the product operation for the second input matrix (B).

[0192] As an example extending from the second embodiment above, the master node (100) has an operation speed (R) that satisfies a preset formula based on noise on the network. ij ) as a lattice equation ( ) is decoded and the result of the product operation for each of the first and second sub-matrices (C ij , i = 1, · · ·, m, j = 1, · · ·, n) can be restored.

[0193] An embodiment of the present invention may also be implemented in the form of a non-transitory storage medium containing computer-executable instructions, such as a program module executed by a computer. Computer-readable media may be any available media that can be accessed by a computer, and includes both volatile and nonvolatile media, removable and non-removable media. Furthermore, computer-readable media may include all computer storage media. Computer storage media includes both volatile and nonvolatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, or other data.

[0194] The foregoing description of the present invention is for illustrative purposes only, and those skilled in the art will readily appreciate that the present invention can be readily modified into other specific forms without altering the technical spirit or essential characteristics of the present invention. Therefore, the embodiments described above should be understood as illustrative in all respects and not restrictive. For example, each component described as a single entity may be implemented in a distributed manner, and similarly, components described as distributed may be implemented in a combined manner.

[0195] The scope of the present invention is indicated by the claims described below rather than the detailed description above, and all changes or modifications derived from the meaning and scope of the claims and their equivalent concepts should be interpreted as being included in the scope of the present invention.

Claims

1. A method for performing distributed matrix multiplication in a distributed computing system comprising a master node and multiple worker nodes, a) A step in which a master node generates a first partition function and a second partition function based on a military number according to a partition order (pm) of a first input matrix and a partition order (pn) of a second input matrix, and broadcasts the same through a network having a multi-access channel structure; b) a step of generating a first encoding matrix and a second encoding matrix by encoding different first sub-matrices and second sub-matrices forming part of the first input matrix and the second input matrix, respectively, based on the first partitioning function and the second partitioning function received through the network by at least pmn worker nodes; c) a step in which each worker node performs an encoding matrix multiplication operation using the first encoding matrix and the second encoding matrix generated by each worker node; and d) A distributed matrix multiplication operation method, comprising a step of the master node collecting the results of the operation from each worker node and restoring the results of the multiplication operation for the first input matrix and the second input matrix.

2. In paragraph 1, Step a) above, A distributed matrix multiplication operation method, comprising a step of generating the first partition function and the second partition function by linearly combining the master node, wherein the master node constructs a cyclic group composed of elements corresponding to each sub-matrix into which the first input matrix and the second input matrix are partitioned, and maps each sub-matrix to an arbitrary matrix structure based on the cyclic group.

3. In paragraph 1, Step a) above, A distributed matrix multiplication operation method, comprising the step of the master node dividing the preset transmission time of the first division function and the second division function into pmn sections having equal intervals, and transmitting coefficients set differently corresponding to the first sub-matrix and the second sub-matrix for each of the divided sections.

4. In paragraph 1, Step a) above, The above master node comprises a step of quantizing and broadcasting the first division function and the second division function using a preset quantization step size, A distributed matrix multiplication operation method, wherein the quantization unit is set based on the transmission time and channel capacity of the first division function and the second division function, and the division order and dimension of the first input matrix and the second input matrix.

5. In paragraph 1, Step b) above, Each of the above worker nodes includes a step of converting coefficients of the first sub-matrix and the second sub-matrix, which are each responsible for the distributed multiplication operation in the first partition function and the second partition function, into complex values. A distributed matrix multiplication operation method, wherein the remainder of dividing the degree applied to the above complex value by the above pmn is different for each worker node.

6. In paragraph 1, Step c) above, Including a step of each worker node sequentially quantizing the result of the encoding matrix multiplication operation performed by each worker node for each set transmission section and transmitting it to the master node. The above transmission section is a distributed matrix multiplication operation method in which the total transmission time taken to transmit the results of all encoded matrix multiplication operations is divided into equal intervals by the pmn number.

7. In paragraph 1, Step c) above, A distributed matrix multiplication operation method, comprising a step of linearly combining the results of each encoded matrix multiplication operation performed by each worker node while transmitting them through the network and reaching the master node.

8. In paragraph 7, Step d) above, The result of each of the above encoding matrix multiplication operations is linearly combined into an equation in which lattice-based complex values ​​according to the partitioning order of the first input matrix and the second input matrix are applied as coefficients, A distributed matrix multiplication operation method, comprising a step of the master node decoding the equation at a computational speed satisfying a preset formula based on noise on the network and restoring the result of the multiplication operation for each of the first sub-matrix and the second sub-matrix.

9. In a distributed computing system that performs a distributed matrix multiplication operation method, A master node that processes the product operation of the first input matrix and the second input matrix; and A plurality of worker nodes configured to process in parallel a portion of the product operation of the first input matrix and the second input matrix distributed by the master node and transmit it to the master node, The process of performing the above distributed matrix multiplication operation method is as follows: The above master node generates a first partition function and a second partition function based on the number of troops according to the partition order (pm) of the first input matrix and the partition order (pn) of the second input matrix, and broadcasts them through a network having a multi-access channel structure. At least pmn worker nodes encode different first sub-matrices and second sub-matrices forming part of the first input matrix and the second input matrix, respectively, based on the first partitioning function and the second partitioning function received through the network to generate a first encoding matrix and a second encoding matrix, Each worker node performs a multiplication operation of the encoding matrix using the first encoding matrix and the second encoding matrix that it has generated. A distributed computing system, wherein the master node collects the results of the operations from each worker node and restores the results of the product operation for the first input matrix and the second input matrix.

10. In paragraph 9, The above master node is, A distributed computing system, wherein a cyclic group is constructed by elements corresponding to each sub-matrix into which the first input matrix and the second input matrix are divided, and the first partitioning function and the second partitioning function are generated by linearly combining the sub-matrices by mapping them to an arbitrary matrix structure based on the cyclic group.

11. In paragraph 9, The above master node is, A distributed computing system, wherein the preset transmission time of the first partition function and the second partition function is divided into pmn sections having equal intervals, and coefficients set differently corresponding to the first sub-matrix and the second sub-matrix are transmitted for each of the divided sections.

12. In paragraph 9, The above master node is, The first partition function and the second partition function are quantized and broadcast using preset quantization units. A distributed computing system, wherein the quantization unit is set based on the transmission time and channel capacity of the first partitioning function and the second partitioning function, and the partitioning order and dimension of the first input matrix and the second input matrix.

13. In paragraph 9, Each of the above worker nodes, In the above first partition function and the above second partition function, the coefficients of the first submatrix and the second submatrix, which are each responsible for the variance product operation, are converted into complex values. A distributed computing system, wherein the remainder of dividing the degree applied to the complex value by the pmn is different for each worker node.

14. In paragraph 9, Each of the above worker nodes, The result of the encoding matrix multiplication operation performed for each sequentially set transmission section is quantized and transmitted to the master node. The above transmission section is a distributed computing system in which the total transmission time taken to transmit the results of all encoded matrix multiplication operations is divided into equal intervals by the pmn number.

15. In paragraph 9, A distributed computing system, wherein the results of each encoded matrix multiplication operation performed by each worker node are linearly combined while being transmitted through the network and reach the master node.

16. In paragraph 15, The result of each of the above encoding matrix multiplication operations is linearly combined into an equation in which grid-based complex values ​​according to the partitioning order of the first input matrix and the second input matrix are applied as coefficients, A distributed computing system, wherein the master node decrypts the equation at a computational speed that satisfies a preset formula based on noise on the network, and restores the result of the product operation for each of the first sub-matrix and the second sub-matrix.

Citation Information

Patent Citations

  • Faucet with blow dryer

    KR1020240002904A

  • Instrument-based distributed computing systems

    US20110167425A1

  • Processing device and related products

    US20200057652A1

  • A method of performing a matrix operation in a distributed processing system

    WO2015004421A1