Computational network and methods for use in the computational network
The computational network efficiently determines the symmetric properties of large distributed matrices by reducing communication overhead through a distributed antisymmetric function application, thereby improving the scalability and performance of HPC and AI applications.
Patent Information
- Application Number
- PCT/EP2023/086927
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-20
- Publication Date
- 2025-06-26
AI Technical Summary
Existing methods for determining the symmetric properties of large distributed matrices are inefficient due to high communication overhead, which prevents effective utilization of symmetric matrix solvers in high-performance computing (HPC) and artificial intelligence (AI) applications.
A computational network architecture that distributes a matrix across multiple nodes, where an orchestrator controller applies an antisymmetric function to each node's partial matrix and summarizes the results modulo a prime constant, reducing communication volume and enabling efficient determination of matrix symmetry.
This approach significantly reduces communication overhead, allowing for scalable and efficient identification of symmetric properties in large matrices, thereby enhancing the performance of HPC and AI applications by minimizing communication among nodes.
Smart Images

Figure EP2023086927_26062025_PF_FP_ABST
Abstract
Description
[0001]COMPUTATIONAL NETWORK AND METHODS FOR USE IN THE COMPUTATIONAL NETWORKTECHNICAL FIELDThe present disclosure relates generally to the field of high-performance computing; and more specifically, to a computationalnetwork, an orchestrator node in the computational network, a node in the computational network, and methods for the computational network, the orchestrator node in the computational network, and the node in the computational network, respectively. BACKGROUND Nowadays, there has been an exponential growth in complexity and scale of computational applications across diverse domains, spanning from high-performance computing (HPC) to artificial intelligence (AI). This surge in demand for computational resources has led to a natural progression towards distributing the computational applications across multiple nodes orprocessors to effectively handle the intensive workloads. The computational applications, such as HPC or AI include variousarithmetic operations including matrix operations. The matrix operations, such as solving large linear sets of equations ^^ =^, are used intensively in HPC or AI applications and dominating overall run time as well. Various math libraries, such as a portable, extensible toolkit for scientific computations (PETSc) include multiple linear solvers, where the solver performanceis highly dependable on specific matrix properties. One of the key properties of interest for accelerating the linear solverperformance is symmetry of the matrix. However, in order to identify if a large distributed matrix is symmetric or not, the identification requires significant communication between various nodes which hold different portions of the matrix. Therequired communication scales linearly with the matrix size, which introduces an overhead that prevents the HPC applicationsand linear solvers from attempting to classify whether a specific matrix is indeed symmetric and consequently, employ generalsolvers which do not take advantage of the symmetric property of the matrix. Moreover, the PETSc and other math librariesalready have efficient solvers for symmetric as well as structural symmetric matrices, but for efficient utilization of theseproperties, there is requirement of a computationally simple method to determine whether a given matrix is symmetric or not.Currently, certain attempts have been made to determine symmetric properties of a given matrix, such as use of a deterministicalgorithm. The deterministic algorithm causes the required communication volume to increase linearly with the matrix size which becomes a significant bottleneck in case of large matrices. Thus, there exists a technical problem of how to efficiently determine the symmetric properties of large matrices that are distributed among many computing devices. Therefore, in light of the foregoing discussion, there exists a need to overcome the aforementioned drawbacks associated withthe conventional ways of determining the symmetric properties of large matrices.SUMMARYThe present disclosure provides a computational network, an orchestrator node in the computational network, a node in thecomputational network, and methods for the computational network, the orchestrator node in the computational network, andthe node in the computational network, respectively. The present disclosure provides a solution to the existing problem of howto efficiently determine the symmetric properties of large matrices that are distributed among many computing devices. An aimof the present disclosure is to provide a solution that overcomes at least partially the problems encountered in prior art, andprovide an improved computational network, an improved orchestrator node in the computational network, an improved nodein the computational network, and improved methods for the computational network, the orchestrator node in the computational network, and the node in the computational network, respectively.The object of the present disclosure is achieved by the solutions provided in the enclosed independent claims. Advantageousimplementations of the present disclosure are further defined in the dependent claims.In one aspect, the present disclosure provides a computational network comprising an orchestrator controller and a plurality ofnodes each comprising a node controller, where the computational network is configured to execute a distributed application on the plurality of nodes, where the orchestrator controller is configured to determine whether a matrix, M, is a Symmetric Matrix, SM, wherein the matrix, M, is distributed into partial matrices, each partial matrix being on one of the plurality ofnodes, whereby the orchestrator controller is configured to determine that the matrix, M, is the SM by causing each nodecontroller to apply an antisymmetric function, f, to the cells of its partial matrix and summarize the results of the application of the antisymmetric function, f, thereby providing a partial matrix sum, whereby the orchestrator controller is further configuredto determine whether a total sum of all the partial matrix sums is zero and if so determine that the matrix is the SymmetricMatrix, where each node controller is further configured to summarize the results of the application of the antisymmetric function, f, by summarizing the results of the application of the antisymmetric function, f, modulo a prime constant P, P being a prime raised to the power of a constant x. The computational network manifests a significantly reduced communication volume between each of the plurality of nodes that is required in determining the symmetric properties of the matrix. The communication volume required in the computational network is proportional to the natural logarithm of the matrix size. The reduced communication volume results in improved scalability and utilization of large matrices. By virtue of an efficient identification of symmetric properties of thematrix, the computational network is suitable for HPC and AI applications by dividing the computations over the plurality ofnodes while minimizing the communication overhead among each of the plurality of nodes. Moreover, the computational network supports a trade-off between prediction probability of the matrix and the communication volume among each of the plurality of nodes. Also, less memory is required to save the symmetric matrix (or the structurally symmetric matrix).In an implementation form, each node controller is further configured to determine modulo P also for the partial matrix sum.The use of the modulo P for summarizing the partial matrix sums leads to a reduced communication volume.In a further implementation form, each node controller is further configured to apply the antisymmetric function, f, to the non-zero cells of its partial matrix. The application of the antisymmetric function, f, to the non-zero cells of the partial matrix results in reduced computational as well as communication overhead in the computational network.In another aspect, the present disclosure provides a method for a computational network. The method comprises determiningwhether a matrix, M, is a Symmetric Matrix, SM, wherein the matrix, M, is distributed into partial matrices, each partial matrix being on one of a plurality of nodes, whereby the method comprises determining that the matrix, M, is a SM by applying an antisymmetric function, f, to cells of the partial matrices, summarizing the results of the application of the antisymmetric function, f, thereby providing a partial matrix sum and determining whether a total sum of all the partial matrix sums is zero and if so determining that the matrix is a Symmetric Matrix, where the method further comprises each node controller summarizing the results of the application of the antisymmetric function, f, by summarizing the results of the application of the antisymmetric function, f, modulo a prime constant P, P being a prime raised to the power of a constant x.The method achieves all the advantages and technical effects of the computational network of the present disclosure.In a yet another aspect, the present disclosure provides a method for an orchestrator node in a computational network. Themethod comprises determining whether a matrix, M, is a Symmetric Matrix, SM, by receiving the matrix, M, distributing the matrix into partial matrices onto nodes, receiving partial matrix sums from the nodes partial and determining whether a total sum of all the partial matrix sums is zero and if so determining that the matrix is a Symmetric Matrix. The disclosed method enables an efficient determination of the symmetric properties of the matrix and hence, supports fast parallel computations in the computational network.In a yet another aspect, the present disclosure provides an orchestrator node in a computational network for determining whethera matrix, M, is a Symmetric Matrix, SM. The orchestrator node comprises an orchestrator controller configured to receive the matrix, M, distribute the matrix into partial matrices onto nodes, receive partial matrix sums from the nodes partial and determine whether a total sum of all the partial matrix sums is zero and if so determine that the matrix is a Symmetric Matrix,wherein each node controller is further configured to summarize the results of the application of the antisymmetric function, f,by summarizing the results of the application of the antisymmetric function, f, modulo a prime constant P, P being a prime raised to the power of a constant x.The orchestrator node achieves all the advantages and technical effects of the method for the orchestrator node after executionof the method.In a yet another aspect, the present disclosure provides a method for a node in a computational network. The method comprises determining whether a matrix, M, is a Symmetric Matrix, SM, by receiving a partial matrix, applying an antisymmetric function, f, to the cells of the partial matrix, summarizing the results of the application of the antisymmetric function, f, thereby providing a partial matrix sum, where the method further comprises summarizing the results of the application of the antisymmetric function, f, by summarizing the results of the application of the antisymmetric function, f, modulo a prime constant P, P being a prime raised to the power of a constant x. The method is based on summarizing the results of application of the antisymmetric function, f, over the modulo P, which significantly reduces the communication volume and leads to an efficient determination of symmetric properties of the matrix. In a yet another aspect, the present disclosure provides a node in a computational network for determining whether a matrix, M, is a Symmetric Matrix, SM. The node comprises a node controller configured to receive a partial matrix, apply an antisymmetric function, f, to the cells of the partial matrix, summarize the results of the application of the antisymmetricfunction, f, thereby providing a partial matrix sum, where the node controller is further configured to summarize the results ofthe application of the antisymmetric function, f, by summarizing the results of the application of the antisymmetric function, f, modulo a prime constant P, P being a prime raised to the power of a constant x. The node achieves all the advantages and technical effects of the method for the node in the computational network after execution of the method. It is to be appreciated that all the aforementioned implementation forms can be combined. It has to be noted that all devices, elements, circuitry, units and means described in the present application could be implemented in the software or hardware elements or any kind of combination thereof. All steps which are performed by the various entities described in the present application as well as the functionalities described to be performed by the various entities are intended to mean that the respective entity is adapted to or configured to perform the respective steps and functionalities. Even if, in the following description of specific embodiments, a specific functionality or step to be performed by external entities is notreflected in the description of a specific detailed element of that entity which performs that specific step or functionality, it should be clear for a skilled person that these methods and functionalities can be implemented in respective software or hardware elements, or any kind of combination thereof. It will be appreciated that features of the present disclosure are susceptible to being combined in various combinations without departing from the scope of the present disclosure as defined by the appended claims. Additional aspects, advantages, features and objects of the present disclosure would be made apparent from the drawings and the detailed description of the illustrative implementations construed in conjunction with the appended claims that follow. BRIEF DESCRIPTION OF THE DRAWINGS The summary above, as well as the following detailed description of illustrative embodiments, is better understood when readin conjunction with the appended drawings. For the purpose of illustrating the present disclosure, exemplary constructions ofthe disclosure are shown in the drawings. However, the present disclosure is not limited to specific methods andinstrumentalities disclosed herein. Moreover, those skilled in the art will understand that the drawings are not to scale. Whereverpossible, like elements have been indicated by identical numbers. Embodiments of the present disclosure will now be described, by way of example only, with reference to the following diagrams wherein:FIG. 1 is a network environment diagram of a computational network comprising an orchestrator controller and a pluralityof nodes, in accordance with an embodiment of the present disclosure; FIG.2 is a flowchart of a method for a computational network, in accordance with an embodiment of the present disclosure; FIG.3 is a flowchart of a method for an orchestrator node in a computational network, in accordance with an embodiment of the present disclosure; FIG.4 is a block diagram that illustrates various exemplary components of an orchestrator node, in accordance with an embodiment of the present disclosure; FIG.5 is a flowchart of a method for a node in a computational network, in accordance with an embodiment of the present disclosure; FIG.6 is a block diagram that illustrates various exemplary components of a node, in accordance with an embodiment of the present disclosure; andFIG. 7 is a graphical representation that depicts probability of true-negative classification for various matrix sizes, inaccordance with an embodiment of the present disclosure. In the accompanying drawings, an underlined number is employed to represent an item over which the underlined number ispositioned or an item to which the underlined number is adjacent. A non-underlined number relates to an item identified by aline linking the non-underlined number to the item. When a number is non-underlined and accompanied by an associated arrow, the non-underlined number is used to identify a general item at which the arrow is pointing. DETAILED DESCRIPTION OF EMBODIMENTS The following detailed description illustrates embodiments of the present disclosure and ways in which they can beimplemented. Although some modes of carrying out the present disclosure have been disclosed, those skilled in the art wouldrecognize that other embodiments for carrying out or practicing the present disclosure are also possible.FIG. 1 is a network environment diagram of a computational network comprising an orchestrator controller and a plurality ofnodes, in accordance with an embodiment of the present disclosure. With reference to FIG.1, there is shown a computationalnetwork 100 comprising an orchestrator controller 102 and a plurality of nodes 104, such as a first node 104A, a second node 104B, up to a Nth node 104N. Each of the plurality of nodes 104 comprises a node controller, for instance, the first node 104Acomprises a first node controller 106A, the second node 104B comprises a second node controller 106B and the Nth node 104Ncomprises a Nth node controller 106N. Furthermore, the orchestrator controller 102 and each of the plurality of nodes 104 are connected through a communication network 108.The computational network 100 may be configured to execute a distributed application on the plurality of nodes 104. Thecomputational network 100 may include hundreds or thousands of nodes which are connected to the orchestrator controller 102via the communication network 108. The computational network 100 may also be referred to as a distributed architecture.The orchestrator controller 102 may include suitable logic, circuitry, interfaces and / or code that is configured to determinewhether a matrix, M, is a Symmetric Matrix, SM. In an implementation, the orchestrator controller 102 may be referred to asa master controller. In another implementation, the orchestrator controller 102 may be referred to as a server. Examples of theorchestrator controller 102 may include, but are not limited to, a cloud server, a web server, an application server, a storageserver, or a combination thereof. Moreover, the orchestrator controller 102 may either be a single hardware server or a plurality of hardware servers operating in a parallel or distributed architecture to determine the symmetric properties of the matrix, M. In an implementation, the orchestrator controller 102 may be comprised by an orchestrator node, described in detail, forexample, in FIG. 4. For AI applications, the orchestrator controller 102 may comprise an AI model configured to executevarious Machine Learning (ML) algorithms, such as supervised ML algorithms, unsupervised ML algorithms, Deep Learning, DL, algorithms, Artificial Neural Network, ANN algorithms, and the like.Each of the plurality of nodes 104 may include suitable logic, circuitry, interfaces and / or code that is configured to process asubset of the matrix, M. Examples of each of the plurality of nodes 104 may include, but are not limited to, a computing devicein a computer cluster (e.g., massively parallel computer clusters), a communication apparatus including a portable or non- portable electronic device, and the like. Each of the plurality of nodes 104 comprises the node controller (e.g., 106A, 106B, up to 106N). Examples of the node controller may include, but are not limited to, a processor, an integrated circuit, a co-processor, a microprocessor, a microcontroller, a complex instruction set computing (CISC) processor, an application-specific integrated circuit (ASIC) processor, a reduced instruction set (RISC) processor, a very long instruction word (VLIW) processor, a central processing unit (CPU), a data processing unit, and other processors or circuits. Moreover, the node controller may refer to one or more individual processors, processing devices, a processing unit that is part of a machine.The communication network 108 may include suitable logic, circuitry, interfaces and / or code that is configured to connect theorchestrator controller 102 to each of the plurality of nodes 104. Examples of the communication network 108 may include,but are not limited to, a cellular network (e.g., a 5G, or 5G NR network, such as sub 6 GHz, cmWave, or mmWave communication network), a wireless sensor network (WSN), a cloud network, a Local Area Network (LAN), a vehicle-to- network (V2N) network, a Metropolitan Area Network (MAN), and / or the Internet.In operation, the computational network 100 is configured to execute a distributed application on the plurality of nodes 104,where the orchestrator controller 102 is configured to determine whether a matrix, M, is a Symmetric Matrix, SM. The matrix(M) is distributed into partial matrices, each partial matrix being on one of the plurality of nodes 104. For execution of thedistributed application, the computational network 100 is configured to utilize a probabilistic algorithm which significantly reduces the inter-communication between the plurality of nodes 104 required in determining the symmetric properties of the matrix, M. The matrix, M is such a large matrix which is distributed among the plurality of nodes 104 (e.g., computers orcomputing devices) in the computational network 100. For parallel computation, the matrix, M is distributed among theplurality of nodes 104 in form of partial matrices where, one partial matrix is stored on one of the plurality of nodes 104. Thepartial matrix may also be referred to as a subset of the matrix, M. For instance, the matrix, M may be represented as A andeach partial matrix (or a subset of matrix, M) may be represented as Aloc. The matrix, M, may also be referred to as a symmetricmatrix, SM, if^[^][^] = 0 ↔ ^[^][^] = 0where, [^] represents ith row of the matrix A and [^] represents jth column of the matrix A.The orchestrator controller 102 is configured to determine that the matrix, M, is the SM by causing each node controller toapply an antisymmetric function, f, to the cells of its partial matrix and summarize the results of the application of theantisymmetric function, f, thereby providing a partial matrix sum. The orchestrator controller 102 causes each node controllercomprised by each of the plurality of nodes 104 to apply the antisymmetric function, f, to each cell of its respective partialmatrix. After applying the antisymmetric function, f, to each cell of the respective partial matrix, each node controller is further configured to summarize the results of application of the antisymmetric function and provide the partial matrix sum.In accordance with an embodiment, the antisymmetric function is selected so that f(i,j) = -f(j,i). The antisymmetric function, f,is considered antisymmetric if f(i,j) = -f(j,i) for every i and j, where i, j represent row and column of matrix, A, respectively. In accordance with an embodiment, each node controller is further configured to apply the antisymmetric function, f, to thenon-zero cells of its partial matrix. The antisymmetric function, f, is applied to the non-zero cells of the respective partialmatrix, resulting in a reduced computational as well as communication overhead in the computational network 100.In accordance with an embodiment, the antisymmetric function is selected so that it has a unique sum V The antisymmetric function, f, is selected in such a way that the application of the antisymmetric function, f, over each partialmatrix provides a unique sum. Instead of the function, f, other functions may also be used as long as they are antisymmetricand have unique sum.In an implementation, when the antisymmetric function, f, is applied to each cell of the partial matrix (i.e., Aloc), themathematical notations may be:Given ^^^^(^, ^) = ^, the mapping function where (^ ≠ ^) will satisfy the following requirements: (i) ^ isantisymmetric, i.e., ^(^, ^, ^) = −^(^, ^, ^) (ii) any partial sum of ^ is unique such that ≠ ∑^,^∈^^ ^(^, ^, ^) , where^, ^ are any positive integers denoting the row and column, and ^^^^(^, ^) = ^.In accordance with an embodiment, the antisymmetric function is selected so that^(^, ^, ^) = ^^^^(^ − ^) ∙ 2^⋅^^^(^,^) ⋅ V , where: ^^^(^, ^) is a unique 2D^1D mapping, such that ^^^(^, ^) = ^^^(^, ^) ↔(^, ^) = (^, ^), wherein f places all r bits of V in a unique bit-offset determined by 2^(r⋅ind(i, j) ) to satisfy the unique sumproperty. In an implementation, each value (i.e., V) of the partial matrix (i.e., ^^^^(^, ^) = ^) may be represented using r bits,and the antisymmetric function, f may be selected in such a way that^(^, ^, ^) = ^^^^(^ − ^) ∙ 2^⋅^^^(^,^) ⋅ V, where ^^^(^, ^) is a unique 2D^1D mapping, such that ^^^(^, ^) = ^^^(^, ^) ↔(^, ^) = (^, ^), as follow: The aforementioned antisymmetric function, f, replaces all r bits of V in the unique bit-offset determined by 2^(r⋅ind(i,j)) tosatisfy the unique sum property. Moreover, the ^^^^(^ − ^) ensures that the function, f, satisfies the antisymmetric property.In accordance with an embodiment, the antisymmetric function is selected so that r =1. In order to identify that the matrix, Mis structurally symmetric, it is sufficient to use a single bit representation (i.e., r = 1) for each V to indicate its value as zero, 0.However, in any other case, r may have different value. For example, if all the matrix values of the matrix (i.e., A) are integer values that means there is requirement of 4 bytes which is equivalent to 32 bits, hence, value of r will be 32. In accordance with an embodiment, the antisymmetric function is selected so that|f(i,j)| ≤ 2^(r⋅matrixSize-1) and ∑|f(i,j)| ≤ 2^(r⋅matrixSize). The antisymmetric function, f, is selected in such a way that thefunction, f, possesses the following properties:|f(i,j)| ≤ 2^(r⋅matrixSize-1) and ∑|f(i,j)| ≤ 2^(r⋅matrixSize). The equation |f(i,j)| ≤ 2^(r⋅matrixSize-1) involves the absolute valueof the function f(i, j). Moreover, this equation states that the absolute value of the function at position (i, j) in the matrix (i.e., A) must be less than or equal to 2^(r⋅matrixSize-1), where r is a constant and matrixSize is the size of the matrix (i.e., A). Theequation ∑|f(i,j)| ≤ 2^(r⋅matrixSize) involves the sum of the absolute values of the function over all positions (i, j) in the matrix(i.e., A). This equation states that the sum of the absolute values of the function for all elements in the matrix must be less than or equal to 2^(r⋅matrixSize). The equation |f(i,j)| ≤ 2^(r⋅matrixSize-1) sets an upper bound on magnitude of individual elements of the matrix (i.e., A), while the equation ∑|f(i,j)| ≤ 2^(r⋅matrixSize) sets an upper bound on sum of magnitudes of all elementsin the matrix (i.e., A). The parameter r influences the scale of these bounds. However, the antisymmetric function, f, is notlimited to aforementioned bounds and may have different values of bounds depending on different implementation scenarios. The orchestrator controller 102 is further configured to determine whether a total sum of all the partial matrix sums is zero and if so determine that the matrix is a Symmetric Matrix, where each node controller is further configured to summarize the results of the application of the antisymmetric function, f, by summarizing the results of the application of the antisymmetric function,f, modulo a prime constant P, P being a prime raised to the power of a constant x. The orchestrator controller 102 is furtherconfigured to determine the total sum of all the partial matrix sums received from each of the plurality of nodes 104 is zero and if the total sum is determined as zero, the matrix, M is determined as the symmetric matrix, SM. For computing the partial matrix sum at each node, each node controller is configured to summarize the results of application of the antisymmetricfunction, f, by summarizing the results of the application of the antisymmetric function, f, over modulo P, where P is raised tothe power of the constant x, where x = 1, 2, …, n. P is the power of a prime number. Typically, a prime number is defined as anatural number that is divisible only by 1 and itself. According to a typical prime number theorem, for a given number, n, there are approximately n / (log(n)) prime numbers in the range 1, …, n. The fundamental theorem of arithmetic states that every integer larger than 1 can be written as a product of one or more primes (also known as prime factorization), for example:34866=2×3^2×13×149 and such a product is unique that means there is exactly one prime factorization for each number.Furthermore, the prime lemma defines that if ^^ , ^^ are prime number factors in the factorization of ^, then ^^ ∙ ^^|^ (i.e., ^is divisible by the product ^^ ∙ ^^).In accordance with an embodiment, each node controller is further configured to determine modulo P also for the partial matrix sum. Each node controller is configured to summarize the results of application of the antisymmetric function, f, by use of themodule P. For instance, the partial matrix sum for each of the plurality of nodes 104 may be denoted as Gloc, ^^^^ =∑^,^∈^^^^ [^(^, ^, ^) ^^^^^^ ^] ^^^^^^ ^ where, P is a power of a prime number. The modulo operation reduces the numberof bits required to represent the partial matrix sum (i.e., ^^^^). Although, the modulo operation introduces a probability of error in classifying a non-symmetric matrix as the symmetric matrix (i.e., true-negative). The probability of error can be controlledvia the selection of the constant x. The unique sum property of the antisymmetric function, f, provides an upper bound for theprobability of true-negative classification, regardless of the specific matrix values. In accordance with an embodiment, the prime constant P is selected to be less than 2 raised to the power of the number of bitsrequired to store the partial matrix sum. The prime constant P is selected such that P < 2x, and x is the number of bits selectedfor storing and communicating the partial matrix sum of each node.For determining the total sum of all the partial matrix sums received from each of the plurality of nodes 104, the orchestratorcontroller 102 is configured to apply a global reduction operation (may be denoted as on all the partial matrix sums as ^^^^^ denotes the set of nodes (i.e., the plurality of nodes 104) that store the various parts of thematrix (i.e., A). The global reduction operation requires communication of up to log(P) bits from every node and represent -P<Furthermore, the orchestrator controller 102 is configured to check whether the total sum of all the partial matrix sums is zero (i.e., ^^^^^^^= 0) to determine the matrix (i.e., A) as the symmetric matrix. The antisymmetric property of the function, f, ensures that the total sum of all the partial sums is zero (i.e., ^^^^^^^= 0) for any symmetric matrix.For example, given a matrix, A which is partitioned in four partial matrices (i.e., ^^, ^^, ^^, and ^^) and each node controlleris assigned one partial matrix. The orchestrator controller 102 is configured to determine whether the matrix, A, is symmetric matrix (or structurally symmetric matrix). In order to ease the process, the matrix, A, is converted into a binary matrix, where 1 represents a non-zero entry of the matrix, A, and 0 represents a zero entry of the matrix, A. The orchestrator controller 102 causes each node controller to apply the antisymmetric function, f, over the cells of the partial matrix and compute the partial matrix sum. The antisymmetric function, f, is applied on the non-zero cells of the partial matrix. ^(^^) = −2^ − 2^ = −6 In this example, 2^bits are required for storing and communicating the partial matrix sum of each node. Moreover, all the numbers in this example, are so small therefore no modulo P is taken. In another example, where the partial matrix sums are of very large values then, the prime number, P, is selected randomly.Thus, the computational network 100 manifests a significantly reduced communication volume between each of the pluralityof nodes 104 that is required in determining the symmetric properties of the matrix (i.e., A). The communication volumerequired in the computational network 100 is proportional to the natural logarithm of the matrix size. For instance, if the matrix(i.e., A) has n elements then, the communication volume required is O(log n) which, is achieved by use of the modulo P duringcomputation of the total sum of the all the partial matrix sums. The modulo P is <2x, where x is the number of bits selected forstoring and communicating the partial matrix sum. The reduced communication volume results in improved scalability andutilization of large matrices. By virtue of an efficient identification of symmetric properties of the matrix (i.e., A), thecomputational network 100 is suitable for HPC and AI applications by dividing the computations over the plurality of nodes104 while minimizing the communication overhead among each of the plurality of nodes 104. Moreover, the computationalnetwork 100 supports a trade-off between prediction probability of the matrix (i.e., A) and the communication volume amongeach of the plurality of nodes 104. Also, less memory is required to save the symmetric matrix (or the structurally symmetricmatrix, i.e., A). The use of the symmetric matrix leads to fast computations of the partial matrix at each of the plurality of nodes 104.FIG. 2 is a flowchart of a method for a computational network, in accordance with an embodiment of the present disclosure.FIG. 2 is described in conjunction with elements from FIG. 1. With reference to FIG. 2, there is shown a method 200 thatincludes steps 202 to 208. The method 200 is executed by the orchestrator controller 102, and the node controller of each ofthe plurality of nodes 104 of the computational network 100 (of FIG.1).There is provided the method 200 for the computational network 100 (of FIG. 1). The method 200 is based on a probabilisticalgorithm in order to determine whether a matrix is symmetric or not. The utilization of the probabilistic algorithm allows the communication volume to increase logarithmically with the matrix size while introducing a negligible change of true-negative classification which differentiates the probabilistic algorithm from a typical deterministic algorithm. The use of the typical deterministic algorithm causes the communication volume to increase linearly with the matrix size and therefore, introduces a significant communication overhead. The introduction of the significant communication overhead is a bottleneck indetermining the symmetric properties of large matrices, which deters the use of the deterministic algorithm. Therefore, theprobabilistic algorithm is preferred over the deterministic algorithm.At step 202, the method 200 comprises determining whether a matrix, M, is a Symmetric Matrix, SM, wherein the matrix, M,is distributed into partial matrices, each partial matrix being on one of a plurality of nodes. The matrix, M, is such a large matrixwhich is distributed to many computing devices or nodes (i.e., the plurality of nodes 104) in the computational network 100.The matrix, M, is distributed into partial matrices and each partial matrix is stored on each of the plurality of nodes 104 in the computational network 100.At step 204, the method 200 further comprises determining that the matrix, M, is a SM by applying an antisymmetric function,f, to the cells of the partial matrices. In order to determine whether the matrix, M is the symmetric matrix, SM, the antisymmetricfunction, f, is applied to all the cells of each partial matrix stored on each of the plurality of nodes 104.At step 206, the method 200 further comprises summarizing the results of the application of the antisymmetric function, f,thereby providing a partial matrix sum. After applying the antisymmetric function, f, to all the cells of each partial matrix, theresults of the application of the antisymmetric function, f, are summarized to obtain the partial matrix sum for each of theplurality of nodes 104. The details have been provided, for example, in FIG. 1.At step 208, the method 200 further comprises determining whether a total sum of all the partial matrix sums is zero and if sodetermining that the matrix is a Symmetric Matrix, wherein the method 200 further comprises each node controllersummarizing the results of the application of the antisymmetric function, f, by summarizing the results of the application of theantisymmetric function, f, modulo a prime constant P, P being a prime raised to the power of a constant x. Furthermore, it isdetermined that the total sum of all the partial matrix sums is zero and if so, it is determined that the matrix is the symmetricmatrix. The results of the application of the antisymmetric function, f, are summarized by summarizing the results of theapplication of the antisymmetric function, f, modulo the prime constant P, P being the prime raised to the power of the constant x, described in detail, for example, in FIG.1.The steps 202 to 208 are only illustrative and other alternatives can also be provided where one or more steps are added, oneor more steps are removed, or one or more steps are provided in a different sequence without departing from the scope of the claims herein.In one aspect, the present disclosure provides a computer program product comprising program instructions for performing themethod 200, when executed by one or more processors (e.g., the orchestrator controller 102 and each of the plurality of nodes104) in a computational network (e.g., the computational network 100, of FIG. 1). In a yet another aspect, the present disclosureprovides a computer-readable storage medium comprising instructions which, when executed by a computer, cause the computer to carry out the method 200 for the computational network 100.FIG. 3 is a flowchart of a method for an orchestrator node in a computational network, in accordance with an embodiment ofthe present disclosure. FIG. 3 is described in conjunction with elements from FIGs. 1 and 2. With reference to FIG. 3, there isshown a method 300 that includes steps 302 to 308. There is provided the method 300 for an orchestrator node in a computational network (e.g., the computational network 100).At step 302, the method 300 comprises determining whether a matrix, M, is a Symmetric Matrix, SM, by receiving the matrix,M.At step 304, the method 300 further comprises distributing the matrix into partial matrices onto nodes. The matrix is dividedinto various subsets, named as partial matrices and each partial matrix is distributed onto each of the nodes (i.e., the pluralityof nodes 104 of FIG. 1). Thereafter, processing is done on each partial matrix at each node, such as applying an antisymmetricfunction, f, over all the cells of each partial matrix and summarize the results of application of the antisymmetric function, f, toobtain a partial matrix sum from each node.At step 306, the method 300 further comprises receiving partial matrix sums from the nodes partial. For example, the partialmatrix sums computed at each node is received at the orchestrator controller 102 (of FIG.1).At step 308, the method 300 further comprises determining whether a total sum of all the partial matrix sums is zero and if sodetermining that the matrix is a Symmetric Matrix. After receiving the partial matrix sum from each node, the total sum of allthe partial matrix sums is computed and checked whether the total sum is zero or not. In case of having the total sum as zero, the matrix is determined as the symmetric matrix. The symmetric matrix is helpful in performing the computations in an efficient way when used in HPC or AI applications.The steps 302 to 308 are only illustrative and other alternatives can also be provided where one or more steps are added, oneor more steps are removed, or one or more steps are provided in a different sequence without departing from the scope of theclaims herein.In one aspect, the present disclosure provides a computer program product comprising program instructions for performing themethod 300, when executed by one or more processors (e.g., the orchestrator controller 102) in a computational network (e.g.,the computational network 100, of FIG. 1). In a yet another aspect, the present disclosure provides a computer-readable storagemedium comprising instructions which, when executed by a computer, cause the computer to carry out the method 300 for the orchestrator controller 102 of the computational network 100. FIG. 4 is a block diagram that illustrates various exemplary components of an orchestrator node, in accordance with an embodiment of the present disclosure. FIG.4 is described in conjunction with elements from FIGs.1, 2, and 3. With referenceto FIG. 4, there is shown an orchestrator node 402 that comprises an orchestrator controller 404, a memory 406 and a networkinterface 408. The orchestrator node 402 comprising the orchestrator controller 404 is configured to execute the method 300of FIG.3.The orchestrator node 402 may include suitable logic, circuitry, interfaces and / or code that is configured to determine whethera matrix, M, is a symmetric matrix, SM. The orchestrator node 402 may be used in a computational network, for example, thecomputational network 100, of FIG. 1. The orchestrator node 402 may also be referred to as a master node comprising theorchestrator controller 404. The orchestrator controller 404 corresponds to the orchestrator controller 102 of FIG.1.The memory 406 may include suitable logic, circuitry, interfaces and / or code that is configured to store machine code and / orinstructions executable by the orchestrator controller 404. Examples of implementation of the memory 406 may include, butare not limited to, an Electrically Erasable Programmable Read-Only Memory (EEPROM), Random Access Memory (RAM), Read Only Memory (ROM), Hard Disk Drive (HDD), Flash memory, a Secure Digital (SD) card, Solid-State Drive (SSD), acomputer readable storage medium, and / or CPU cache memory. The memory 406 may store an operating system and / or acomputer program product to operate the orchestrator node 402. A computer readable storage medium for providing a non-transient memory may include, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing.The network interface 408 may include suitable logic, circuitry, interfaces, or code that is communicatively coupled with thememory 406 and the orchestrator controller 404. Examples of the network interface 408 include, but are not limited to, a dataterminal, a transceiver, a facsimile machine, and the like.In operation, the orchestrator node 402 comprising the orchestrator controller 404 is configured to receive the matrix, M,distribute the matrix into partial matrices onto nodes, receive partial matrix sums from the nodes partial and determine whether a total sum of all the partial matrix sums is zero and if so determine that the matrix is a Symmetric Matrix, wherein each node controller is further configured to summarize the results of the application of the antisymmetric function, f, by summarizing the results of the application of the antisymmetric function, f, modulo a prime constant P, P being a prime raised to the powerof a constant x. The orchestrator node 402 is configured to receive the matrix M which is distributed over multiple nodes (e.g.,the plurality of nodes 104 of FIG.1) in terms of the partial matrices where, one partial matrix is stored on one of the multiple nodes. Each nodes comprises a node controller which is configured to process the partial matrix by applying an antisymmetric function, f, over all the cells of the respective partial matrix and then, summarize the results of application of the antisymmetric function, f, to generate a partial matrix sum for each node. The orchestrator node 402 is configured to receive the partial matrix sum from each node and compute the total sum of all the partial matrix sums and further check whether the total sum is zero or not. In case of determining the total sum as zero, the orchestrator node 402 is configured to determine that the matrix is the symmetric matrix. The antisymmetric function, f, and the summarizing the results of application of the antisymmetric function, f, modulo the prime constant P, P being the prime raised to the power of the constant x, have been described in detail, for example, in FIG.1.FIG. 5 is a flowchart of a method for a node in a computational network, in accordance with an embodiment of the presentdisclosure. FIG.5 is described in conjunction with elements from FIGs.1, 2, 3 and 4. With reference to FIG.5, there is showna method 500 that includes steps 502 to 508. The method 500 is executed by each of the plurality of nodes 104 of thecomputational network 100 (of FIG. 1).There is provided the method 500 for a node (e.g., each of the plurality of nodes 104) in a computational network (e.g., thecomputational network 100, of FIG. 1).At step 502, the method 500 comprises determining whether a matrix, M, is a Symmetric Matrix, SM, by receiving a partialmatrix. In order to determine the matrix, M as the symmetric matrix, SM, the matrix, M is divided into a number of partialmatrices.At step 504, the method 500 further comprises applying an antisymmetric function, f, to the cells of the partial matrix. Theantisymmetric function, f, is applied on each of the partial matrices and all the cells of the partial matrix.At step 506, the method 500 further comprises summarizing the results of the application of the antisymmetric function, f,thereby providing a partial matrix sum. The results of the application of the antisymmetric function, f, are summarized to obtainthe partial matrix sum.At step 508, the method 500 further comprises summarizing the results of the application of the antisymmetric function, f, bysummarizing the results of the application of the antisymmetric function, f, modulo a prime constant P, P being a prime raisedto the power of a constant x. The results of the application of the antisymmetric function, f, are summarized by summarizingthe results of the application of the antisymmetric function, f, modulo the prime constant P, P being the prime raised to the power of the constant x, where x = 1, 2, …, n.The steps 502 to 508 are only illustrative and other alternatives can also be provided where one or more steps are added, oneor more steps are removed, or one or more steps are provided in a different sequence without departing from the scope of the claims herein.In one aspect, the present disclosure provides a computer program product comprising program instructions for performing themethod 500, when executed by one or more processors (e.g., each node of the plurality of nodes 104) in a computationalnetwork (e.g., the computational network 100, of FIG. 1). In a yet another aspect, the present disclosure provides a computer-readable storage medium comprising instructions which, when executed by a computer, cause the computer to carry out the method 500 for each node of the plurality of nodes 104 of the computational network 100.FIG. 6 is a block diagram that illustrates various exemplary components of a node, in accordance with an embodiment of thepresent disclosure. FIG.6 is described in conjunction with elements from FIGs.1, 2, 3, 4 and 5. With reference to FIG.6, thereis shown a node 602 that comprises a node controller 604, a memory 606 and a network interface 608. The node 602 comprisingthe node controller 604 is configured to execute the method 500 of FIG.5.The node 602 may include suitable logic, circuitry, interfaces and / or code that is configured to determine whether a matrix, M,is a Symmetric Matrix, SM. The node 602 corresponds to each node of the plurality of nodes 104 of the computational network100, of FIG. 1. Similarly, the node controller 604 corresponds to the node controller comprised by each of the plurality of nodes104. The memory 606 may include suitable logic, circuitry, interfaces and / or code that is configured to store machine code and / orinstructions executable by the node controller 604. Examples of implementation of the memory 606 are similar to the examplesof implementation of the memory 406 (of FIG.4). The network interface 608 may include suitable logic, circuitry, interfaces, or code that is communicatively coupled with thememory 606 and the node controller 604. Examples of the network interface 608 are similar to the examples of the networkinterface 408 (of FIG.4). In operation, the node 602 comprising the node controller 604 is configured to receive a partial matrix, apply an antisymmetric function, f, to the cells of the partial matrix, summarize the results of the application of the antisymmetric function, f, thereby providing a partial matrix sum, where the node controller 604 is further configured to summarize the results of the application of the antisymmetric function, f, by summarizing the results of the application of the antisymmetric function, f, modulo a primeconstant P, P being a prime raised to the power of a constant x. The node 602 comprising the node controller 604 is configuredto receive the partial matrix, (i.e., subset of the matrix, M) and then, apply the antisymmetric function, f, over all the cells of the partial matrix and summarize the results of application of the antisymmetric function, f, to compute the partial matrix sum. The results of application of the antisymmetric function, f, are summarized by use of modulo prime constant P with the results of the application of the antisymmetric function, f, where P is prime raised to the power of the constant x.FIG. 7 is a graphical representation that depicts probability of true-negative classification for various matrix sizes, in accordancewith an embodiment of the present disclosure. FIG.7 is described in conjunction with elements from FIGs.1, 2, 3, 4, 5, and 6. With reference to FIG.7, there is shown a graphical representation 700 that depicts probability of true-negative classificationfor various matrix sizes, assuming each variable V requires ^ = 32 bits representation.In the graphical representation 700, there is an X-axis 702 that represents bits and a Y-axis 704 that represents chances of true-negative in logarithmic scale. In the graphical representation 700, this may be observed that the number of bits, x, required to communicate to ensure negligible probability of true-negative classification is significantly lower when compared to the typical deterministic algorithm for which the communication among the multiple nodes is proportional to the matrix size. Modifications to embodiments of the present disclosure described in the foregoing are possible without departing from the scope of the present disclosure as defined by the accompanying claims. Expressions such as "including", "comprising", "incorporating", "have", "is" used to describe and claim the present disclosure are intended to be construed in a non-exclusivemanner, namely allowing for items, components or elements not explicitly described also to be present. Reference to thesingular is also to be construed to relate to the plural. The word "exemplary" is used herein to mean "serving as an example, instance or illustration". Any embodiment described as “exemplary” is not necessarily to be construed as preferred or advantageous over other embodiments and / or to exclude the incorporation of features from other embodiments. The word "optionally" is used herein to mean "is provided in some embodiments and not provided in other embodiments". It is appreciated that certain features of the present disclosure, which are, for clarity, described in the context of separate embodiments, may also be provided in combination in a single embodiment. Conversely, various features of the present disclosure, which are, for brevity, described in the context of a single embodiment, may also be provided separately or in any suitable combination or as suitable in any other described embodiment of the disclosure.
Claims
CLAIMS1. A computational network (100) comprising an orchestrator controller (102) and a plurality of nodes (104) each comprisinga node controller, wherein the computational network (100) is configured to execute a distributed application on the pluralityof nodes (104), wherein the orchestrator controller (102) is configured to determine whether a matrix, M, is a SymmetricMatrix, SM, wherein the matrix, M, is distributed into partial matrices, each partial matrix being on one of the plurality of nodes (104), whereby theorchestrator controller (102) is configured to determine that the matrix, M, is the SM by causing each node controller toapply an antisymmetric function, f, to cells of its partial matrix and summarize the results of the application of the antisymmetric function, f, thereby providing a partial matrixsum, whereby the orchestrator controller (102) is further configured todetermine whether a total sum of all the partial matrix sums is zero and if so determine that the matrix is the Symmetric Matrix, wherein each node controller is further configured to summarize the results of the application of the antisymmetric function, f, by summarizing the results of the application of the antisymmetric function, f, modulo a primeconstant P, P being a prime raised to the power of a constant x.
2. The computational network (100) according to claim 1, wherein the antisymmetric function is selected so thatf(i, j) = -f(j, i).
3. The computational network (100) according to claim 2, wherein the antisymmetric function is selected so that it has a uniquesum V ⋁(^>^,^_^>^_^) ^(^,^)=^^^(^^, nk)→^=1 ^^^ (^,^)=(^1, n1)4. The computational network (100) according to claim 3, wherein the antisymmetric function is selected so that⋅ V , where: ^^^(^, ^) is a unique 2D^1D mapping, such that ^^^(^, ^) = ^^^(^, ^) ↔(^, ^) = (^, ^), wherein f places all r bits of V in a unique bit-offset determined by 2^(r⋅ind(i,j) ) to satisfy the unique sumproperty.
5. The computational network (100) according to claim 4, wherein the antisymmetric function is selected so that r =1.
6. The computational network (100) according to claim 4 or 5, wherein the antisymmetric function is selected so that|f(I,j)| ≤ 2^(r⋅matrixSize-1) and ∑|f(i,j)| ≤ 2^(r⋅matrixSize).
7. The computational network (100) according to any preceding claim, wherein each node controller is further configured todetermine modulo P also for the partial matrix sum.
8. The computational network (100) according to claim 7, wherein the prime constant P is selected to be less than 2 raised tothe power of the number of bits required to store the partial matrix sum.
9. The computational network (100) according to any preceding claim, wherein each node controller is further configured toapply the antisymmetric function, f, to non-zero cells of its partial matrix.
10. A method (200) for a computational network (100), the method (200) comprising determining whether a matrix, M, is aSymmetric Matrix, SM, wherein the matrix, M, isdistributed into partial matrices, each partial matrix being on one of a plurality of nodes (104), whereby the method (200)comprises determining that the matrix, M, is a SM by applying an antisymmetric function, f, to cells of the partial matrices, summarizing the results of the application of the antisymmetric function, f, thereby providing a partial matrix sum and determining whether a total sum of all the partial matrix sums is zero and if so determining that the matrix is a Symmetric Matrix, wherein the method (200) further comprises each node controllersummarizing the results of the application of the antisymmetric function, f, by summarizing the results of the application of theantisymmetric function, f, modulo a prime constant P, P being a prime raised to the power of a constant x.
11. A method (300) for an orchestrator node (402) in a computational network (100), the method (300) comprising determiningwhether a matrix, M, is a Symmetric Matrix, SM, by receiving the matrix, M, distributing the matrix into partial matrices onto nodes, receiving partial matrix sums from the nodes partial and determining whether a total sum of all the partial matrix sums is zero and if so determining that the matrix is a Symmetric Matrix.12 An orchestrator node (402) in a computational network (100) for determining whether a matrix, M, is a Symmetric Matrix,SM, the orchestrator node (402) comprising an orchestrator controller (404) configured toreceive the matrix, M, distribute the matrix into partial matrices onto nodes, receive partial matrix sums from the nodes partial and determine whether a total sum of all the partial matrix sums is zero and if so determine that the matrix is a Symmetric Matrix, wherein each node controller is further configured to summarize the results of the application of the antisymmetric function, f, by summarizing the results of the application of the antisymmetric function, f, modulo a prime constant P, P being a prime raised to the power of a constant x.
13. A method (500) for a node (602) in a computational network (100), the method (500) comprising determining whether amatrix, M, is a Symmetric Matrix, SM, by receiving a partial matrix, applying an antisymmetric function, f, to cells of the partial matrix, summarizing the results of the application of the antisymmetric function, f, thereby providing a partial matrix sum,wherein the method (500) further comprises summarizing the results of the application of the antisymmetric function, f, bysummarizing the results of the application of the antisymmetric function, f, modulo a prime constant P, P being a prime raised to the power of a constant x.
14. A computer program product comprising program instructions for performing the method (100; 300; 500) according to anyof claims 10, 11 or 13, when executed by one or more processors in a computational network (100).
15. A node (602) in a computational network (100) for determining whether a matrix, M, is a Symmetric Matrix, SM, the node(602) comprising a node controller (604) configured toreceive a partial matrix,apply an antisymmetric function, f, to cells of the partial matrix, summarize the results of the application of the antisymmetric function, f, thereby providing a partial matrix sum,wherein the node controller (604) is further configured to summarize the results of the application of the antisymmetric function,f, by summarizing the results of the application of the antisymmetric function, f, modulo a prime constant P, P being a primeraised to the power of a constant x.
Citation Information
Patent Citations
Processor for Large Graph Algorithm Computations and Matrix Operations
US20110307685A1