Data transmission optimization method and system based on network coding
By adopting a network encoding-based data transmission optimization method in the MapReduce cluster and combining with the greedy grouping strategy, the communication bottleneck problem caused by the different byte lengths of the Map intermediate value is solved, and the optimization of data transmission volume and the improvement of computing performance is achieved.
Patent Information
- Application Number
- CN202510212551.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-05-16
AI Technical Summary
The prior art In MapReduce cluster, due to the limitation of network bandwidth, the communication bottleneck problem during data reshuffle is caused. Especially when the byte lengths of the middle value of Map is not equal, the encoding technology cannot effectively solve it.
The data transmission optimization method based on network encoding is adopted, and the data transmission volume in the Shuffle stage is optimized by calculating the encoding and decoding operations between nodes, and a greedy grouping strategy is proposed to adapt to the byte length of the intermediate value of different output functions.
It effectively reduces the data transmission volume and algorithm design complexity in the Shuffle stage, improves the reliability and security of data transmission, shortens the running time of data shuffle, and thus optimizes the data shuffle process.
Smart Images

Figure CN120017652A_ABST
Abstract
Description
Background Art
[0002] With the rapid development of Internet services, data centers, as the infrastructure supporting Internet services, are also growing in size. Data centers store large amounts of data, communicate frequently between servers, and perform collaborative computing to provide services to the outside world. In the data center cluster computing process, a large amount of data exchange and data shuffling are often required between nodes. However, in a data center cluster system, bandwidth resources are usually relatively limited, and a large amount of data communication will constitute a performance bottleneck for the overall computing. Therefore, it is of great practical significance to focus on the data shuffling optimization problem in data centers.
[0003] MapReduce is a paradigm for parallel processing of big data. Its main processes include the Map phase, the Shuffle phase, and the Reduce phase. Shuffle requires data to be transferred between map and reduce tasks, which will result in excessive network consumption. When the size of the MapReduce cluster grows, the communication bottleneck problem caused by excessive network consumption becomes more serious.
[0004] The communication bottleneck problem in distributed systems can be overcome through coding technology. In 2017, Li et al. proposed the CDC (Coded Distributed Computing) strategy applied to the MapReduce framework to compress the amount of data transmitted in the Shuffle phase through coding, and characterized the trade-off between communication overhead and computing overhead. However, when designing the CDC algorithm, it made an idealized assumption about the intermediate values output in the Map phase, and only designed the encoding scheme with the best performance for the case where the byte lengths of all Map intermediate values are equal. In addition, there is no relevant research on the encoding technology for the case where the byte lengths of Map intermediate values are unequal. Summary of the invention
[0005] The technical problem to be solved by the present invention is to provide a data transmission optimization method and system based on network coding in view of the deficiencies in the above-mentioned prior art, so as to solve the technical problem of poor network data transmission performance.
[0006] The present invention adopts the following technical solutions:
[0007] A data transmission optimization method based on network coding comprises the following steps:
[0008] Map phase: In a MapReduce cluster consisting of K distributed computing nodes, computing node k calculates the input file set M stored on itself. k , k∈{1,…,K}; computing node k executes the Map function to obtain Q Map calculation intermediate values, and each input file is stored on at least one node;
[0009] Shuffle phase: Assign computing node k to be responsible for outputting function set W k The output result of
[0010] Reduce phase: computing node k is responsible for outputting the function set W k After the Shuffle phase, the computing node k receives the intermediate value code, and then decodes it to obtain the output function set W in combination with the locally known intermediate value. k Each output function φ q All N Maps calculate the intermediate value v q,1 ,…,v q,N ; Finally, computing node k executes the Reduce function to calculate the output result u q .
[0011] Preferably, the input file set M k There are N input files ω1,ω2,…,ω N , corresponding to Q output functions φ1, φ2, …, φ Q The output is u q =φ q (ω1,ω2,…,ω N ), q∈{1,2,…,Q}, output result u q Decomposed into the Map phase to calculate the intermediate value and the Reduce phase to merge and calculate the intermediate value output results; the Map phase calculates the intermediate value g q,1 ,…,g q,N For the corresponding input files ω1, ω2, …, ω N The calculated intermediate values; all N intermediate values merged in the Reduce phase are passed through the Reduce function h q Calculate the output result.
[0012] Preferably, the Map function calculates the input file ω n , the input file ω n Q Map calculation intermediate values are obtained by calculation.
[0013] Preferably, the Reduce function h q Responsible for statistical output function φ Q To calculate the result, we first need to obtain all the N required Map calculation intermediate values v through Shuffle data exchange. q,1 ,…,v q,N , then calculate the output result u q =h q (v q,1 ,…,v q,N ), q∈{1,2,…,Q}, v q,1 ,…,vq,N Calculate the intermediate value for Map, h q is the reduce function.
[0014] Preferably, in the Shuffle phase:
[0015] Node k needs all N Map intermediate values v q,1 ,…,v q,N ; Some intermediate values are known through calculation in their own Map phase, and the remaining intermediate values Obtained through other node network transmission;
[0016] For each node k, encode the locally known intermediate value of the Map calculation to obtain the intermediate value code X, and the encoding function is recorded as ψ k ; Afterwards, node k broadcasts X k to all other nodes.
[0017] Preferably, in the Shuffle phase, for any given MapReduce model, i.e., given a set of input files M k , output function set W k , encoding function ψ k , decoding function And the computational overhead r, we get the communication overhead L, and form a computation-communication relationship pair (r, L).
[0018] Preferably, the computation-communication function relationship is:
[0019]
[0020] Among them, l is the byte length of all map intermediate values, N is the number of input files, Q is the number of output functions, K is the number of distributed computing nodes, and r is the computing overhead.
[0021] Preferably, the computational cost r is:
[0022]
[0023] Among them, M k is the set of input files that node k is responsible for.
[0024] Preferably, the communication overhead L is:
[0025]
[0026] in, is node k, k∈{1,…,K}, is the coded data X to be sent k The total length in bytes.
[0027] In a second aspect, an embodiment of the present invention provides a data transmission optimization system based on network coding, including:
[0028] The Map module calculates the input file set M stored on itself through computing node k in a MapReduce cluster consisting of K distributed computing nodes. k , k∈{1,…,K}; computing node k executes the Map function to obtain Q Map calculation intermediate values, and each input file is stored on at least one node;
[0029] Shuffle module, which assigns computing node k to be responsible for outputting function set W k The output result of
[0030] Reduce module, computing node k is responsible for outputting the function set W k After the Shuffle phase, the computing node k receives the intermediate value code, and then decodes it to obtain the output function set W in combination with the locally known intermediate value. k Each output function φ q All N Maps calculate the intermediate value v q,1 ,…,v q,N ; Finally, computing node k executes the Reduce function to calculate the output result u q .
[0031] In a third aspect, a computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above-mentioned network coding-based data transmission optimization method when executing the computer program.
[0032] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, including a computer program, which, when executed by a processor, implements the steps of the above-mentioned network coding-based data transmission optimization method.
[0033] Compared with the prior art, the present invention has at least the following beneficial effects:
[0034] A data transmission optimization method based on network coding has high feasibility in improving network data transmission performance by increasing the computational overhead of network nodes. It uses network coding to improve network data transmission performance and shortens the running time of data shuffling, thereby optimizing the data shuffling process. On the premise of maintaining network coding to improve network data transmission performance, the data transmission process based on network coding is optimized, and the data transmission delay and computational overhead in coding are effectively controlled appropriately, which has important theoretical and practical significance for promoting the large-scale application of network coding.
[0035] Furthermore, network coding is combined with multipath routing protocol to form a coding-aware multipath routing protocol, so that network coding processing can be distributed throughout the network, rather than only existing on the source node or a path in the network. At the same time, according to the coding conditions, each node only needs to obtain part of the network topology information to perform network coding operations, which makes network coding independent of the network topology structure and improves the adaptability of network coding.
[0036] It can be understood that the beneficial effects of the second aspect mentioned above can be found in the relevant description of the first aspect mentioned above, and will not be repeated here.
[0037] In summary, the present invention applies network coding to the data shuffling process, which can not only compress the transmission volume of data packets, reduce communication overhead, and save network bandwidth resources, but also improve the reliability of data transmission, and the security of data transmission is also guaranteed, thereby ensuring high-quality QoS services.
[0038] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 It is a schematic diagram of the process of the present invention;
[0040] Figure 2 It is the execution process of Map function and Reduce function;
[0041] Figure 3 It is the execution process of MapReduce;
[0042] Figure 4 It is the cumulative distribution function graph under normal distribution;
[0043] Figure 5 It is the cumulative distribution function graph under exponential distribution;
[0044] Figure 6 It is the cumulative distribution function graph under exponential distribution;
[0045] Figure 7 A schematic diagram of a computer device provided by an embodiment of the present invention;
[0046] Figure 8 The present invention is a block diagram of an electronic device provided according to an embodiment. DETAILED DESCRIPTION
[0047] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0048] In the description of the present invention, it should be understood that the terms “include” and “comprises” indicate the presence of described features, wholes, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or collections thereof.
[0049] It should also be understood that the terms used in the present specification are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the present specification and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include plural forms.
[0050] It should be further understood that the term "and / or" used in the present specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, and includes these combinations. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in the present invention generally indicates that the associated objects are in an "or" relationship.
[0051] It should be understood that, although the terms first, second, third, etc. may be used to describe preset ranges, etc. in the embodiments of the present invention, these preset ranges should not be limited to these terms. These terms are only used to distinguish preset ranges from each other. For example, without departing from the scope of the embodiments of the present invention, the first preset range may also be referred to as the second preset range, and similarly, the second preset range may also be referred to as the first preset range.
[0052] The word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining" or "in response to detecting", depending on the context. Similarly, the phrases "if it is determined" or "if (stated condition or event) is detected" may be interpreted as "when it is determined" or "in response to determining" or "when detecting (stated condition or event)" or "in response to detecting (stated condition or event)", depending on the context.
[0053] Various structural schematic diagrams of the embodiments disclosed in the present invention are shown in the accompanying drawings. These figures are not drawn to scale, and some details are magnified and some details may be omitted for the purpose of clear expression. The shapes of various regions and layers shown in the figures and the relative sizes and positional relationships therebetween are only exemplary, and may deviate in practice due to manufacturing tolerances or technical limitations, and those skilled in the art may additionally design regions / layers with different shapes, sizes, and relative positions according to actual needs.
[0054] The present invention provides a data transmission optimization method based on network coding. Taking into account the situation that the byte lengths of the intermediate values of the Map are not equal, in order to reduce the data transmission volume and the complexity of the algorithm design in the Shuffle stage, the coding idea of the CDC strategy is continued to be used, but the output function allocation scheme is changed. The data calculation amount in the Map stage is represented by the computational overhead □, and the data volume transmitted in the Shuffle stage is represented by the communication overhead L. When the byte lengths of the intermediate values of each output function Map are not equal, we propose a grouping strategy based on the CDC coding strategy, and characterize the functional relationship between the communication overhead LGroupC and the computational overhead r. In addition, to solve the optimal value of the communication overhead, we are limited by the complexity of the functional relationship solution space, and we propose a simpler and more generalized greedy strategy.
[0055] The functional relationship between the communication overhead LGroupC and the computational overhead r is as follows:
[0056]
[0057] Among them, K is the number of distributed computing nodes, N is the number of input files, Q is the number of output functions, φ i is the i-th output function, i∈{1,2,3,…,Q}, ω i is the i-th input file, r is the computational cost, L is the communication cost, S i is the set of byte lengths of the intermediate values of the output function responsible for node i, l i For the set S i The sum of the byte lengths of the upper middle values.
[0058] Example 1
[0059] See also Figure 1 The present invention provides a data transmission optimization method based on network coding, comprising the following steps:
[0060] In a MapReduce cluster consisting of K distributed computing nodes (servers), N input files need to be calculated through any Q output functions to obtain output results, where And N≥K.
[0061] Suppose there are N input files ω1,ω2,…,ω N , for Q output functions φ1, φ2, …, φ Q Then the output result is φq=φq(ω1,ω2,…,ω N ), where q∈{1,2,…,Q}. From the execution process of MapReduce, we can see that the calculation result of the output function is decomposed into the calculation of intermediate values in the Map phase and the calculation output result of the merged intermediate values in the Reduce phase, that is, φq(ω1,ω2,…,ω N )=h q (g q,1 (ω1),…,g q,N (ω N )); where g q,1 (ω1),…,g q,N (ωN) is the Map intermediate value calculation function g q,1 ,…,g q,N The intermediate values calculated for the corresponding input files ω1, ω2, …, ωN are obtained respectively; in the Reduce stage, all the N intermediate values merged are passed through the Reduce function h q The output result is obtained by calculation. The execution process of Map function and Reduce function is as follows Figure 2 shown.
[0062] Map Function Responsible for calculating the input file ω n , which will input the file ω n Calculate Q Maps to calculate the intermediate value v 1,n =g 1,n (ω n ),…,v Q,n =g Q,n (ω n ), n∈{1,2,…,N}.
[0063] The Reduce function hq is responsible for counting the calculation results of the output function φq. First, it is necessary to obtain all the required N Map intermediate values v through Shuffle data exchange. q,1 ,…,v q,N , then calculate the output result u q =h q (v q,1 ,…,v q,N ), q∈{1,2,…,Q}.
[0064] Since the above calculations are completed by K distributed computing nodes (servers), the nodes are interconnected through a network that supports broadcast / multicast. Next, the calculation process of N input files and Q output functions is mapped to K computing nodes. The process can be divided into three stages: Map stage, Shuffle stage and Reduce stage.
[0065] Map phase:
[0066] Node k, k∈{1,…,K} computes the set of input files stored on itself For the input file set M k Each input file ω n , node k executes the Map function to obtain Q Map calculation intermediate values, that is, In order to obtain the complete output result value, each input file is stored on at least one node, that is,
[0067] Shuffle phase:
[0068] Assign nodes k, k∈{1,…,K} to be responsible for outputting the function set The output result. Without considering the cascade process, in order to reduce the network load, it is required that each output function is only responsible for one node, that is, and
[0069] In order to calculate the output function set W k The output result u of each output function φq in q , node k needs all N Map intermediate values v q,1 ,…,v q,N ; Some of the intermediate values can be known through the calculation of the Map stage itself, and the remaining intermediate values It needs to be transmitted from other nodes.
[0070] For each node k, the locally known intermediate values can be encoded to obtain X k , the encoding function is denoted as ψ k ,Right now Afterwards, node k broadcasts X k to all other nodes.
[0071] CDC strategy based on heterogeneous map intermediate values
[0072] The CDC strategy is mainly used to solve the communication bottleneck problem in the Shuffle phase. It is currently the best algorithm in the ideal case where all intermediate values have equal byte lengths, but the performance is not so significant when the intermediate values have unequal byte lengths. In this section, the CDC strategy process is introduced for the case where the intermediate value byte lengths of each output function are unequal but the intermediate values within the same output function are equal, and the functional relationship between the computational overhead and the communication overhead is characterized, so as to find the optimization point of the CDC process to improve the communication performance when the intermediate value byte lengths are unequal.
[0073] In order to simplify the model and facilitate analysis, assume that for the same output function φq,q∈{1,2,…,Q}, N Map intermediate values v q,1 ,…,v q,N The byte length is b q .
[0074] For the case where the intermediate value byte lengths are unequal, since the output function average distribution scheme cannot achieve the communication performance optimization effect when the intermediate value byte lengths are equal, it is necessary to change the output function average distribution scheme in the CDC strategy to make the communication performance in the Shuffle stage better.
[0075] In MapReduce, the input file placement scheme, intermediate value encoding scheme, and intermediate value decoding scheme are the same as the CDC strategy. The algorithm process with only a different output function allocation scheme is called a grouping strategy.
[0076] If the output function set W of node k is known k , and the corresponding communication overhead is calculated. For the distribution of output functions, there are K Q The optimization goal of the solution is to Q Find a solution that minimizes the communication overhead among the output function schemes of orders of magnitude.
[0077] Greedy Algorithm for Output Function Allocation
[0078] There are Q output functions and K distributed computing nodes. The intermediate values of the Q output functions are arranged in ascending order, namely b1, b2, …, b Q , that is, b1≤b2≤…≤b Q . And the output function is also matched in this order, b i is the i-th output function φ i The length in bytes of the intermediate value.
[0079] Suppose that the byte length of the intermediate value of the Map of Q=6 output functions is {b1,b2,b3,b4,b5,b6}={1,1,1,4,5,6}; these 6 output functions need to be assigned to K=3 distributed computing nodes. First, assign output function φ6 to node 3, output function φ5 to node 2, and output function φ4 to node 1. At this time, the total byte length of the intermediate value of the output functions assigned to nodes 1, 2, and 3 is 4, 5, and 6 respectively;
[0080] Therefore, for the output function φ3, the byte length of its intermediate value is 1, and it is assigned to node 1. At this time, the total byte lengths of the intermediate values of the output functions assigned to nodes 1, 2, and 3 are 5, 5, and 6 respectively.
[0081] For the output function φ2, the byte length of its intermediate value is 1, and it is assigned to node 1. At this time, the total byte lengths of the intermediate values of the output function assigned to nodes 1, 2, and 3 are 6, 5, and 6 respectively.
[0082] For the output function φ1, the byte length of its intermediate value is 1, and it is assigned to node 2. At this time, the total byte lengths of the intermediate values of the output functions assigned to nodes 1, 2, and 3 are 6, 6, and 6 respectively.
[0083] The output function allocation is completed. Node 1 is responsible for the output function set {φ4, φ3, φ2}, node 2 is responsible for the output function set {φ5, φ1}, and node 3 is responsible for the output function set {φ6}.
[0084] After using the greedy allocation strategy, the execution process of MapReduce is as follows Figure 3 As shown, the communication overhead in the Shuffle phase is only 6+6+6=18.
[0085] Reduce phase:
[0086] Node k, k∈{1,…,K} is responsible for outputting the function set After the Shuffle phase, ignoring packet loss, node k has received X1, X2, …, X K , combined with the locally known intermediate values, the output function set W can be decoded k Each output function φ q All N Map intermediate values v q,1 ,…,v q,N ; The decoding function is recorded as Right now
[0087] Finally, node k executes the Reduce function to calculate the output result u q =h q(v q,1 ,…,v q,N ),φq∈W k .
[0088] Calculate the cost r:
[0089] The ratio of the number of input files stored on all K compute nodes to the number of original input files N.
[0090]
[0091] in,
[0092] It can also be understood that each input file is stored on r nodes. Since each input file is stored on at most K nodes and at least one node, 1≤r≤K.
[0093] Communication overhead L:
[0094] The sum of the byte lengths of the data sent by all nodes in the Shuffle phase. In the Shuffle phase, the coded data X sent by node k, k∈{1,…,K} k If the sum of the byte lengths is but
[0095] For any given MapReduce process, that is, given the input file set M k , output function set W k , encoding function ψ k , decoding function As well as the computational overhead r, the communication overhead L can be theoretically calculated to form a computation-communication relationship pair (r, L).
[0096] The CDC strategy states that when the byte length of all Map intermediate values is l, the communication overhead value of the Shuffle phase is unique, and the computation-communication function relationship is:
[0097]
[0098] But obviously the CDC strategy does not take into account the situation where the byte lengths of the intermediate values in the Map are not equal. In order to simplify the model and facilitate analysis, it is assumed that for the same output function φ q ,q∈{1,2,…,Q},N Map intermediate values v q,1 ,…,v q,N The byte length of is bq. Therefore, our optimization goal is to design an encoding strategy for a certain computational cost L so that the communication cost L in the Shuffle phase is as small as possible when the byte lengths of the intermediate values of each output function are different but the intermediate values in the same output function are equal.
[0099] It will be appreciated by those skilled in the art that various aspects of the present invention may be implemented as systems, methods or program products. Therefore, various aspects of the present invention may be specifically implemented in the following forms, namely: complete hardware implementation, complete software implementation (including firmware, microcode, etc.), or a combination of hardware and software implementations, which may be collectively referred to herein as "circuits", "modules" or "platforms".
[0100] Example 2
[0101] A data transmission optimization system based on network coding is provided, which can be used to implement the above-mentioned data transmission optimization method based on network coding. Specifically, the data transmission optimization system based on network coding includes a Map module, a Shuffle module and a Reduce module.
[0102] Among them, the Map module, in a MapReduce cluster consisting of K distributed computing nodes, calculates the input file set M stored on itself through computing node k k , k∈{1,…,K}; computing node k executes the Map function to obtain Q Map calculation intermediate values, and each input file is stored on at least one node;
[0103] Shuffle module, which assigns computing node k to be responsible for outputting function set W k The output result of
[0104] Reduce module, computing node k is responsible for outputting the function set W k After the Shuffle phase, the computing node k receives the intermediate value code, and then decodes it to obtain the output function set W in combination with the locally known intermediate value. k Each output function φ q All N Maps calculate the intermediate value v q,1 ,…,v q,N ; Finally, computing node k executes the Reduce function to calculate the output result u q .
[0105] Example 3
[0106] The present invention provides a terminal device, which includes a processor and a memory, wherein the memory is used to store a computer program, the computer program includes program instructions, and the processor is used to execute the program instructions stored in the computer storage medium. The processor may be a central processing unit (CPU), or other general-purpose processors, graphics processors (GPU), tensor processors (TPU), digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc., which are the computing core and control core of the terminal, which are suitable for implementing one or more instructions, and are specifically suitable for loading and executing one or more instructions to implement corresponding method processes or corresponding functions; the processor described in the embodiment of the present invention can be used for the operation of the data transmission optimization method based on network coding, including:
[0107] Map phase: In a MapReduce cluster consisting of K distributed computing nodes, computing node k calculates the input file set M stored on itself. k , k∈{1,…,K}; computing node k executes the Map function to obtain Q Map calculation intermediate values, and each input file is stored on at least one node;
[0108] Shuffle phase: Assign computing node k to be responsible for outputting function set W k The output result of
[0109] Reduce phase: computing node k is responsible for outputting the function set W k After the Shuffle phase, the computing node k receives the intermediate value code, and then decodes it to obtain the output function set W in combination with the locally known intermediate value. k Each output function φ q All N Maps calculate the intermediate value v q,1 ,…,v q,N ; Finally, computing node k executes the Reduce function to calculate the output result u q .
[0110] See also Figure 7, the terminal device is a computer device, and the computer device 60 of this embodiment includes: a processor 61, a memory 62, and a computer program 63 stored in the memory 62 and executable on the processor 61. When the computer program 63 is executed by the processor 61, the data transmission optimization method based on network coding in the embodiment is implemented. To avoid repetition, it is not described one by one here. Alternatively, when the computer program 63 is executed by the processor 61, the functions of each model / unit in the data transmission optimization system based on network coding in the embodiment are implemented. To avoid repetition, it is not described one by one here.
[0111] The computer device 60 may be a computing device such as a desktop computer, a notebook, a PDA, or a cloud server. The computer device 60 may include, but is not limited to, a processor 61 and a memory 62. Those skilled in the art will appreciate that Figure 7 This is only an example of the computer device 60 and does not constitute a limitation of the computer device 60. It may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, the computer device may also include input and output devices, network access devices, buses, etc.
[0112] The processor 61 may be a central processing unit (CPU), or other general-purpose processors, graphics processing units (GPU), tensor processing units (TPU), digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.
[0113] The memory 62 may be an internal storage unit of the computer device 60, such as a hard disk or memory of the computer device 60. The memory 62 may also be an external storage device of the computer device 60, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc., equipped on the computer device 60.
[0114] Furthermore, the memory 62 may include both an internal storage unit of the computer device 60 and an external storage device. The memory 62 is used to store computer programs and other programs and data required by the computer device. The memory 62 may also be used to temporarily store data that has been output or is to be output.
[0115] See also Figure 8 The terminal device is an electronic device 600, which is in the form of a general-purpose computing device. The components of the electronic device may include, but are not limited to: at least one processing unit 610, at least one storage unit 620, a bus 630 connecting different platform components (including the storage unit 620 and the processing unit 610), a display unit 640, etc.
[0116] The storage unit stores program codes, which can be executed by the processing unit 610, so that the processing unit 610 performs the steps according to various exemplary embodiments of the present invention described in the above method section of this specification. For example, the processing unit 610 can perform the following steps: Figure 1 Follow the steps shown in .
[0117] The storage unit 620 may include a readable medium in the form of a volatile storage unit, such as a random access memory unit (RAM) 6201 and / or a cache memory unit 6202 , and may further include a read-only memory unit (ROM) 6203 .
[0118] The storage unit 620 may also include a program / utility 6204 having a set (at least one) of program modules 6205, such program modules 6205 including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment.
[0119] Bus 630 may represent one or more of several types of bus structures, including a memory unit bus or memory unit controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus architectures.
[0120] The electronic device 600 may also communicate with one or more external devices 700 (e.g., keyboards, pointing devices, Bluetooth devices, etc.), may also communicate with one or more devices that enable a user to interact with the electronic device 600, and / or communicate with any device that enables the electronic device 600 to communicate with one or more other computing devices (e.g., routers, modems). Such communication may be performed via an input / output interface 650. In addition, the electronic device 600 may also communicate with one or more networks (e.g., local area networks, wide area networks, and / or public networks, such as the Internet) via a network adapter 660. The network adapter 660 may communicate with other modules of the electronic device 600 via a bus 630. It should be understood that, although not shown in the figure, other hardware and / or software modules may be used in conjunction with the electronic device 600, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage platforms, etc.
[0121] Example 4
[0122] The present invention also provides a storage medium, specifically a computer-readable storage medium, which is a memory device in a terminal device for storing programs and data. It is understandable that the computer-readable storage medium here can include both the built-in storage medium in the terminal device and the extended storage medium supported by the terminal device, and can be any tangible medium containing or storing a program, which can be used by an instruction execution system, device or device or used in combination with it. The computer-readable storage medium provides a storage space, which stores the operating system of the terminal. In addition, one or more instructions suitable for being loaded and executed by a processor are also stored in the storage space, and these instructions can be one or more computer programs (including program codes). It should be noted that more specific examples of the computer-readable storage medium here include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory, an optical fiber, a portable compact disk read-only memory, an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0123] Computer readable storage media also include data signals propagated in baseband or as part of a carrier wave, which carry readable program code. Such propagated data signals can take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The readable storage medium can also be any readable medium other than a readable storage medium, which can send, propagate or transmit a program for use by or in conjunction with an instruction execution system, device or device. The program code contained on the readable storage medium can be transmitted using any appropriate medium, including but not limited to wireless, wired, optical cable, radio frequency, etc., or any suitable combination of the above.
[0124] Program code for performing the operations of the present invention may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, as a separate software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user computing device through any type of network, including a local area network or a wide area network, or may be connected to an external computing device (e.g., through the Internet using an Internet service provider).
[0125] The processor may load and execute one or more instructions stored in a computer-readable storage medium to implement the corresponding steps of the data transmission optimization method based on network coding in the above embodiment; the processor may load and execute the following steps of one or more instructions in the computer-readable storage medium:
[0126] Map phase: In a MapReduce cluster consisting of K distributed computing nodes, computing node k calculates the input file set M stored on itself. k , k∈{1,…,K}; computing node k executes the Map function to obtain Q Map calculation intermediate values, and each input file is stored on at least one node;
[0127] Shuffle phase: Assign computing node k to be responsible for outputting function set W k The output result of
[0128] Reduce phase: computing node k is responsible for outputting the function set W k After the Shuffle phase, the computing node k receives the intermediate value code, and then decodes it to obtain the output function set W in combination with the locally known intermediate value. kEach output function φ q All N Maps calculate the intermediate value v q,1 ,…,v q,N ; Finally, computing node k executes the Reduce function to calculate the output result u q .
[0129] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. The components of the embodiments of the present invention described and shown in the drawings here can usually be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0130] Communication cost L based on greedy grouping strategy GroupC The average communication cost L with the CDC strategy CDC And the minimum communication overhead L MinCDC Make a comparison.
[0131] Consider a MapReduce scenario with K=10, Q=40, and r=2, where the byte lengths of the intermediate values of the Q=40 output functions satisfy normal distribution, exponential distribution, and uniform distribution, respectively. Use MATLAB to write a program, repeat 10,000 times of numerical analysis, and obtain the cumulative distribution function graphs of the communication overhead and performance improvement percentage under the three distributions.
[0132] Depend on Figure 4 It can be seen that under normal distribution, the communication overhead of the greedy grouping strategy is the smallest, and its average communication overhead can be increased by 10% to 25% compared with the CDC strategy, and its minimum communication overhead can be increased by 3% to 10% compared with the CDC strategy.
[0133] Depend on Figure 5 It can be seen that under the exponential distribution, the greedy grouping strategy has the smallest communication overhead, and its average communication overhead can be increased by 20% to 35% compared with the CDC strategy, and its minimum communication overhead can be increased by 5% to 18% compared with the CDC strategy.
[0134] Depend on Figure 6 It can be seen that under uniform distribution, the communication overhead of the greedy grouping strategy is the smallest, and its average communication overhead can be increased by 14% to 30% compared with the CDC strategy, and its minimum communication overhead can be increased by 3% to 10% compared with the CDC strategy.
[0135] In summary, the greedy grouping strategy has the smallest communication overhead, and its average communication overhead can be increased by 10% to 35% compared with the CDC strategy, and its minimum communication overhead can be increased by 3% to 18% compared with the CDC strategy.
[0136] In summary, the data transmission optimization method and system based on network coding of the present invention have the following advantages:
[0137] Network coding can improve the fault tolerance and robustness of data shuffling. When a random network coding scheme is used, even if some nodes or links in the network fail, the destination node can still recover the original data through data decoding as long as it receives a certain amount of encoded data; without the need for complex encryption algorithms, network coding improves the security of the network to a certain extent.
[0138] Network coding can be used to reduce the communication overhead of data shuffling. To address the communication bottleneck problem, the excess storage or computing power on the computing nodes can be used to store redundant data or perform redundant calculations so that the computing nodes have part of the data of other nodes, and then encode their own data and redundant data, and broadcast the encoded data to other nodes. After receiving the encoded data, other nodes can decode the required data in combination with local data. In this way, a small additional storage or computing cost can be exchanged for a significant reduction in communication load, thereby solving the communication bottleneck problem to a certain extent.
[0139] In the MapReduce cluster system, based on the different byte lengths of the intermediate values of the output functions, an output function allocation scheme is proposed, and the amount of data grouping in the data shuffle is greatly compressed by using network coding. The byte lengths of the intermediate values of Map are not necessarily equal in MapReduce applications, such as query systems and data processing. For the case where the byte lengths of the intermediate values of each output function are unequal but the intermediate values in the same output function are equal, the functional relationship between the computational overhead and the communication overhead is characterized. In order to achieve the near-optimal communication overhead of data shuffle, a greedy strategy is designed for the allocation of output functions to improve the communication performance under the condition of unequal byte lengths of intermediate values.
[0140] The technicians in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In practical applications, the above-mentioned function allocation can be completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated in a processing unit, or each unit can exist physically separately, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, which will not be repeated here.
[0141] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0142] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed in the present invention can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.
[0143] In the embodiments provided by the present invention, it should be understood that the disclosed devices / terminals and methods can be implemented in other ways. For example, the device / terminal embodiments described above are only schematic. For example, the division of the modules or units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0144] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0145] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0146] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present invention implements all or part of the processes in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and the computer program can implement the steps of the above-mentioned various method embodiments when executed by the processor. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, U disk, mobile hard disk, disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium, etc. It should be noted that the content contained in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electric carrier signals and telecommunication signals.
[0147] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0148] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0149] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0150] The above contents are only for explaining the technical idea of the present invention and cannot be used to limit the protection scope of the present invention. Any changes made on the basis of the technical solution in accordance with the technical idea proposed by the present invention shall fall within the protection scope of the claims of the present invention.
Claims
1. A data transmission optimization method based on network coding, characterized in that: The following steps are involved: Map phase: In a MapReduce cluster consisting of K distributed computing nodes, computing node k calculates the input file set Mk stored on itself, k∈{1,…,K}; computing node k executes the Map function to obtain Q Map calculation intermediate values, and each input file is stored on at least one node; Shuffle phase: Assign computing node k to be responsible for outputting function set W k The output result of Reduce phase: computing node k is responsible for outputting the function set W k After the Shuffle phase, the computing node k receives the intermediate value code, and then decodes it to obtain each output function φ in the output function set Wk in combination with the locally known intermediate value. q All N Maps calculate the intermediate values vq,1,…,vq,N; finally, computing node k executes the Reduce function to calculate the output result u q .
2. The data transmission optimization method based on network coding according to claim 1, characterized in that: The input file set Mk has N input files ω1, ω2, …, ωN, corresponding to Q output functions φ1, φ2, …, φQ, and the output result is uq = φq(ω1, ω2, …, ωN), q∈{1,2, …,Q}, and the output result u q Decompose it into the Map phase to calculate the intermediate values and the Reduce phase to merge and calculate the intermediate values and output the results; The Map phase calculates the intermediate values gq,1,…,gq,N respectively for the corresponding input files ω1,ω2,…,ωN. In the Reduce phase, all N intermediate values merged are processed by the Reduce function h. q Calculate the output result.
3. The data transmission optimization method based on network coding according to claim 2, characterized in that: The Map function calculates the input file ωn and converts the input file ω n Q Map calculation intermediate values are obtained by calculation.
4. The data transmission optimization method based on network coding according to claim 2, characterized in that: Reduce function h q Responsible for statistical output function φ Q To calculate the result, we first need to obtain all the N required Map calculation intermediate values v through Shuffle data exchange. q,1 ,…,v q,N , then calculate the output result u q =h q (v q,1 ,…,v q,N ), q∈{1,2,…,Q}, v q,1 ,…,v q,N Calculate the intermediate value for Map, h q is the reduce function.
5. The data transmission optimization method based on network coding according to claim 1, characterized in that: Shuffle phase: Node k needs all N Map intermediate values v q,1 ,…,v q,N ; Some intermediate values are known through calculation in their own Map phase, and the remaining intermediate values Obtained through other node network transmission; For each node k, encode the locally known intermediate value of the Map calculation to obtain the intermediate value code X k , the encoding function is denoted as ψ k ; Afterwards, node k broadcasts X k to all other nodes.
6. The data transmission optimization method based on network coding according to claim 5, characterized in that: In the Shuffle phase, for any given MapReduce model, that is, given a set of input files M k , output function set W k , encoding function ψ k , decoding function And the computational overhead r, we get the communication overhead L, and form a computation-communication relationship pair (r, L).
7. The data transmission optimization method based on network coding according to claim 6, characterized in that: The computation-communication function relationship is: Among them, l is the byte length of all map intermediate values, N is the number of input files, Q is the number of output functions, K is the number of distributed computing nodes, and r is the computing overhead.
8. The data transmission optimization method based on network coding according to claim 7, characterized in that: The computational cost r is: Among them, M k is the set of input files that node k is responsible for.
9. The data transmission optimization method based on network coding according to claim 6, characterized in that: The communication overhead L is: L=l1+l2+…+l k Among them, l k is node k, k∈{1,…,K}, is the coded data X to be sent k The total length in bytes.
10. A data transmission optimization system based on network coding, characterized in that: include: The Map module calculates the input file set M stored on itself through computing node k in a MapReduce cluster consisting of K distributed computing nodes. k , k∈{1,…,K}; computing node k executes the Map function to obtain Q Map calculation intermediate values, and each input file is stored on at least one node; Shuffle module, which assigns computing node k to be responsible for outputting function set W k The output result of Reduce module, computing node k is responsible for outputting the function set W k After the Shuffle phase, the computing node k receives the intermediate value code, and then decodes it to obtain the output function set W in combination with the locally known intermediate value. k Each output function φ q All N Maps calculate the intermediate value v q,1 ,…,v q,N ; Finally, computing node k executes the Reduce function to calculate the output result u q .
Citation Information
Patent Citations
Coding MapReduce method for median length isomerism
CN111490795A
Code distributed calculation method based on Partition structure
CN112769522A
Code distributed computing method based on MapReduce framework
CN113434299A
Faucet with blow dryer
KR1020240002904A