A distributed collaborative processing method for weak networks and full dynamics
Through hierarchical architecture and a variety of innovative technologies, the existing distributed computing framework is solved inefficient in weak networks and fully dynamic scenarios, and efficient and adaptive distributed collaborative computing is achieved.
Patent Information
- Application Number
- CN202411563320.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-05
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2044-11-05
AI Technical Summary
The existing distributed computing framework is difficult to efficiently collaboratively calculate in weak networks and fully dynamic scenarios, and its performance is degraded or it cannot work properly.
It adopts a hierarchical architecture of perception layer, support layer, scheduling layer and processing layer, combining event-triggered incremental data synchronization, inter-node data propagation of epidemic and gossip algorithms, task scheduling based on greed and dynamic programming, and operator execution mechanism of lightweight RPC.
It improves the distributed collaborative computing efficiency under weak networks and full dynamic conditions, realizes adaptive and insensitive collaborative computing, and adapts to fully dynamic application scenarios such as Internet of Vehicles and UAV clusters.
Smart Images

Figure CN119383118B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of distributed computing, and in particular to a distributed collaborative processing method for weak networks and full dynamics. Background Art
[0002] With the rapid development of emerging technologies such as the Internet of Things, mobile Internet, and edge computing, distributed collaborative processing faces new application scenarios and technical challenges. In the fields of industrial Internet, Internet of Vehicles, mobile crowdsourcing, etc., a large number of intelligent devices need to collaborate efficiently in a weak network and fully dynamic environment to complete complex computing tasks. However, existing distributed computing frameworks such as Hadoop and Spark are mainly designed and optimized for high-performance stable network environments, and often fail to perform as expected in weak network and highly dynamic scenarios.
[0003] Existing frameworks usually assume that the network bandwidth is high enough, the latency is low, and the node locations are relatively fixed. However, in actual weak network scenarios, due to unstable wireless links and severe signal fading, the network conditions are often difficult to meet the ideal assumptions, resulting in a significant drop in the performance of existing frameworks, and in severe cases, they may even fail to work properly.
[0004] Most existing frameworks are optimized for specific network topologies and node configurations and lack dynamic adaptability. When nodes move, join or leave dynamically, existing frameworks often need to be manually reconfigured, and cannot achieve adaptive, seamless collaborative computing, making it difficult to adapt to fully dynamic application scenarios such as Internet of Vehicles and drone swarms. Summary of the invention
[0005] In response to the problem of low efficiency of distributed collaborative computing under weak network and fully dynamic conditions in the prior art, the present application provides a distributed collaborative processing method for weak network and fully dynamic conditions, which improves the efficiency of distributed collaborative computing under weak network and fully dynamic conditions through the layered architecture design of perception layer, support layer, scheduling layer and processing layer, combined with event-triggered incremental data synchronization, data propagation between nodes combining epidemic and gossip, task scheduling based on greedy and dynamic programming, and lightweight RPC operator execution mechanism.
[0006] The purpose of this application is achieved through the following technical solutions.
[0007] The present application provides a distributed collaborative processing method for weak networks and full dynamics, which is used in a distributed collaborative processing system, including: the distributed collaborative processing system includes a perception layer, a support layer, a scheduling layer and a processing layer; the perception layer obtains the network status data, computing power data and environmental data of the local node respectively; and obtains the network connection indicators of the local node and the surrounding nodes by sending detection packets, and the network connection indicators include bandwidth, delay, packet loss rate and jitter; and the collected network status data, computing power data, environmental data and network connection indicators are standardized to obtain standardized data, and the standardized data is reported to the support layer; the support layer receives the standardized data reported by the perception layer, and encodes and compresses the data through a transmission algorithm; the support layer transmits the compressed standardized data between each node to achieve data synchronization between nodes; the support layer sends the data to each node based on the same The compressed standardized data of the first step is used to construct a neighborhood data support table, which includes a local data support table for storing local node data and a neighborhood data support table for storing neighborhood node data; the scheduling layer calculates the node's metric indicators based on the neighborhood data support table of each node, and constructs a node scheduling vector table based on the metric indicators; the scheduling layer uses a greedy algorithm or a dynamic programming algorithm to generate a collaborative computing node list based on the user's computing tasks and the node scheduling vector table; the processing layer divides the user's computing tasks into multiple independently executed task units through task processing, and encapsulates each task unit into an operator with a unified input and output interface; the UDP-based lightweight remote procedure call LightRPC method is used to transfer the operator to the specified node in the collaborative computing node list for execution; when the operator is executed on the collaborative computing node, distributed task processing is performed through the parallel computing framework.
[0008] Among them, network status data refers to a set of indicator data that reflects the quality and performance of network connections, and is often used to evaluate and monitor the health of the network. Network status data usually includes indicators such as bandwidth, delay, packet loss rate, and jitter. Bandwidth: Indicates the maximum data transmission rate of a network link, usually in bits per second (bps). Delay: Also known as latency, it indicates the time required for a data packet to be transmitted from the source node to the destination node, usually in milliseconds (ms). Packet loss rate: Indicates the proportion of data packets lost during data transmission to the total number of data packets, usually expressed as a percentage. Jitter: Indicates the change in the delay for a data packet to reach the destination node, reflecting the stability of network transmission, usually in milliseconds (ms).
[0009] Computing capacity data refers to a set of indicator data that reflects the computing performance and resource capacity of a node, and is used to evaluate the node's ability to perform computing tasks. Computing capacity data usually includes indicators such as CPU frequency, memory capacity, storage capacity, and I / O speed. CPU frequency: Indicates the clock frequency of the node processor, reflects the computing speed of the node, and the unit is usually Hertz (Hz). Memory capacity: Indicates the size of the node's random access memory (RAM), reflects the richness of the node's memory resources, and the unit is usually bytes (B). Storage capacity: Indicates the size of the node's persistent storage device (such as a hard disk), reflects the richness of the node's storage resources, and the unit is usually bytes (B). I / O speed: Indicates the data transmission rate of the node when performing input and output operations, reflects the communication efficiency between the node and external devices, and the unit is usually bytes per second (B / s).
[0010] In distributed computing, environment data refers to the software environment that distributed computing relies on, such as operating systems, drivers, and dependent libraries. Specifically, operating system information: Operating system type: the name of the operating system used by the node, such as Windows, Linux, MacOS, etc. Operating system version: the specific version number of the operating system used by the node, such as Windows 10, Ubuntu 20.04, CentOS 7, etc. Kernel version: the version number of the operating system kernel, which reflects the underlying functions and features of the operating system. Software dependency information: Runtime environment: the runtime environment of the programming language installed on the node, such as Java Runtime Environment (JRE), Python interpreter, etc. Dependent library version: the version number of the key dependent library installed on the node, such as database driver, communication middleware, parallel computing framework, etc. System library version: the version number of the system-level library installed on the node, such as glibc, OpenSSL, etc. Hardware driver information: GPU driver version: if the node is equipped with a GPU acceleration device, you need to record the version number of the GPU driver. Network card driver version: the version number of the driver used by the node's network adapter. Other device driver versions: Version numbers of drivers used by other key hardware devices on the node (such as FPGA, coprocessor, etc.).
[0011] UDP is the abbreviation of User Datagram Protocol, which is a connectionless transport layer protocol. Compared with TCP, UDP has the following characteristics: Connectionless: The two parties in UDP communication do not need to establish a connection in advance, and the sending and receiving of data are independent. Unreliable: UDP does not guarantee the reliable transmission of data packets, and data packets may be lost, duplicated, or arrive out of order. Datagram-oriented: UDP encapsulates application layer data into datagrams, each of which is independent and contains information such as the destination address and port number. No congestion control: UDP does not have a built-in congestion control mechanism, so data packets may be lost when the network is congested.
[0012] High transmission efficiency: Since the UDP protocol is simple and does not require connection establishment and maintenance status, the transmission efficiency is high. It is suitable for scenarios with high real-time requirements but relatively low reliability requirements, such as video streaming and voice calls.
[0013] Furthermore, the data is encoded and compressed through the transmission algorithm, including: using a custom message format based on the UDP protocol to encapsulate and transmit standardized data. Considering that the TCP protocol is prone to packet loss and retransmission, connection interruption and other problems in a weak network environment, which affects the transmission efficiency, this solution chooses the more lightweight UDP protocol as the transport layer protocol. On this basis, a custom message format is designed, which contains metadata fields such as data type, data length, sequence number, priority, checksum, etc., to achieve reliable data transmission and parsing.
[0014] A variety of data compression algorithms are used to compress the encapsulated standardized data, including data dictionary, character replacement, dynamic encoding, etc. First, according to the characteristics and distribution patterns of the data, a global data dictionary is constructed to map high-frequency data into short codes to reduce the data volume. Secondly, the repeated characters or strings in the data are replaced with shorter symbols. Thirdly, dynamic encoding technology is used to dynamically adjust the encoding table according to the contextual information and statistical laws of the data to further improve the compression ratio. Finally, a general lossless data compression algorithm (such as LZ77, DEFLATE, etc.) can be used to minimize the amount of transmitted data while ensuring data integrity.
[0015] According to the network bandwidth and data priority, the compressed standardized data is divided into multiple data slices, and the priority queue mechanism is used to sort and schedule the data slices. Specifically, the network perception module of the support layer detects the current network bandwidth in real time, and sets different priorities based on the timeliness and importance of the data. In the data slice process, high-priority data is allocated to smaller slices to ensure that it can be transmitted as quickly as possible; low-priority data is allocated to larger slices to avoid excessive bandwidth resources. At the same time, the sender maintains a priority queue, sorts the slices according to their priority, and gives priority to transmitting data slices with high importance and strong timeliness, thereby increasing the arrival speed of key data.
[0016] Adopting an incremental transmission strategy, only the changed data fragments are transmitted to reduce data transmission redundancy. In specific implementation, each node maintains a local data record table to record the data version number and hash value of the last synchronization. In a new data synchronization cycle, the node first compares the local data record table to identify the changed data items, and then encodes, compresses and fragments the changed data items according to the above scheme, and finally selectively transmits the changed data fragments. This incremental transmission method can significantly reduce the transmission of duplicate data and save network bandwidth.
[0017] Furthermore, constructing a neighborhood data support table includes: initializing a local data support table. When each node starts, it first creates and initializes a local data support table to record various status data of the node itself. These data include: network status data (such as network delay, bandwidth, packet loss rate, etc.), computing power data (such as CPU model and main frequency, memory size, storage capacity, etc.), and environmental data (such as geographic location, power, etc.). These data are collected in real time through various sensors and monitoring tools on the node, and filled into the local data support table according to a certain data format.
[0018] Discover and record neighbor nodes. Nodes discover neighbor nodes in a variety of ways, including: Network adapter: The node's network adapter can discover other nodes in the same network segment through link layer protocols (such as ARP) or network layer protocols (such as ICMP); Broadcast and monitoring: The node can send a discovery message in a specific format to a predetermined broadcast address. After receiving the message, other nodes reply with their own node identification and address information; Central node: In some scenarios, a central node (such as a registration center) can be deployed, and other nodes register with the central node after startup, thereby indirectly discovering each other. The discovered neighbor node information is recorded in the neighborhood data support table, including basic information such as node ID, IP address, port number, etc.
[0019] Communicate with neighboring nodes, collect and update the status data of neighboring nodes. Nodes establish network connections (such as TCP connections) with neighboring nodes and regularly exchange each other's status data. Specifically, the node sends its own local data support table to the neighboring node, and receives their local data support table sent by the neighboring node. After receiving the data from the neighboring node, the node extracts the network status, computing power, environmental data and other fields, and fills them into the corresponding fields of its own neighborhood data support table, thereby obtaining detailed status information of the neighboring node.
[0020] Dynamically adjust the selection of neighbor nodes according to changes in network status. In a weak network environment, the connectivity between a node and its neighbor nodes may change at any time. Therefore, the node needs to continuously monitor the network quality with neighbor nodes and dynamically adjust the selection of neighbor nodes according to changes in network status. Specifically, the node periodically measures indicators such as RTT delay and packet loss rate with each neighbor node. When the network quality of a neighbor node drops below a certain threshold, the node removes it from the neighborhood data support table; conversely, if a new neighbor node with good network quality appears, it is added to the neighborhood data support table. Through this dynamic adjustment mechanism, nodes can adapt to frequent changes in network topology and always maintain a collaborative relationship with nodes with the best network conditions.
[0021] Update the data support table periodically or triggered by events. In order to ensure that the status data recorded in the data support table is synchronized with the actual status of the node, the data items in the table need to be updated regularly or when key events occur. For example, the node can set a fixed time interval (such as 5 seconds) to periodically re-collect and calculate various status data and update the local data support table. In addition, when the node's own network interface status, computing load, geographical location, etc. change significantly, an update of the data support table is also triggered. Similarly, when the node senses that a neighbor node has joined or left, or the network topology has changed, the neighborhood data support table is also updated immediately to ensure that the neighbor node status data recorded therein is accurate and valid.
[0022] Furthermore, the support layer uses an event-triggered update mechanism to update incremental data, including: an event-triggered incremental update mechanism. Compared with periodic full updates, event-triggered incremental updates can significantly reduce unnecessary data transmission and processing overhead. Specifically, an event listener is set inside each node to monitor various types of status data in the local data support table and the neighborhood data support table. When the listener detects the following situations, an update event is triggered: the node's own network status data (such as latency, bandwidth, etc.), computing power data (such as CPU utilization, memory usage, etc.) or environmental data (such as operating system, file system, driver and dependency library, etc., the remaining power can also be used as environmental data) changes, and the change exceeds a certain threshold (such as 10%); the neighboring node reports that its status data has changed, and the change exceeds the threshold.
[0023] Once an update event is triggered, the node extracts the changed data items and forms an incremental update package. If the change occurs in the local node, the incremental update package is directly applied to the data items of the corresponding node in the local data support table; if the change occurs in the neighboring node, the incremental update package is applied to the data items of the corresponding neighboring node in the neighborhood data support table. This incremental update method can avoid frequent full table synchronization and reduce communication and computing burdens.
[0024] Adaptive information dissemination mechanism. In order for nodes to perceive each other's state changes, incremental update packages need to be propagated between nodes. However, in a weak network environment, the cost of full network broadcast is very high, which can easily cause network congestion and data loss. Therefore, this application adopts an adaptive dissemination mechanism that combines epidemic and gossip to dynamically adjust the range and speed of information dissemination according to network conditions:
[0025] In a local area, an epidemic algorithm is used to propagate information. Specifically, when a node generates an incremental update package, it sends it to all directly connected neighboring nodes. The neighboring nodes that receive the incremental update package also immediately forward it to their own neighboring nodes, so that the incremental update package is quickly propagated in the local area and local consistency is achieved as soon as possible.
[0026] In the entire distributed network, the gossip algorithm is used to spread information. Unlike epidemic, the gossip algorithm does not require that messages be sent to all neighbor nodes every time, but randomly selects one or several neighbor nodes to send. The node that receives the message also continues to forward it with a certain probability, rather than forwarding it unconditionally. Although this propagation method is slow, it can effectively control the network load and avoid congestion. The node dynamically adjusts the forwarding probability and forwarding range of the gossip algorithm according to the current network conditions (such as average delay and packet loss rate), and makes a trade-off between propagation efficiency and network load. Through the collaboration of epidemic and gossip, incremental update packages can quickly reach consistency locally and eventually reach consistency globally, adapting to dynamic changes in weak network environments. At the same time, nodes can adjust propagation parameters according to real-time network conditions, reduce redundant propagation under poor network conditions, and speed up propagation when network conditions improve, thereby improving the efficiency and reliability of information synchronization.
[0027] Furthermore, the epidemic algorithm propagates information in a local area, including: in a distributed collaborative processing system, assuming that there are N nodes in total, each node maintains a neighbor node list and an update information queue. The neighbor node list records the ID, IP address, port number and other information of the current neighbor nodes of the corresponding node; the update information queue is used to temporarily store the update data items that need to be propagated to the neighbor nodes.
[0028] When the local data support table or neighborhood data support table of a node (such as node i) is updated (such as the network delay changes by more than 10%), an update event is triggered. Node i puts the changed data item (such as the new delay value) and the corresponding identification information (such as the name of the data item, version number, etc.) into the update information queue. The identification information is used to uniquely identify a data item so that other nodes can distinguish whether it is a new update.
[0029] Node i selects a neighbor node j from its neighbor node list through certain strategies (such as random selection, polling, etc.). In order to control the propagation range, the selected neighbor nodes are usually within one or two hops of node i and belong to the local area. Node i packages all data items in the update information queue, together with their identification information, into an update message and sends it to the selected neighbor node j through the network. The update message is transmitted using the UDP protocol to reduce overhead.
[0030] When node j receives the update message from node i, it first unpacks the message and extracts the data items and identification information. Then, node j compares each received data item with the corresponding data item in its local data support table and the neighborhood data support table to see if the identification information is consistent. If the received data item is a new update (identification information is inconsistent), it is placed in node j's own update information queue and continues to propagate to node j's neighboring nodes; if the received data item already exists (identification information is consistent), it is ignored and no longer propagated.
[0031] For new data items placed in the update information queue, node j applies them to the local data support table and the neighborhood data support table to update the corresponding status data. In this way, node j completes an incremental update, making the local state synchronized with node i. Node j repeats steps 3 to 6 for the data items in the update information queue, that is, randomly selects one of its neighbor nodes k, sends the update information to node k, and node k continues to propagate it. This process is performed recursively, allowing the update information to spread rapidly in the local area. Repeat the above propagation process until all nodes in the local area have received the update information and completed the update of the local data support table and the neighborhood data support table. At this point, the node states in the local area have reached consistency.
[0032] Furthermore, the gossip algorithm is used to propagate information in the entire distributed network, including: in the initial state of the gossip algorithm, each node i maintains a local view to record the information of other nodes obtained by node i through information exchange. At the initial moment, the local view of node i only contains the information of node i's own neighbor nodes, that is, the ID, IP address, port number, etc. of the one-hop or two-hop neighbor nodes obtained through the neighbor discovery mechanism. In each round of information exchange, node i first randomly selects a node j from its own local view as the object of this round of information exchange. Unlike the epidemic algorithm that directly selects neighbor nodes, the communication object of the gossip algorithm can be any node in the local view, so its propagation range can cover the entire network.
[0033] Node i determines the amount of node information to be exchanged with node j in this round based on the preset propagation probability. The propagation probability is a parameter in the interval (0, 1], which is used to control the communication overhead of the gossip algorithm. For example, if the propagation probability is 0.5, and there are 100 nodes in the local view of node i, then the information of 50 nodes will be randomly selected to be exchanged with node j in this round. Node i randomly selects the node information corresponding to the calculated amount of data from its local view, packages it into an information exchange message, and sends it to node j through the network. At the same time, node j also uses the same method to package the randomly selected node information in its local view and send it to node i. This process is bidirectional, and nodes i and j each send a message and each receive a message.
[0034] After completing this round of information exchange, nodes i and j will merge the new node information obtained from each other into their respective local views. If the received information contains nodes that do not yet exist in the local view, these new nodes will be added to the local view; if the received information contains updated information of existing nodes in the local view (such as changes in IP address, status, etc.), the information of the corresponding node will be updated. Through this update of the local view, each node gradually expands its cognitive scope of the entire network. All nodes repeat the execution in parallel, i.e., randomly select communication objects, calculate the amount of data, exchange part of the local view, and update their own local view. This process is recursively performed for multiple rounds, so that the node status information is diffused and updated throughout the entire network.
[0035] As information exchange proceeds, the local view of each node will converge to the global view of the entire network, that is, all nodes will eventually obtain complete network node information. When the preset convergence conditions are met (such as reaching the maximum number of rounds, the change in the local view is lower than the threshold, etc.), the information exchange process stops. At this point, all nodes in the distributed network have reached a global state consistency.
[0036] Furthermore, the scheduling layer calculates the node's metric index based on the neighborhood data support table of each node, and constructs a node scheduling vector table based on the metric index, including: the scheduling layer calculates the network status metric index Net between nodes based on the network status data of each node in the neighborhood data support table by the following formula:
[0037] Net = f(Bandwidth, Latency, PacketLoss, Option), where Bandwidth represents network bandwidth, Latency represents network delay, PacketLoss represents network packet loss rate, Option represents other candidate indicators that affect network status, and f is the network status evaluation function. The network status evaluation function is calculated based on the weights of different influencing factors, among which bandwidth, delay and packet loss rate are the main influencing factors, and others such as routing protocols and network jitter are secondary influencing factors (corresponding to option). The evaluation function f is a weighted function. Before weighted calculation, all factors need to be normalized and set between [0, 1]. The expression of function f can be:
[0038] f(B,L,P,O)=w b o(B)+w l o(L)+w p o(P)+w o o(O), where B is the network bandwidth, L is the network delay, P is the packet loss rate, and O is other optional indicators; w b ,w l ,w p ,w o Respectively represent the weights of the four indicator parameters. Satisfy: w b +w d +w p +w r =1; the o function is a normalization function. Normalization can map the values of each parameter to between [0, 1] to ensure that parameters of different dimensions can be directly added. The normalization method can be calculated according to the maximum tolerance method. For example, if the maximum tolerable packet loss rate is x and the actual packet loss rate is y, the normalization function o can be expressed as: The normalization method for other influencing factors is similar.
[0039] The scheduling layer calculates the computing power measurement index Compute of the node according to the computing power data of each node in the neighborhood data support table through the following formula: Compute = g (CPU, NPU, GPU, Memory, Option); where CPU, NPU, GPU and Memory represent the CPU performance, NPU performance, GPU performance and memory capacity of the node respectively, and g is the computing power evaluation function; the computing power evaluation function g is different from the network evaluation function. Different computing tasks have different requirements for the type of computing power (some are biased towards CPU, some are biased towards GPU, etc.), that is, different computing tasks have different weights on factors, so they cannot be processed in a unified weighted manner. Each influencing factor in the calculation function represents a dimension of computing performance. Different computing tasks can calculate some or all of the dimensions in multiple dimensions with different weights. Therefore, we need to generate a multi-dimensional computing power vector to independently evaluate the computing power performance of each dimension. Define the computing power vector C as: Where C is the comprehensive computing power evaluation vector. Each item C in the vector is an independent evaluation vector of each factor. If we do not consider the more detailed preference of the computing task for each factor (such as the CPU cache, etc.), it can be simplified to a simpler calculation model: Here, C is still the comprehensive computing power evaluation vector, and each item O in the vector is a scalar in the range of 0 to 1 calculated after normalization.
[0040] The scheduling layer calculates the node's environmental matching metric Node according to the environmental data of each node in the domain data support table through the following formula: Node = h (Hardware, OS, feature); where Hardware represents the hardware features of the node, OS represents the operating system features of the node, feature represents the environmental features of the node, and h is the environmental matching evaluation function. The environmental evaluation function h is similar to the calculation evaluation function g, and is also a multi-dimensional vector evaluation model. The major aspects include three dimensions: hardware, operating system, and features. Here, we refine one level and extract more specific factors from the three major dimensions of hardware, operating system, and features, including: Operating system type OS: Different operating systems (such as Linux, Windows, and macOS) have better support for certain tasks. This item evaluates the compatibility, reliability, and stability of the operating system in the target task. Kernel version K: The operating system kernel version has an important impact on the support of the underlying hardware, driver compatibility, and security, especially in systems such as Linux. File system FS: The file system type (such as ext4, NTFS, and XFS) has a greater impact on storage and I / O intensive tasks, so the choice of the file system is scored. Driver compatibility D: Driver compatibility and version are crucial to hardware performance, especially in GPU and NPU intensive tasks. Software dependencies and library versions L: Includes the library versions (such as CUDA, OpenCL, etc.) that the task depends on and other dependencies required by the task. Based on the above specific factors, the environmental evaluation function can be expressed as: Different computing tasks have different requirements for the above different factor dimensions. Here we define a task weight vector W E :
[0041] W E =[w os ,w k ,w fs ,w d ,w l ]; where w os 、w k 、w fs 、w d 、w l They respectively represent the weight coefficients of a computing task on the operating system, kernel version, file system, driver compatibility, and software dependency library version.
[0042] According to the above definition, for a specific task's environment evaluation function ht (where h represents the environment evaluation function and t represents the specific task), the environment vector E and the task weight vector W can be used E Calculated by the dot product: ht = W E E = w os o(os)+wk o(k)+w fs o(fs)+w d o(d)+w l o(l); where o is a normalization function, and o(os), o(k), o(fs), o(d), and o(l) are the normalized values of the corresponding factors. When inferring the environment evaluation function, the computing task is considered. In order to be compatible with the different environmental requirements of different computing tasks, the environment evaluation vector and the task weight vector need to be calculated simultaneously. This evaluation model uses a multi-dimensional, weighted scoring method to enable different computing tasks to flexibly select the best operating platform based on the system environment.
[0043] The scheduling layer combines the network status metric, computing power metric, and environment matching metric of each node into a multidimensional metric vector M i , as the comprehensive performance representation of node i, M i Expressed as: M i =S(Net i ,Compute i ,Node i )=[v1,v2,v3,......,vn]; where Net i ,Compute i ,Node i They represent the network performance metric, computing performance metric, and environmental metric of node i respectively; M i is the metric vector of node i, expressed as [v1,v2,v3,......,vn], where [v1,v2,v3,......,vn] is M i Component values in each dimension; The scheduling layer's measurement vector M for all nodes i Normalization is performed to generate a node scheduling vector table, which contains the number of each node, the node type, the relationship with other nodes, and the normalized metric vector of the corresponding node.
[0044] Furthermore, the scheduling layer generates a list of collaborative computing nodes using a greedy algorithm or a dynamic programming algorithm based on the computing tasks and node scheduling vector table of the user, including: the scheduling layer generates a task measurement vector T according to the computing task requirements through the following formula: T = S (T) = [t1, t2, t3, ..., tn]; wherein t1 represents the computing task's requirements for the network, t2 represents the computing task's requirements for the node's computing power, t3 represents the computing task's requirements for the node's environment, S (T) is the task requirement evaluation function, T is the generated task measurement vector, and T is expressed as [t1, t2, t3]; network demand measurement t1, the computing task's requirements for the network are mainly composed of bandwidth, delay and packet loss rate and some other factors (Option), and different computing tasks have different requirements for each factor. Combined with the measurement function of the network indicators mentioned above, the network demand function t1 can be expressed as: Among them B req , D req , P req , O req They represent the requirements of the computing task for bandwidth, delay, packet loss rate and other factors respectively, and W represents the weight coefficient of each item.
[0045] Computational demand metric t2, the computing demand of computing tasks is mainly composed of four factors: CPU, GPU, NPU, and memory size. Different from the network, each computing factor itself is multi-dimensional, and the demand of computing tasks for each computing factor is also multi-dimensional. For example, some tasks prefer the CPU's main frequency, while others prefer the number of cores. Therefore, t2 can be expressed as: Where Ccpu req Indicates the demand of computing tasks for the number of CPU cores, cache, main frequency, etc., and is quantified by a comprehensive computing power demand indicator C. It is reflected in the structure as a regular array, and each item in the array represents a factor of the CPU. req 、Cnpu req 、Cmem req They represent the comprehensive computing power requirements of GPU, NPU, and MEM respectively. W represents the weight coefficient of each item.
[0046] Environmental requirement metric t3. In the previous environmental evaluation function, in order to calculate the environmental evaluation function under a specific task, the calculation task weight vector W is defined E , combined with the impact factor and weight coefficient, t3 can be expressed as: Among them, OS req , K req , FS req , D req , L reqRespectively represent the demand for operating system, kernel version, file system, driver and dependent library, and W represents the weight system of each demand. The demand measurement of kernel version, file system, startup and dependent library here is different from the numerical measurement of network and computing. These indicators are qualitative measurements and can be judged by whether they are satisfied. If satisfied, it is set to 1, and if not satisfied, it is set to 0. The sum of each weight coefficient should be 1.
[0047] The scheduling layer uses a vector matching algorithm to calculate the task metric vector T and the metric vector M of each node in the node scheduling vector table through the following formula: i The matching degree d(M i ,T):
[0048] Among them, M i1 ,M i2 ,......,M in Represents the metric vector M of node i i The dimensional components of T 1 ,T 2 ,......,T n Represents the components of each dimension of the task metric vector T; the scheduling layer calculates the matching degree of all nodes and obtains the node matching degree list L; the scheduling layer divides the computing tasks into small batch tasks and large batch tasks according to the requirements of the computing tasks on the computing scale; for small batch tasks, the scheduling layer adopts a greedy algorithm, and selects nodes with matching degrees higher than the threshold according to the node matching degree list L, and adds them to the collaborative computing node list C; for large batch tasks, the scheduling layer adopts a dynamic programming algorithm to obtain the optimal collaborative computing node list C*; the collaborative computing list C or C* is output as the task scheduling result.
[0049] Furthermore, for small batch tasks, the scheduling layer adopts a greedy algorithm to select nodes with matching degrees higher than the threshold according to the node matching degree list L and add them to the collaborative computing node list C, including: at the beginning of the algorithm, a matching degree threshold a is first set to screen nodes with sufficiently high matching degrees. The value of threshold a can be dynamically adjusted according to factors such as the type and complexity of the task and the demand for computing resources to balance the efficiency and cost of collaborative computing. For example, for computing-intensive tasks, a can be set higher to select nodes with stronger computing power; for I / O-intensive tasks, a can be set lower to select more nodes for parallel processing. In order to store nodes with matching degrees higher than the threshold, the algorithm maintains a collaborative computing node list C. At the beginning of the algorithm, list C is initialized to be empty. The scheduling layer maintains a node matching degree list L, which records the measurement vectors of all available nodes in the current distributed network. The algorithm traverses each node i in the list L and calculates its matching degree with the task.
[0050] For each node i in the node matching list L, through the function d(M i ,T) Calculate the metric vector M of node i i The degree of match between the task's measurement vector T. The measurement vector can include multiple dimensions of the node's computing power, storage capacity, network bandwidth, load status, etc. The matching function d can be calculated using a variety of methods, such as cosine similarity, Euclidean distance, etc., to measure the similarity or proximity between two vectors. For example, when using cosine similarity, the matching calculation formula is: Among them, ||M i || and ||T|| respectively represent the modulus of the two vectors. The value range of the matching degree is [0, 1]. The larger the value, the higher the matching degree between node i and the task.
[0051] For each node i, its matching degree d(M i ,T) is compared with the threshold a. If the matching degree is greater than the threshold, that is, d(M i ,T)>a, then node i is added to the collaborative computing node list C. This process embodies the idea of the greedy strategy, that is, each time the local optimal node (the one with the highest matching degree) is selected, hoping to approach the global optimal solution through the local optimal selection.
[0052] After traversing all nodes in the node matching list L, the collaborative computing node list C already contains all nodes with matching degrees higher than the threshold a. The algorithm returns list C as the output result to the scheduling layer. The scheduling layer can distribute small batches of tasks to the selected nodes based on the node information in list C to start the collaborative computing process.
[0053] Furthermore, for large-scale tasks, the scheduling layer uses a dynamic programming algorithm to obtain the optimal collaborative computing node list C*, including: quantifying the computing scale requirements of large-scale tasks into the required total computing power Σ(T). Where T is the task measurement vector, which contains the demand description for various node capability indicators. The calculation formula of Σ(T) is: Σ(T) = t2, where t2 is the component in the task measurement vector T that represents the computing task's demand for node computing power. Through this formula, the computing scale of the task can be mapped to a scalar value, which is convenient for subsequent optimization solutions.
[0054] In order to measure the matching degree between candidate nodes and tasks, the algorithm defines the node matching function d(C i ,T). For the i-th node C in the candidate node list C i , and its matching degree with the task metric vector T is calculated by the following formula: in, Candidate node C i The measurement vector of , T is the task measurement vector, and T j Represent candidate nodes C i The jth component of the metric vector of the task and the metric vector of the task. This formula uses the cosine similarity calculation method. The larger the cosine value of the angle between the metric vectors, the higher the matching degree between the node and the task.
[0055] To calculate the sum of the computing power of the candidate node list C, ∑C i , the algorithm defines a function to sum the node computing capabilities. For all nodes in the candidate node list C, the sum of their computing capabilities is calculated using the following formula: in, Candidate node C i The metric vector In the formula, the computing power of the candidate node list C can be accumulated and summed as one of the constraints of the optimization problem.
[0056] Based on the node matching list L, a dynamic programming algorithm is used to solve the following optimization problem to obtain the optimal collaborative computing node list C*: maximize:∑(d(C i ,T)); Constraints:∑(C i )≥∑(T);|C|=min(C); where C is the candidate collaborative computing node list, and |C| represents the number of nodes included in the candidate node list C. The optimization goal is to maximize the sum of the matching degrees of all nodes in the candidate node list C and the task, and the constraint condition is the sum of the computing power of the candidate node list C∑(C i ) is not less than the computing power ∑(T) required by the task, while minimizing the number of nodes |C|.
[0057] After constructing the mathematical model of the optimization problem, the dynamic programming algorithm is used to solve it. The core idea of dynamic programming is to decompose the original problem into several sub-problems, and gradually approach the optimal solution of the original problem by solving the sub-problems and recording the intermediate results. In this scheme, the problem can be divided into multiple stages, each stage corresponding to the selection of a candidate node. By recursively calculating the optimal state value of each stage, the global optimal solution is finally obtained, that is, the optimal collaborative computing node list C* with the maximum sum of node matching degrees and the minimum number of nodes under the constraints. The optimal collaborative computing node list C obtained by the dynamic programming algorithm is output as the task scheduling result of large-scale tasks. The scheduling layer can decompose large-scale tasks into multiple sub-tasks according to the node information in list C, and distribute them to the corresponding nodes for collaborative computing, thereby realizing efficient and scalable task processing.
[0058] Furthermore, a UDP-based lightweight remote procedure call LightRPC method is used to transfer the operator to the specified node in the collaborative computing node list for execution, including: splitting large-scale computing tasks into multiple independent computing units, each computing unit focusing on a sub-problem or processing stage of the task. Then, each computing unit is encapsulated as a standardized operator. The operator defines the input, output, processing logic and interface specifications of the computing unit, and has the characteristics of independence, reusability and composability. Through operator encapsulation, complex computing tasks can be modularized to improve the parallelism and flexibility of the task.
[0059] In order to transfer the encapsulated operators to the collaborative computing nodes for execution, this solution adopts the LightRPC method based on the UDP protocol. LightRPC is an optimized RPC communication mechanism with the advantages of high transmission efficiency, low resource overhead, and strong real-time performance. It is particularly suitable for distributed collaborative computing scenarios.
[0060] The original task node sends an execution request to the target collaborative node through LightRPC. The request content includes the operator to be executed, the initial state of the operator, and the expected return result format. LightRPC serializes the request message and sends it to the target collaborative node in the form of a UDP datagram. Compared with the TCP protocol, the UDP protocol does not require a connection to be established, has a smaller transmission overhead, and can support a higher number of concurrent requests.
[0061] After receiving the UDP datagram sent by LightRPC, the target collaborative node deserializes it to obtain the original execution request. According to the operator definition and interface standard contained in the request, the collaborative node loads the operator and initializes the operator state, and then starts the execution of the operator. During the execution of the operator, the local resources of the collaborative node, such as CPU, memory, storage, etc., can be accessed to complete a sub-problem or processing stage of the computing task.
[0062] After the collaborative node completes the execution of the operator, it encapsulates the execution result in the format specified in the request, serializes the result message through LightRPC, and returns it to the original task node in the form of a UDP datagram. Thanks to the simplicity and efficiency of the UDP protocol, the execution results can be quickly transmitted to the original node, reducing the communication delay of distributed tasks.
[0063] Since the UDP protocol is connectionless and does not guarantee reliable transmission of data packets, LightRPC has designed a confirmation response and retransmission mechanism to ensure the reliability of request and result messages. When the original node sends a request, if it does not receive a confirmation response from the collaborative node within a certain period of time, the request retransmission is triggered; similarly, when the collaborative node returns the execution result, if it does not receive a confirmation response from the original node within a certain period of time, the result retransmission is triggered. Through the confirmation response mechanism, LightRPC can detect and resend lost messages, improving the reliability of communication.
[0064] The original task node summarizes and post-processes the calculation results of the subtasks according to the execution results returned by the collaborative computing node. Since the task is split into multiple operators for parallel execution, the original node needs to assemble, merge and convert the execution results of each operator according to the logical order of the task, and finally obtain the complete task processing results. The result summary process can further apply optimization methods such as data compression, filtering, and aggregation to reduce the transmission and storage overhead of the result data.
[0065] Compared with the prior art, the advantages of this application are:
[0066] By building a comprehensive node status information support table at the perception layer, including multi-dimensional indicators such as network environment, computing environment, and environmental environment, and designing a customized weak connection transmission protocol and encoding compression algorithm at the support layer, the data transmission bottleneck in a weak network environment is solved. At the same time, a lightweight RPC mechanism based on UDP is used to distribute and execute operators to reduce communication overhead. These technologies ensure that the system can maintain stable and efficient collaborative computing capabilities even in poor network conditions.
[0067] The scheduling layer adopts an intelligent task scheduling strategy, comprehensively considers factors such as network conditions, node computing power, node status and task characteristics, and uses heuristic algorithms to dynamically generate the optimal combination of collaborative computing nodes, while meeting the task computing requirements and achieving full utilization of node computing power and load balancing. The processing layer supports flexible splitting of complex tasks into operators and distributing them to the optimal nodes for parallel processing, significantly improving the system's computing efficiency and throughput.
[0068] The support layer realizes the incremental update of node status through event triggering mechanism, and uses epidemic and gossip algorithms to intelligently adjust the range and rate of information dissemination; the scheduling layer perceives the dynamic changes of system topology and node status by updating the node scheduling vector table in real time. Based on these mechanisms, the system can adapt to dynamic changes such as node movement, frequent joining and leaving, without manual reconfiguration, and maintain the continuity and efficiency of the computing process.
[0069] Supports elastic scaling and load balancing of nodes. Thanks to the layered decoupled architecture design and full consideration of node heterogeneity, this system can achieve smooth node access and removal. When a new node joins, it only needs to collect its status information in the perception layer and quickly synchronize it to the support layer to join the scheduling candidate set without modifying the working mechanism of the scheduling and processing layer. When a node exits, the scheduling layer can dynamically update the scheduling vector table and reallocate tasks to other optimal nodes to avoid computing interruptions. At the same time, combined with the task splitting and migration mechanism, this system can dynamically migrate operators to nodes with lighter loads under local high load conditions, thereby achieving load balancing. These features make this system have excellent scalability and flexibility, and can adapt to distributed environments with dynamically changing scales. BRIEF DESCRIPTION OF THE DRAWINGS
[0070] The present application will be further described in the form of exemplary embodiments, which will be described in detail by the accompanying drawings. These embodiments are not restrictive, and in these embodiments, the same number represents the same structure, wherein:
[0071] Figure 1 is a schematic diagram of a completely peer-to-peer topology structure according to some embodiments of the present application;
[0072] Figure 2 is a schematic diagram of a central topology structure according to some embodiments of the present application;
[0073] Figure 3 is a schematic diagram of a tree topology structure according to some embodiments of the present application;
[0074] Figure 4 is a schematic diagram of a hybrid topology structure according to some embodiments of the present application;
[0075] Figure 5 is a schematic diagram of a distributed collaborative computing framework according to some embodiments of the present application;
[0076] Figure 6 is a schematic diagram of network perception according to some embodiments of the present application;
[0077] Figure 7 It is a neighborhood information support relationship diagram shown in some embodiments of the present application;
[0078] Figure 8 This is a schematic diagram of operator drift according to some embodiments of the present application. DETAILED DESCRIPTION
[0079] The method and system provided in the embodiments of the present application are described in detail below with reference to the accompanying drawings.
[0080] A distributed collaborative processing method for weak networks and full dynamics, used in a distributed collaborative processing system, comprising: the distributed collaborative processing system comprises a perception layer, a support layer, a scheduling layer and a processing layer; the perception layer respectively obtains network status data, computing power data and environmental data of the local node; and obtains network connection indicators between the local node and the surrounding nodes by sending detection packets, the network connection indicators including bandwidth, delay, packet loss rate and jitter; and standardizes the collected network status data, computing power data, environmental data and network connection indicators to obtain standardized data, and reports the standardized data to the support layer; the support layer receives the standardized data reported by the perception layer, and encodes and compresses the data through a transmission algorithm; the support layer transmits the compressed standardized data between each node to achieve data synchronization between nodes; the support layer is based on synchronization at each node The compressed standardized data is used to construct a neighborhood data support table, which includes a local data support table for storing local node data and a neighborhood data support table for storing neighborhood node data; the scheduling layer calculates the node's metrics based on the neighborhood data support table of each node, and constructs a node scheduling vector table based on the metrics; the scheduling layer uses a greedy algorithm or a dynamic programming algorithm to generate a list of collaborative computing nodes based on the user's computing tasks and the node scheduling vector table; the processing layer divides the user's computing tasks into multiple independently executed task units through task processing, and encapsulates each task unit into an operator with a unified input and output interface; the UDP-based lightweight remote procedure call LightRPC method is used to transfer the operator to the specified node in the collaborative computing node list for execution; when the operator is executed on the collaborative computing node, distributed task processing is performed through the parallel computing framework.
[0081] This application is applicable to the application scenarios of multi-node distributed collaborative computing under weak connection and full dynamic conditions. For the topological structure composed of multiple nodes in the system, there are mainly the following types: Figure 1 As shown in the figure, the complete equivalence equation means that nodes 1, 2, 3, and 4 are completely equal, with no center and no hierarchy. Figure 2 As shown in the figure, in the central type, node 1 is the central node, and nodes 2, 3, 4, and 5 are the child nodes of the central node. Ordinary nodes can establish relationships with the central node or with each other. Figure 3 As shown in the figure, the tree structure is equivalent to a multi-level structure of multiple central structures. The top node 1 is the root node, nodes 2 and 5 are child nodes of node 1, nodes 3 and 4 are child nodes of node 2, and nodes 6 and 7 are child nodes of node 5. Figure 4As shown, hybrid, full peer-to-peer, central, and tree types can be combined into other types of topological structures. For any node in any topological structure, its relationship is always composed of upper nodes, peer nodes, and lower nodes. For node 1, node 4 is its upper node, node 7 is its peer node, and nodes 2 and 3 are its lower child nodes. In this application, this description relationship is used for any node in the system.
[0082] The overall framework of this application consists of four layers, from bottom to top: the perception layer for perceiving the network and collecting information, the support layer for providing information synchronization and neighboring nodes, the scheduling layer based on the metric model and vector matching, and the execution processing layer using LightRPC and operator drift, such as Figure 5 As shown in the figure, the perception layer includes three modules: environment perception, network perception, and computing power perception. Environment perception mainly refers to the hardware environment such as battery, storage space, etc., operating system environment such as kernel version and file system, etc. In some cases, it can also include some business environment, such as the role and function of the node. These three aspects of perception correspond to the three aspects of node network, computing and environmental characteristics. Any node will obtain these three aspects of information from the local as the basis for the subsequent process.
[0083] The support layer is built on top of the perception layer. On the one hand, it provides weak connection transmission services. The core purpose of the support layer is to exchange and synchronize data through the network, and finally build a neighborhood information support table on the local node.
[0084] The scheduling layer is built on the supporting layer. According to the neighborhood information support table provided by the supporting layer, a mathematical model for measuring the key elements of distributed collaborative computing and a mathematical model for distributed task scheduling are established to form an indicator measurement model including a network measurement model, a computing power measurement model, and an environmental measurement model. The task scheduling model is then implemented based on the indicator measurement model.
[0085] The perception layer includes three modules: network perception, computing perception, and environmental perception. It is responsible for collecting information about the node's network status (type, bandwidth, latency, packet loss rate, etc.) and computing power (CPU, GPU, NPU, memory, etc.), as well as environmental information such as the node's hardware environment, operating system environment, and upper-layer application environment (such as computing volume, real-time requirements, etc.). This information is standardized and regularly reported to the support layer.
[0086] like Figure 6As shown in the figure, network perception, on the one hand, obtains the local network type and status, and secondly obtains the network connection quality with surrounding nodes by sending detection packets. Network quality includes multiple indicators such as bandwidth, delay, packet loss rate, jitter, etc. The type, environment and operation status information of the local network are obtained through system calls provided by the local OS and adapters or drivers for various networks. The detection packets are sent in various forms such as unicast packets, multicast packets, broadcast packets, etc., and key information such as timestamps and sequence numbers are carried in the detection packets. After receiving the detection packet reply, real-time calculation is performed to obtain network quality data.
[0087] Computational awareness includes obtaining local CPU, GPU, NPU, memory, and hard disk information. The level of detail of the information can be selected according to the situation. In general, it generally includes the following: Detailed information on the CPU, including processor utilization, core information, etc. The specific details are as follows: CPU usage: real-time utilization of the system as a whole and each core (usually expressed as a percentage) Number of cores: physical cores and logical cores (multiple logical cores can be used through hyperthreading technology). CPU frequency: the frequency at which the current CPU runs, whether it is in energy-saving mode or acceleration mode (Turbo Boost). Temperature monitoring: the temperature of each core, which can be used to detect whether the CPU is under high load or poor heat dissipation. Cache size: L1, L2, and L3 cache size, cache hit rate, etc., which will affect data access speed. Instruction set architecture: such as x86, ARM, and whether special instruction sets are supported (such as AVX, SSE, etc.).
[0088] Detailed information about the GPU (graphics processing unit) includes: GPU usage: overall usage and usage of each core. Video memory usage: total amount of video memory, used amount, free amount, and video memory overflow monitoring. GPU frequency: similar to the CPU, monitors the frequency and power consumption management of the GPU core. Temperature monitoring: the temperature of the GPU core. Video memory bandwidth: video memory read and write rate, the speed at which data is transferred between the video memory and the core. GPU model and architecture: obtain the specific model and architecture of the GPU (such as the number of CUDA cores, the number of Tensor cores, etc.). Multi-GPU configuration: in a multi-GPU system (multiple computing cards), the load balancing between the GPUs.
[0089] The detailed information of the NPU includes: NPU usage rate: the current workload and operating status of the NPU. Power consumption: the power consumption performance of the NPU when performing inference or training tasks. NPU temperature: monitor the temperature of the NPU to ensure stability during task execution. Data throughput rate: the speed at which the NPU processes neural network calculations (how many operations are performed per second, such as TOPS). NPU architecture and type: obtain the architecture and type information of the NPU, such as whether tensor operations, matrix multiplication optimization, etc. are supported.
[0090] Detailed memory information includes: Total memory capacity: the physical memory capacity of the system. Used memory and free memory: the memory usage of the current system, the memory allocated to the application and the remaining memory. Memory bandwidth: the read and write rate of the memory, the efficiency of data transmission. Memory page swap (swap) status: when the physical memory is insufficient, whether the swap memory is triggered, and the usage of the swap partition. Cache memory: cache memory used for the file system or other system processes. Memory type: the type of memory (such as DDR4, LPDDR5, etc.), number of memory channels, frequency, etc.
[0091] Environmental perception includes multiple aspects, which can be generally divided into three categories: hardware environment, operating system environment, and computing application environment. The hardware environment refers to other important environments besides the computing environment, such as battery power, hardware interface, node geographical location, etc. The operating system environment includes the operating system type, kernel version, system architecture, and file system, etc. The application environment includes the application computing type, such as AI computing, encryption and decryption, compression, data stream processing, etc. Different computing types have different computing requirements. The content of environmental perception can be scaled according to actual conditions. Some environmental items can be added when needed, and some can be reduced when not needed.
[0092] The support layer transmits information between nodes through efficient transmission implementation and synchronization algorithm, and forms a lightweight neighborhood information support table (Neighbor Computation Support Table, NCST). Each node maintains information about itself and neighbor nodes as well as relationship information. The local information support table (Local Computation Support Table, LCST) refers to the network, computing and environment information of the local node itself, and NCST is the network, computing and environment information of other nodes in the neighborhood. The design of NCST can significantly reduce the pressure of information transmission and information synchronization, especially in weak network and fully mobile environments. In the design of the neighborhood information support table, neighbors are judged based on network indicators. Under a relatively broad design, as long as the nodes are bidirectionally reachable on the network, they can be used as neighbor nodes; if in some high-demand situations, the requirements for network indicators can be higher. Neighbor nodes can be peer nodes, upper nodes or lower nodes. The neighbors here are only reflected in network capabilities. An example of the neighborhood information support table is shown in Table 1:
[0093] Table 1
[0094]
[0095] Each row in the table represents a neighbor node. The number is the node number, the type is the node type, and the relationship is the relationship with the local node. It is divided into three types: peer (same level), center (superior), and sub (subordinate). Then there are three key environments: network, computing, and environment. The three environments are composed of a one-dimensional array, and each item in the array represents an environment item. Each environment item has a custom structure and can include multiple sub-items. Based on the neighborhood information support table, a neighborhood information support relationship diagram that is easier to understand can be generated, such as Figure 7 As shown, Figure 7 It is a relationship diagram in a multi-hop network, so there will be multiple links. In this example diagram, No. 3 is the local node, and the other nodes are all nodes in the neighborhood. The network, computing, and environmental information of each node can be obtained, and the characteristics of the nodes are marked in the diagram based on the content of this information.
[0096] In the case of weak network and full dynamic, if the efficient synchronization of information between nodes is the most critical link, the core implementation of the support layer mainly includes the following three aspects: low-latency and high-compression transmission implementation, in order to cope with the low bandwidth and unstable network conditions of the network, a low-latency and high-compression transmission algorithm is designed. The core idea of the algorithm is to reduce the size and delay of the transmitted data as much as possible under the premise of ensuring the accuracy of the information, which is mainly reflected in the following aspects: message protocol: based on the more efficient UDP transmission protocol. Data compression: using efficient data compression mechanism and compression algorithm to compress the transmitted data to reduce the network load. Data compression includes data dictionary, character replacement, dynamic encoding and compression algorithm. Information fragmentation: splitting larger data packets into multiple small pieces to adapt to the transmission capacity of low-bandwidth networks. Differential transmission: only transmit data that has changed since the last synchronization, rather than the entire data set, to reduce unnecessary data transmission. Priority queue: assign different transmission priorities to different data to ensure that key information can be transmitted first. Congestion control: implements an adaptive congestion control mechanism to dynamically adjust the data transmission rate according to network conditions.
[0097] Neighborhood information table construction: Each node constructs a neighborhood information table (NCST) based on the data of the perception layer. The algorithm flow for constructing NCST is as follows: Initialization: Each node initializes its own local information support table (LCST), including its own network, computing and environment information. Neighbor discovery: Based on network adapters, broadcast and monitoring mechanisms, central nodes and other mechanisms, neighbor nodes are discovered and recorded. Based on network adapters, it means that for some special network devices, the interface provided by the network device can be directly called to obtain the list of nodes on the network. The broadcast and monitoring mechanism is a self-discovery mechanism that discovers the list of nodes on the network by sending broadcast or multicast packets. The central node mechanism refers to obtaining the list of online nodes from the central node. Information collection: Collect and update the network, computing and environment information of neighbors through communication with neighbor nodes. Periodic update: Update the information in NCST regularly or when changes are detected.
[0098] The information synchronization and update mechanism adopts an event-based update mechanism. When the local LCST changes by more than a threshold, the relevant information will be updated to the NCST. The specific steps are as follows: Event detection: The node continuously monitors its own LCST information and the change information received from neighboring nodes. Event response: Once a change is detected, the node immediately updates its own LCST and NCST and notifies the affected neighboring nodes. Information propagation: An information propagation algorithm combining epidemic and gossip is used to ensure that the change information can be quickly propagated to all relevant nodes without consuming too many network resources. Through the above measures, the support layer can efficiently synchronize information between nodes and ensure that each node can quickly obtain neighbor information, thereby providing accurate data support for upper-layer node scheduling and task allocation. This design not only improves the system's response speed and scalability, but also reduces network load, and is particularly suitable for operation under weak network and full mobility conditions.
[0099] The scheduling layer is built on the supporting layer. Based on the neighborhood information support table, it builds the measurement model of key elements of distributed computing and the node scheduling vector table. Then, based on the node scheduling vector table and the vector matching algorithm, it implements the scheduling model and selection algorithm of the task.
[0100] The node network is the most important parameter in distributed computing. Based on the three core indicators of bandwidth, latency and packet loss rate and an optional indicator, a mathematical model of the network performance between two nodes is established. The network status is evaluated by the following formula: Net = f (Bandwidth, Latency, Packet Loss, Option)
[0101] Among them, Bandwidth represents bandwidth, Latency represents delay, PacketLoss represents packet loss rate, and Option represents other influential indicators besides the three core indicators, such as network type, bearer protocol cluster, etc. The f function can be calculated based on the above four factors to obtain a metric Net representing the network status between two nodes.
[0102] The computing power of the node is another important parameter in distributed computing. For the measurement of the computing power of the node, the following mathematical model is established, and the formula is as follows:
[0103] Compute=g(CPU,NPU,GPU,Memory,Option)
[0104] The key information of the node, such as CPU, NPU, GPU, Memory, etc., are all involved in the calculation. Option is also optional. If there are other indicators that affect the computing power, such as hard disk IO, cache size, OS configuration, etc., they can be added to Option. The g function combines these parameters to give the node computing power measurement indicator Compute.
[0105] The node environment includes other aspects of the node information, including remaining power, storage size, OS type, software environment, etc. Hardware features and OS features will be reflected in the matching degree of computing tasks. Different computing tasks have different computing requirements, some for hardware and some for operating systems. In general, the measurement of node environment mainly includes hardware, software and other features: Node = h (Hardware, OS, feature)
[0106] Among them, Hardware represents hardware features, which may include battery power, storage space, etc.; OS represents operating system features, which may include OS type, kernel version, file system, software environment, etc.; feature represents other features, such as computing features: floating-point computing, real-time computing, AI inference computing, IO-intensive computing, etc.
[0107] Metric Vector Table,In distributed computing, the core of node scheduling lies in effectively evaluating the,performance and availability of each node.,Based on the three sets of scalars output by the node's network,computational,and environmental metrics, a multidimensional metric space is generated through,function S. The metric space of node i in the neighborhood node list can be,represented as a vector Mi:
[0108] M i =S(Net i ,Compute i ,Node i )=[v1,v2,v3,......,vn]
[0109] Where Neti represents the network performance metric of node i, which is a set of scalars. Computei represents the computing performance metric of node i, which is a set of scalars. Nodei represents the environmental metric of node i, which is a set of scalars. The calculated Mi is a vector, which can be expressed as [v1, v2, v3, ..., vn], indicating that the components of project Mi in each scheduling direction are v1, v2, v3, ..., vn respectively.
[0110] The vector Mi represents the comprehensive performance of node i in the neighborhood. By standardizing these metrics, we ensure that the metric vectors of different nodes can be compared in the same metric space. All the node sets N = {N1, N2, N3, ..., Nn} in the neighborhood node list are vectorized and standardized. The vectorized and standardized Mi set is the scheduling vector table V = {M1, M2, M3, ..., Mn} corresponding to the neighborhood information support table, as shown in the following table:
[0111] Node Number Node Type relation Scheduling vector Mi Kn001 c-rk3588 Peer [v1, v2, v3, ...] Kn002 c-ft Peer [v4,v5,v6,...] Kn003 c-intel Center [v7, v8, v9, ...]
[0112] Task scheduling, after completing the scheduling vector table of the neighborhood nodes, the scheduling of computing tasks can be performed based on the scheduling vector table. The scheduling of computing tasks requires a measurement model. The measurement vector T of the task is used to describe the task's requirements for the network, computing resources, and node environment. The measurement vector T of the task is generated by the task's demand function S(T): T = S(T) = [t1, t2, t3, ....., tn], where t1, t2, t3, ..., tn represent the task's requirements for the network, computing power, and node environment, respectively, and the components of the T vector in each demand direction are t1, t2, t3, ..., tn.
[0113] After obtaining the task's metric vector T, the matching degree between the task vector T and the node metric vector Mi is calculated to schedule the task. The matching degree is calculated using the Euclidean distance between the two vectors. The Euclidean distance is calculated as follows:
[0114] Among them, Mi1, Mi2, ..., Mi are the components of the metric vector of node i, and T1, T2, ..., Tn are the components of the task metric vector. The Euclidean distance reflects the actual gap and can reflect the absolute gap between the node and the task in terms of resource requirements. d(Mi, T) is the embodiment of the matching degree between task T and node i. After normalization, the closer to 1, the better the matching degree. After calculating the matching degree, the optimal list of collaborative computing nodes is selected according to the greedy algorithm and the dynamic programming algorithm. For better matching, the dynamic programming algorithm is used. If the matching situation is poor, the greedy algorithm is used. The core functions of the scheduling layer are completed based on the matching degree and algorithm.
[0115] The processing layer is built on the scheduling layer and is the executor of the scheduling results. The processing layer tasks the computing requirements of the upper-layer applications, and based on the scheduling results returned by the scheduling layer, performs tasks such as splitting, distributed execution, and result aggregation. The distributed execution system of tasks consists of two core parts: one is the lightweight remote procedure call technology LightRPC, and the other is the operator encapsulation of distributed tasks. After encapsulating the tasks as operators, efficient cross-node transmission and execution can be achieved through LightRPC.
[0116] In the process of executing distributed computing tasks, operators, execution status or final results are efficiently transmitted between nodes through LightRPC. LightRPC is a lightweight remote procedure access technology based on UDP, which aims to solve the problems of dynamic network connection and efficient data transmission in distributed computing environments. Unlike traditional TCP-based RPC, LightRPC is optimized for scenarios with weak connections, high network latency or unstable networks, making it more suitable for task execution and result return in collaborative computing environments. At the same time, since UDP is a connectionless protocol, LightRPC can better handle dynamic changes between nodes. For example, the addition or departure of nodes or fluctuations in network status will not cause disconnection or reconstruction of connections like TCP, which is suitable for weak connection environments. LightRPC also has a fault-tolerant mechanism, which, combined with a custom retransmission mechanism and confirmation mechanism, can make up for the transmission unreliability of the UDP protocol to a certain extent and ensure accurate return of results.
[0117] The basic workflow of LightRPC is as follows: Request initiation: When the node executes a distributed task, the original task node sends an execution request to the collaborative node through LightRPC. The request content may include the operator, the status of the current operator, and the expected return result. Message passing: LightRPC is based on UDP for message transmission, and the message is sent to the target collaborative node in the form of a datagram in the network. Result return: After completing the task, the collaborative computing node returns the execution result (intermediate state or final result) to the original task node through LightRPC. This return process also uses UDP transmission to reduce network overhead. Confirmation mechanism: LightRPC has designed a lightweight confirmation and retransmission mechanism. When a message is lost or no confirmation is received, the system will trigger a retransmission to ensure that the task execution result is returned accurately.
[0118] Operator drift, such as Figure 8 As described, according to the list of scheduling nodes returned by the scheduling layer, the computing tasks are split and encapsulated into operators. The purpose of encapsulation into operators is to standardize distributed computing tasks. All subsequent operations are based on operator standards. All operators can drift through LightRPC technology to achieve cross-node execution. The technology of LightRPC and operator drift is an abstraction and standard definition of distributed computing. Whether it is AI training and reasoning calculations, encryption and decryption calculations, or 3D rendering, engineering simulation and other computing tasks, as long as the computing tasks can be encapsulated according to the operator interfaces and standards, they can all run under a set of systems and standards, and can perform distributed collaborative computing in weakly connected networks and fully dynamic scenarios.
Claims
1. A distributed collaborative processing method for weak networks and full dynamics, used in a distributed collaborative processing system, characterized in that: include: The distributed collaborative processing system includes a perception layer, a support layer, a scheduling layer, and a processing layer; The perception layer obtains the network status data, computing power data and environmental data of the local node respectively; and obtains the network connection indicators of the local node and the surrounding nodes by sending detection packets. The network connection indicators include bandwidth, delay, packet loss rate and jitter; and standardizes the collected network status data, computing power data, environmental data and network connection indicators to obtain standardized data, and reports the standardized data to the support layer; The support layer receives the standardized data reported by the perception layer and encodes and compresses the data through the transmission algorithm; The support layer transmits the compressed standardized data between nodes to achieve data synchronization between nodes; The support layer constructs a neighborhood data support table on each node based on the synchronized compressed standardized data, and the neighborhood data support table includes a local data support table storing local node data and a neighborhood data support table storing neighborhood node data; The scheduling layer calculates the node's metrics based on the neighborhood data support table of each node, and constructs a node scheduling vector table based on the metrics; The scheduling layer generates a list of collaborative computing nodes using a greedy algorithm or a dynamic programming algorithm based on the user's computing tasks and the node scheduling vector table; The processing layer divides the user's computing tasks into multiple independently executed task units through task processing, and encapsulates each task unit into an operator with a unified input and output interface; The UDP-based lightweight remote procedure call LightRPC method is used to transfer the operator to the node specified in the collaborative computing node list for execution; When operators are executed on collaborative computing nodes, distributed task processing is performed through the parallel computing framework.
2. The distributed collaborative processing method for weak networks and full dynamics according to claim 1, characterized in that: Data is encoded and compressed using transmission algorithms, including: Adopt the message format based on UDP protocol to encapsulate and transmit standardized data; Compressing the encapsulated standardized data using at least one of a data dictionary, character replacement, dynamic encoding, and compression algorithm; According to network bandwidth and data priority, the compressed standardized data is divided into multiple data slices, and the data slices are sorted using a priority queue mechanism; Compare the data records of the local node and only transmit the changed data shards.
3. The distributed collaborative processing method for weak networks and full dynamics according to claim 2, characterized in that: Construct neighborhood data support table, including: Initialize the local data support table to record the network status data, computing power data and environment data of the local node; Discover and record neighbor nodes through at least one of a network adapter, a broadcast and monitoring mechanism, and a central node; Communicate with neighboring nodes, collect and update the network status data, computing power data and environmental data of neighboring nodes, and fill in the neighborhood data support table; According to the changes in network status, adjust the selection of neighbor nodes and update the neighborhood data support table; Periodically or when the local network topology changes, the data items in the local data support table and the neighborhood data support table are updated.
4. The distributed collaborative processing method for weak networks and full dynamics according to claim 3, characterized in that: The support layer uses an event-triggered update mechanism to update incremental data, including: When the network status data, computing power data or environment data of the local node or neighboring node changes and the change exceeds the threshold, an update event is triggered to incrementally update the changed data, and only the data items of the corresponding nodes in the local data support table or the neighborhood data support table are updated; An information dissemination mechanism combining epidemic and gossip is adopted to adjust the dissemination range and speed of the change information between nodes according to the network conditions; among them, the epidemic algorithm is used to disseminate information in the local area, and the gossip algorithm is used to disseminate information in the entire distributed network.
5. The distributed collaborative processing method for weak networks and full dynamics according to claim 4, characterized in that: The epidemi algorithm is used to propagate information in a local area, including: N nodes are set in the distributed collaborative processing system, and each node is set with a neighbor node list and an update information queue; When the local data support table or the neighborhood data support table of node i is updated, node i puts the updated data item and the corresponding identification information into the update information queue; Node i randomly selects a neighbor node j from the neighbor node list and sends all data items in the update information queue to node j; When node j receives the update information, it compares the received data items with the corresponding data items in the local data support table and the neighborhood data support table, puts the data items with inconsistent identification information into the update information queue, and propagates it to neighboring nodes; Repeat the above node propagation until all nodes in the local area receive the update information and complete the update of the local data support table and the neighborhood data support table.
6. The distributed collaborative processing method for weak networks and full dynamics according to claim 5, characterized in that: The gossip algorithm is used to spread information throughout the distributed network, including: The initial state of the gossip algorithm is set to each node i maintain a local view to record the information of other nodes obtained by node i through information exchange, and initialize the local view of node i to only contain the neighbor node information of node i; In each round of information exchange, a node j is randomly selected from the local view of node i as the information exchange object, and the amount of data for this round of information exchange is determined. The amount of data is calculated based on the preset propagation probability; In this round of information exchange, node i randomly selects node information corresponding to the amount of data from its local view and sends it to node j; at the same time, node j also sends the node information selected in its local view to node i in the same way; After completing this round of information exchange, nodes i and j will add the new node information obtained from each other to their respective local views; All nodes repeat the above information exchange and local view update steps until the preset convergence conditions are met, and then stop information exchange.
7. The distributed collaborative processing method for weak networks and full dynamics according to claim 6, characterized in that: The scheduling layer calculates the node's metrics based on the neighborhood data support table of each node, and constructs a node scheduling vector table based on the metrics, including: The scheduling layer calculates the network status metric Net between nodes based on the network status data of each node in the neighborhood data support table using the following formula: Net=f(Bandwidth,Latency,PacketLoss,Option) Among them, Bandwidth represents network bandwidth, Latency represents network delay, PacketLoss represents network packet loss rate, Option represents other candidate indicators that affect network status, and f is the network status evaluation function; The scheduling layer calculates the node's computing power measurement indicator Compute based on the computing power data of each node in the neighborhood data support table using the following formula: Compute=g(CPU,NPU,GPU,Memory,Option) Among them, CPU, NPU, GPU and Memory represent the CPU performance, NPU performance, GPU performance and memory capacity of the node respectively, and g is the computing capacity evaluation function; The scheduling layer calculates the node's environment matching metric Node based on the environment data of each node in the domain data support table using the following formula: Node=h(Hardware,OS,feature) Among them, Hardware represents the hardware characteristics of the node, OS represents the operating system characteristics of the node, feature represents the environmental characteristics of the node, and h is the environmental matching evaluation function; The scheduling layer combines the network status metric, computing power metric, and environment matching metric of each node into a multidimensional metric vector M i , as the comprehensive performance representation of node i, M i It is expressed as: M i =S(Net i ,Compute i ,Node i )=[v1,v2,v3,......,vn] Among them, Net i ,Compute i ,Node i They represent the network performance metric, computing performance metric, and environmental metric of node i respectively; M i is the metric vector of node i, expressed as [v1,v2,v3,......,vn], where v1,v2,v3,......,vn are M i The component values in each dimension; The scheduling layer measures the vector M of all nodes i Normalization is performed to generate a node scheduling vector table, which contains the number of each node, the node type, the relationship with other nodes, and the normalized metric vector of the corresponding node.
8. The distributed collaborative processing method for weak networks and full dynamics according to any one of claims 2 to 6, characterized in that: The scheduling layer uses a greedy algorithm or a dynamic programming algorithm to generate a list of collaborative computing nodes based on the user's computing tasks and the node scheduling vector table, including: The scheduling layer generates the task metric vector T according to the requirements of the computing task through the following formula: T=S(T)=[t1,t2,t3,...,tn] Among them, t1 represents the demand of computing tasks on the network, t2 represents the demand of computing tasks on node computing power, t3 represents the demand of computing tasks on node environment, S(T) is the task demand evaluation function, T is the generated task measurement vector, and T is expressed as [t1, t2, t3, ..., tn]; The scheduling layer uses a vector matching algorithm to calculate the task metric vector T and the metric vector M of each node in the node scheduling vector table through the following formula: i The matching degree d(M i ,T): Among them, M i1 ,M i2 ,......,M in Represents the metric vector M of node i i The dimensional components of T1, T2, ..., T n Represents the dimensional components of the task measurement vector T; The scheduling layer calculates the matching degree of all nodes and obtains the node matching degree list L; The scheduling layer divides computing tasks into small batch tasks and large batch tasks according to the computing scale requirements of computing tasks; For small batch tasks, the scheduling layer uses a greedy algorithm to select nodes with matching degrees higher than the threshold according to the node matching degree list L and add them to the collaborative computing node list C; For large batch tasks, the scheduling layer uses a dynamic programming algorithm to obtain the optimal collaborative computing node list C*; The collaborative computation list C or C* is output as the task scheduling result.
9. The distributed collaborative processing method for weak networks and full dynamics according to claim 8, characterized in that: For small batch tasks, the scheduling layer uses a greedy algorithm to select nodes with matching degrees higher than the threshold according to the node matching degree list L and add them to the collaborative computing node list C, including: Set the matching threshold a, initialize an empty collaborative computing node list C, and traverse each node i in the node matching list L; For each node i in the node matching list L, calculate the matching degree d(M i ,T); The matching degree d(M i ,T) nodes whose degree is greater than the threshold a are added to the collaborative computing node list C, the node matching list L is traversed, and the collaborative computing node list C is output.
10. The distributed collaborative processing method for weak networks and full dynamics according to claim 9, characterized in that: The UDP-based lightweight remote procedure call LightRPC method is used to transfer the operator to the node specified in the collaborative computing node list for execution, including: Split the computing task into multiple independent computing units and encapsulate each computing unit into a standardized operator; The LightRPC method based on the UDP protocol is used to transmit the encapsulated operator to the node specified in the collaborative computing node list: The original task node sends an execution request to the target collaborative node through LightRPC. The request content includes the operator, operator status and expected return result; LightRPC sends the request message to the target collaborative node in the form of a UDP datagram; After the target collaborative node completes the operator execution, it returns the execution result to the original task node in the form of a UDP datagram through LightRPC; LightRPC sets up a confirmation and retransmission mechanism, triggering retransmission when a request or result message is lost; After receiving the operator, the collaborative computing node executes the computing task according to the definition and interface standard of the operator, and returns the execution result to the original task node through LightRPC after completion; The original task node summarizes the results based on the execution results returned by the collaborative computing node.
Citation Information
Patent Citations
Multi-hop dynamic ad hoc network method of wide field sensor network
CN101909345A
Fog computing-based spatial information network architecture and method, and readable storage medium
CN109936619A