Data synchronization transmission method based on cloud edge collaboration

By analyzing the computing power usage of network nodes and optimizing the mapping position of data tables using hash rings, the problems of wasted computing resources and poor data coordination in cloud-edge computing clusters are solved, achieving load balancing and efficient data synchronization.

CN120768910BActive Publication Date: 2025-11-21BEIJING GUYU TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511033166.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-25
Publication Date
2025-11-21
Estimated Expiration
2045-07-25

AI Technical Summary

Technical Problem

In cloud-edge computing clusters, the uneven distribution of computing tasks among heterogeneous network nodes leads to waste of computing resources and poor data collaboration during data synchronization.

Method used

By analyzing the computing power utilization efficiency and computing power utilization time of network nodes, the performance consumption index is calculated. Hash rings are used to map and adjust network nodes, optimize the mapping position of data tables, and achieve balanced load distribution.

Benefits of technology

It improves the adaptability of computing task allocation among heterogeneous network nodes, reduces the overhead of data synchronization, and enhances data collaboration and computing resource utilization efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120768910B_ABST
    Figure CN120768910B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of data transmission, and particularly relates to a data synchronization transmission method based on cloud edge cooperation, wherein the present application maps network nodes into a hash ring to obtain hash partitions of the network nodes; analyzes load change conditions of the network nodes to obtain urgency degrees of each data table reported by the network nodes at each step length; accumulates the step length to analyze change trends and accumulated amounts of the urgency degrees, to obtain damage projection weights of the data tables, and then maps the data tables into the hash partitions to obtain initial mapping positions of the data tables; and adjusts the mapping positions of the data tables in combination with position relationships between the data tables and the nearest data tables of the data tables. The present application disperses responsibility ranges of physical network nodes to different positions of the hash ring through multiple virtual network nodes, avoids data skew, improves adaptability of split computing task distribution tables in the hash ring to heterogeneous network nodes, and thus improves data cooperativity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data transmission technology, and more specifically to a data synchronization and transmission method based on cloud-edge collaboration. Background Technology

[0002] Web server expansion primarily includes two methods: vertical expansion and horizontal expansion. Vertical expansion refers to enhancing hardware (such as replacing high-performance computing units), while horizontal expansion refers to increasing the number of hardware devices. Data synchronization and transmission between network nodes based on cloud-edge collaboration are mainly applied in horizontally expanded web servers. Clearly, for cloud-edge computing clusters, increasing the number of hardware devices directly increases the computing power of network nodes.

[0003] As the number of hardware devices in the server increases, the amount of data generated also increases. Server operations only require responses between similar servers, but the operational data needs to be stored in the database to record user actions. Under high-concurrency user saves, changed data needs to be saved. Each time a large amount of changed data is stored, it needs to be synchronized in the database, and each data synchronization needs to be broadcast to all network nodes in the database, resulting in a significant waste of cluster computing power in the data synchronization process. Therefore, after separating the database read and write strategies, it is necessary to synchronize the data stream to minimize the data synchronization overhead.

[0004] In existing technologies, the computational tasks and loads vary among network nodes in the database during data synchronization. This means that the scale of computation log synchronization differs between heterogeneous network nodes. Weighting the number of tasks on each network node determines the insertion position of a virtual network node, thus balancing the target network nodes. However, for heterogeneous computing entities in the cloud, the accuracy of the overall computational task ratio output is poor. This results in insufficient adaptability between the computational task allocation form partitioned in the hash ring and the heterogeneous network nodes, leading to poor data collaboration. Summary of the Invention

[0005] To address the above technical problems, the present invention aims to provide a data synchronization and transmission method based on cloud-edge collaboration.

[0006] According to a first aspect of the present invention, a data synchronization and transmission method based on cloud-edge collaboration is provided, the specific technical solution of which is as follows:

[0007] Based on the network node's work logs, analyze the network node's computing power utilization efficiency and computing power utilization duration to obtain the real-time performance consumption index of each network node.

[0008] Based on the performance consumption index, network nodes are mapped to the hash ring to obtain the hash partitions corresponding to the network nodes;

[0009] On the hash ring, at a unit step size, we analyze the load changes of network nodes to obtain the urgency of each data table reported by the network nodes at each step size.

[0010] The step size is accumulated, the trend of urgency and the amount of accumulation are analyzed, and the damage projection weight of each data table is obtained.

[0011] Based on the damage projection weights, the data table is mapped to the corresponding hash partition to obtain the initial mapping position of the data table;

[0012] In hash partitioning, based on the initial mapping position, the positional relationship between the data table and its nearest neighboring data table is analyzed to obtain the mapping adjustment distance of the data table, and the mapping position of the data table is adjusted accordingly.

[0013] In some embodiments of the present invention, based on the network node's work logs, the computing power utilization efficiency and computing power utilization duration of the network node are analyzed to obtain the real-time performance consumption index of each network node, including:

[0014] Obtain the real-time processor power, processor utilization, and disk read / write latency of the network node at the current reporting time from the network node's work log;

[0015] Calculate the ratio of the processor's real-time power to its rated power, and combine this with the processor utilization rate to obtain the real-time computing power utilization efficiency of the network node.

[0016] Based on the disk read / write latency, combined with the average read / write latency in the prior performance metrics of the network node, the real-time computing power usage time of the network node is obtained.

[0017] By combining computing power utilization efficiency and computing power utilization time, the real-time performance consumption index of each network node is obtained.

[0018] In some embodiments of the present invention, on a hash ring, at a unit step size, the load changes of network nodes are analyzed to obtain the urgency of each data table reported by the network nodes at each step size, including:

[0019] Obtain the data tables reported by network nodes;

[0020] Based on the number of data items in the data table, the processing time of the data table, and the time from the time the previous data table was sent to the current time, the urgency of increasing the load on each data table is evaluated to obtain the load increase urgency factor for each data table.

[0021] Under a unit step size, analyze the difference between the load increase urgency factor of the data table and the previous data table to obtain the urgency growth degree of the data table;

[0022] Based on the urgency of growth, iterate through all data tables at each step size to determine the urgency of growth for each data table at each step size, and report the urgency of each data table to the network node at each step size.

[0023] In some embodiments of the present invention, at a unit step size, the difference between the load increase urgency factor of the data table and its predecessor data table is analyzed to obtain the urgency growth degree of the data table, including:

[0024] Starting from the position of the data table, proceed forward in the order of the data tables reported by the network nodes, and obtain the previous data table in unit steps;

[0025] Calculate the difference between the load increase urgency factor of the data table and the corresponding previous data table, and combine the mean and standard deviation of the load increase urgency factor of all data tables to obtain the urgency increase degree of the data table.

[0026] In some embodiments of the present invention, the step size is accumulated, and the changing trend of urgency is analyzed, including:

[0027] The step lengths are accumulated, and the step lengths where the urgency increases are marked to obtain the number of first steps that are consistent with the direction of the longest urgency change.

[0028] Starting from the position in the data table, obtain the second step number that is continuously contained in the same direction as the urgency change in the data table;

[0029] Assess the changing trend of urgency based on the number of the first and second steps.

[0030] In some embodiments of the present invention, the analysis of the cumulative amount of urgency includes:

[0031] Within the second step range, the urgency of the data table is accumulated at each step size to obtain the accumulated amount of urgency for the corresponding data table.

[0032] In some embodiments of the present invention, the step size is accumulated, the changing trend and accumulation amount of urgency are analyzed, and the damage projection weight of each data table is obtained, including:

[0033] The output synchronization cost of the data table is obtained by weighting the cumulative amount by the trend of change.

[0034] During the accumulation of step size, the maximum value of the output synchronization cost of the data table is obtained and recorded as the damage projection weight of each data table.

[0035] In some embodiments of the present invention, in a hash partition, based on the initial mapping position, the positional relationship between the data table and its nearest neighboring data table is analyzed to obtain the mapping adjustment distance of the data table, and the mapping position of the data table is adjusted, including:

[0036] In hash partitioning, the hash distance between the initial mapping position of the data table and its nearest data table is calculated based on the mapping hash value of the initial mapping position of the data table and the mapping hash value of its nearest data table. The ratio of the mapping hash values ​​between the data table and its nearest data table is also calculated to obtain the mapping adjustment distance of the data table.

[0037] Adjust the mapping distance of the data table to move it away from the nearest corresponding data table, thus completing the adjustment of the mapping position of the data table.

[0038] In some embodiments of the present invention, the method further includes:

[0039] The central server broadcasts the time to all network nodes in real time through the same time synchronization unit to synchronize the timing of network nodes.

[0040] In some embodiments of the present invention, network nodes are mapped to a hash ring based on a performance consumption index to obtain the hash partition corresponding to the network node, including:

[0041] The SHA-256 algorithm is applied to the performance consumption index of the network node to obtain the mapped hash value of the network node;

[0042] Based on the mapping hash value of the network node, the network node is mapped to a hash ring, resulting in the hash partition corresponding to the network node. Compared with existing technologies, the cloud-edge collaborative data synchronization and transmission method provided by this invention has the following advantages:

[0043] This invention analyzes the computing power utilization efficiency and duration of network nodes based on their work logs to obtain a real-time performance consumption index for each node. Based on this index, network nodes are mapped to a hash ring to obtain their corresponding hash partitions, thus utilizing spatial partitioning for the synchronization of newly added data tables. On the hash ring, the load changes of network nodes are analyzed at unit step sizes to obtain the urgency of each data table reported by the network nodes at each step size. The urgency trend and accumulation are analyzed by accumulating the step sizes to obtain the damage projection weight for each data table. Based on this damage projection weight, the data tables are mapped to their corresponding hash partitions to obtain their initial mapping positions. Furthermore, considering the positional relationship between each data table and its nearest neighbor, the mapping positions are adjusted. This involves considering the data table's creation environment and projecting newly created data tables onto low-load network nodes within the hash ring as much as possible, thus creating a load shift from high to low and completing the collaborative load change among network nodes in the cloud. This invention distributes the responsibility of physical network nodes to different positions in the hash ring through multiple virtual network nodes, avoiding data skew and improving the adaptability of the computational task allocation form split in the hash ring with heterogeneous network nodes, thereby improving data collaboration. Attached Figure Description

[0044] To more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0045] Figure 1 A basic flowchart of a cloud-edge collaborative data synchronization and transmission method provided in one embodiment of the present invention is shown below;

[0046] Figure 2 This is a schematic diagram showing the distribution of physical and virtual network nodes in a hash ring before and after adjustment, as provided in an embodiment of the present invention. Detailed Implementation

[0047] To further illustrate the technical means and effects adopted by the present invention to achieve its intended purpose, the following, in conjunction with the accompanying drawings and preferred embodiments, details the specific implementation, structure, features, and effects of the cloud-edge collaborative data synchronization transmission method proposed according to the present invention. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. Furthermore, specific features, structures, or characteristics in one or more embodiments can be combined in any suitable form.

[0048] It should be noted that, in order to ensure that the calculation results are meaningful, when performing fractional operations, if the denominator is 0, a parameter adjustment factor greater than 0 needs to be added to the denominator to prevent the denominator from being 0. The value of the parameter adjustment factor shall be set by the implementer according to the actual situation, and this application does not impose any special restrictions.

[0049] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. Terms such as “comprising,” “including,” or any other variations thereof are intended to cover a non-exclusive inclusion, such that a circuit structure, article, or device comprising a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such an article or device. Without further limitation, an element defined by the phrase “comprising one…” does not exclude the presence of other identical elements in the article or device that includes that element. Relational terms such as “first” and “second” are used merely to distinguish one entity or operation from another and do not necessarily require or imply any such actual relationship or order between these entities or operations.

[0050] The following description, in conjunction with the accompanying drawings, details a specific scheme for a cloud-edge collaborative data synchronization transmission method provided by the present invention.

[0051] Please see Figure 1 This illustrates the basic process of a cloud-edge collaborative data synchronization transmission method provided by an embodiment of the present invention.

[0052] like Figure 1 As shown, an embodiment of the present invention provides a data synchronization and transmission method based on cloud-edge collaboration, which specifically includes:

[0053] S100: Based on the network node's work logs, analyze the network node's computing power utilization efficiency and computing power utilization duration to obtain the real-time performance consumption index of each network node.

[0054] First, each network node sends its local routing table back to the central server. The central server retrieves the MAC addresses of all network nodes within the current network topology from the routing table and establishes a bidirectional directed graph of node-link directions by comparing the routing table data. Additionally, the central server broadcasts its time to all network nodes in real-time via the same time synchronization unit, synchronizing the network nodes' timing. These operations provide the foundation for the subsequent synchronization of data tables reported by each network node.

[0055] Then, the computational performance consumption index of the network nodes is evaluated. Due to the heterogeneity of the network nodes, the evaluation of the output phenomenon is based on their computational power utilization efficiency and computational power utilization time. In the embodiments of the present invention, the computational power utilization efficiency and computational power utilization time of the network nodes are analyzed based on the network node's work logs to obtain the real-time performance consumption index of each network node. Specifically:

[0056] First, obtain the real-time processor power, processor utilization, and disk read / write latency of the network node at the current reporting time from the network node's work log.

[0057] Then, the ratio of the network node's real-time processor power to its rated processor power at the current reporting moment is calculated. Combined with the processor utilization rate, the real-time computing power utilization efficiency of the network node is obtained. This is used to measure the output ratio of the network node's processor utilization rate to determine whether a large-scale single-threaded computing task is causing blocking. In other words, in the case of single-threaded blocking, a situation of low power and high utilization rate will occur, indicating that a single core of the processor is continuously working.

[0058] Next, based on the disk read / write latency of the network node at the current reporting moment, and combined with the average read / write latency in the network node's prior performance metrics, the difference between the disk read / write latency and the average read / write latency is calculated to obtain the real-time computing power usage time of the network node. Since the impact of single-threaded operation blocking caused by excessive disk usage should be greater on performance than the impact of large-scale multi-threaded usage, the actual computing power usage time of the network node is obtained.

[0059] Finally, combining the computing power utilization efficiency and computing power utilization time, the real-time performance consumption index of each network node is obtained as follows:

[0060]

[0061] In the formula, x i U' represents the performance consumption exponent of the i-th network node at time t; t P' represents the processor utilization rate of the i-th network node at time t; t P represents the real-time processor power of the i-th network node at time t; i p represents the rated power of the processor of the i-th network node at time t. i This represents the ratio of the processor's real-time power to its rated power; y' t This represents the disk read / write latency of the i-th network node at time t; This represents the average read / write latency in the prior performance metrics of the i-th network node.

[0062] This represents the computing power utilization efficiency, which is used to measure the output ratio of the processor utilization rate of network nodes. It determines whether there is a large-scale single-threaded computing task causing blocking. In other words, in the case of single-threaded blocking, there will be a situation of low power and high utilization rate, indicating that a single core of the processor is working continuously. The larger the value, the more significant the single-threaded blocking is, and the greater the consumption of network node computing performance. This indicates the real-time computing power usage time of a network node. The larger this value, the more significant the single-threaded blocking situation is, and the greater the consumption of computing performance of the network node. In other words, the impact of single-threaded operation blocking caused by excessive disk usage on performance consumption should be greater than the impact of multi-threaded heavy usage on performance consumption.

[0063] Similarly, obtain the real-time performance consumption index of all network nodes.

[0064] Thus, the evaluation of the real-time performance consumption parameters of network nodes has been completed. Since the linear calculation is performed by directly reading log data, the time complexity is 0, and a unified measurement of the performance of heterogeneous network nodes is achieved while minimizing computational overhead.

[0065] S200: Based on the performance consumption index, network nodes are mapped to the hash ring to obtain the hash partition corresponding to the network node.

[0066] The process of synchronizing newly added data tables reported by network nodes requires projecting the network nodes into a ring hash space, thereby utilizing space partitioning for the synchronization operation of new data tables. Specifically, the SHA-256 algorithm is applied to the performance consumption index of the network node to obtain the mapped hash value of the network node; based on the mapped hash value of the network node, the network node is mapped to the hash ring to obtain the hash partition corresponding to the network node.

[0067] S300: On the hash ring, analyze the load changes of network nodes at unit step size to obtain the urgency of each data table reported by the network nodes at each step size.

[0068] The traditional hash mapping process projects data tables based on the density of network nodes. This process can lead to uneven density of projection points and often only measures the internal content of newly added data tables without evaluating the environment in which the data tables are generated. The same data table content can produce the same network node load changes, but when these changes occur under the influence of high network node load, they can significantly increase the difficulty of network node data synchronization, causing throughput delays and computational blockages.

[0069] Therefore, during the process of mapping data tables after mapping network nodes, by considering the environment in which the data tables are generated, the data is projected onto low-load network nodes in the hash ring as much as possible, thereby forming a load shift trend from high to low and completing the collaborative load change between network nodes in the cloud.

[0070] Based on the above analysis, in an embodiment of the present invention, on the hash ring, at a unit step size, the load change of network nodes is analyzed to obtain the urgency of each data table reported by the network nodes at each step size. Further steps include:

[0071] First, after the network node mapping is completed, the data tables reported by the network nodes are retrieved. Considering the constantly changing nature of the real-time load of network nodes, the load change trend of the data tables within the specified time period should be analyzed for accurate assessment. Therefore, specifically, for network node i, all data tables reported by network node i up to the previous synchronization time period are extracted.

[0072] Then, based on the number of data items in each data table and the processing time of the data table, the urgency of increasing the load on each data table is evaluated, resulting in a load increase urgency factor for each data table. Specifically, for any data table b, the number m of data items in data table b is obtained. b Processing time t for data table b b Compared with the previous data table t b-1 (Starting from position b in data table, proceeding forward in the order of data tables reported by network nodes, and taking unit steps, retrieve the data table preceding data table b) Time difference t from the time of transmission to the current time. bi,bi-1 Calculate the processing time t b With time difference t bi,bi-1 The difference t b -t bi,bi-1 And calculate the number of data items m. b and the difference t b -t b,b-1 The ratio of these values ​​yields the load increase urgency factor for data table b, i.e.:

[0073]

[0074] In the formula, k b This indicates that the load on data table b has increased due to the urgency factor; m b t represents the number of data items in table b; b This indicates the processing time for data table b; t bi,bi-1 This represents the time difference between the time the previous data table was sent and the current time.

[0075] The load increase process is described as follows: When the network node load increases rapidly, the data items in the data table are quickly filled, thus the data table return time for adjacent return sequences gradually shortens (t). b- t bi,bi-1 (reduce), and based on the number of data items m in the data table b As a weight, the accurate assessment of computationally intensive tasks means that the more data items there are, the more different resources network nodes need to use during task scheduling, thereby further extending the processing cost after the load increases.

[0076] Next, based on the changing trend of the urgency factor, the local data synchronization needs are determined, thereby determining the requirements for the distributed mapping of dense data points, ensuring more accurate collaboration between nodes in the cloud. In some embodiments of the present invention, at a unit step size, the difference between the load increase urgency factor of a data table and its predecessor is analyzed to obtain the urgency growth degree of the data table. Specifically, starting from the data table position, proceeding forward in the order of network node reporting data tables, at a unit step size, the predecessor data table is obtained; the difference between the load increase urgency factor of a data table and its predecessor is calculated, and combined with the mean and standard deviation of the load increase urgency factor of all data tables, the urgency growth degree of the data table is obtained as follows:

[0077]

[0078] In the formula, G b Indicates the degree of urgency of growth in data table b; k b This indicates that the load on data table b has increased due to the urgency factor; k b-1 This indicates that the load on data table b-1 has increased by the urgency factor; This represents the average urgency factor for the increase in load across all data tables; σ k This indicates the urgency factor standard deviation representing the load increase across all data tables.

[0079] Finally, based on the urgency of growth, the urgency of growth corresponding to all data tables at each step size is iterated, and the urgency of each data table reported to the network node at each step size is:

[0080]

[0081] In the formula, K b This indicates that the data table b reported by the network node is in n. t Urgency level under step size; n t Indicates the number of steps already acquired; This refers to the data table bt. b The degree of urgency of growth, including data in table bt b This represents the forward step size t from data table b.b A table of data reported by network nodes at any given time.

[0082] By controlling the short-term cumulative urgency growth rate to balance the proportion, the load increase can be assessed in real time. The higher the urgency growth rate value up to the current position, the more steps there are of high growth, and therefore the stronger the load impact appears at the current data table position; similarly, the lower the urgency growth rate value, the smaller the load impact.

[0083] Iterate through all the data tables obtained in the current time period to see the output changes and obtain the urgency of each step size for each data table.

[0084] S400: Accumulate the step size, analyze the changing trend and accumulation amount of urgency, and obtain the damage projection weight of each data table.

[0085] The mapping set of the hash space relies on the fact that the characteristics of the data tables are similar, resulting in similar hash values ​​and thus similar mapping positions. This reflects that the load states of the network nodes described by the data tables used for synchronization between network nodes in the current environment are relatively consistent, meaning that the generation state of the data tables at the current moment is also close to consistent. Therefore, for all data tables obtained in the current time period, the computational task consumption caused by the urgency at each step size is specifically balanced with the data table generation environment.

[0086] The effective urgency of the output data table is evaluated. In other words, the data table with a high urgency level that has accumulated in a short period of time should have large gaps to avoid a large amount of load being concentrated on the same network node.

[0087] Based on the above analysis, in the embodiments of the present invention, by accumulating the step size, analyzing the changing trend and accumulation amount of urgency, the damage projection weight of each data table is obtained. Further:

[0088] The urgency value range is from negative infinity to positive infinity. However, due to the limitations of the data form and the fact that the generated data is still within the limits of network node performance, the output urgency value range is within a controllable range.

[0089] Therefore, in some embodiments of the present invention, the step length is accumulated, and the trend of urgency change is analyzed. Specifically, this includes: accumulating the step length, marking the step lengths where urgency increases, and obtaining the number of steps contained in the longest urgency change direction during the step length accumulation process, i.e., the number of steps contained in the longest continuous increase interval or the longest continuous decrease interval, denoted as... And obtain the number of consecutive steps where the urgency change direction is the same for data tables of different time lengths, that is, the number of consecutive steps where the urgency sign is the same for data tables of different time lengths, denoted as n. tb Based on the first and second steps, assess the changing trend of urgency in the data table, i.e., calculate the ratio of the second step number to the first step number. As an assessment of the changing trend of the urgency of the data table.

[0090] In some embodiments of the present invention, analyzing the accumulated amount of urgency specifically includes: within the second step range, that is, within the range of the number of consecutive step lengths in which the urgency signs of the data table are the same under different step lengths, accumulating the urgency of the data table under each step length to obtain the accumulated amount of urgency of the data table, denoted as sum(K). b ).

[0091] Furthermore, based on the changing trend and cumulative amount of urgency, the output synchronization cost of the data table is obtained. Specifically, the cumulative amount is weighted by the changing trend to obtain the output synchronization cost of the data table as follows:

[0092]

[0093] In the formula, C b Indicates the output synchronization cost of data table b; n tb This indicates the second step number in which the direction of change in urgency of data table b is consecutively the same under different step lengths; This indicates the number of first steps included in the longest urgency change direction during the step-size accumulation process; sum represents the summation function; sum(K) b ) represents the cumulative amount of urgency corresponding to data table b.

[0094] This represents the changing trend of the urgency of data table b. The larger the value, the greater the synchronization difficulty of this data table during the synchronization process, and the greater the output synchronization cost of data table b; sum(K b The urgency level of data table b is represented by the cumulative amount. A larger value indicates a larger synchronization scale for that data table during the synchronization process, and consequently, a higher output synchronization cost for data table b. Therefore, the balance of network nodes corresponding to the data table should be considered when calculating the output synchronization cost. The balance of network nodes corresponding to the data table means that data tables under high load, i.e., those experiencing abnormal load increases in a short period of time, are stored centrally on the same network node to accurately describe synchronization events. However, the triggering conditions of different network nodes should be separated because the performance consumption of network nodes is similar, resulting in similar mapping positions in the hash ring. Therefore, the mapping positions of data tables on different network nodes should be separated.

[0095] Finally, during the accumulation of step sizes, the maximum value of the output synchronization cost of each data table is obtained and recorded as the damage projection weight for each data table. Similarly, the damage projection weights of all data tables to be synchronized are obtained.

[0096] S500: Based on the damage projection weight, the data table is mapped to the corresponding hash partition to obtain the initial mapping position of the data table.

[0097] Based on the damage projection weights, the data table is mapped to the corresponding hash partitions to obtain the initial mapping position of the data table. Specifically, the SHA-256 algorithm is applied to the damage projection weights of the data table, resulting in the mapped hash value H of the data table. b Based on the mapping hash value H of the data table b Map the data table to the corresponding hash partition to obtain the initial mapping position of the data table.

[0098] S600: In a hash partition, based on the initial mapping position, analyze the positional relationship between the data table and its nearest data table to obtain the mapping adjustment distance of the data table, and adjust the mapping position of the data table.

[0099] By transforming the local mapped hash values, the initial mapped position of the data table is reprojected. Therefore, in hash partitioning, based on the initial mapped position, the positional relationship between the data table and its nearest neighbor is analyzed to obtain the mapping adjustment distance of the data table, and the mapped position of the data table is adjusted accordingly. Specifically:

[0100] First, in the hash partition, based on the mapped hash value of the initial mapped position of the data table and the mapped hash value of its nearest neighboring data table, the hash distance between the initial mapped position of the data table and its nearest neighboring data table is calculated, denoted as d(H). b ,H' b ) min , where H b H' represents the mapping hash value of the initial mapping position of data table b. b This represents the hash value of the data table that is closest to data table b.

[0101] Additionally, the mapping hash value ratio between the data table and its nearest neighboring data table is calculated to obtain the mapping adjustment coefficient for the data table:

[0102]

[0103] In the formula, a b h represents the mapping adjustment factor for data table b; b H' represents the mapping hash value indicating the initial mapping position of data table b. bThis represents the hash value of the data table that is closest to data table b; softsign represents the activation function.

[0104] H b With H' b When the deviation is large, the value after mapping with the softsign function is closer to 0, which shows that the need to adjust the two data tables is lower, that is, direct mapping can complete the splitting of the synchronization results.

[0105] Then, based on the hash distance and mapping adjustment coefficient corresponding to the data table, the mapping adjustment distance of the data table is obtained as follows:

[0106] d′=(a b +1)×d(H b ,H′ b ) min

[0107] In the formula, d' is the mapping adjustment distance of data table b; a b This represents the mapping adjustment coefficient for data table b; d(H) b ,H' b ) min This represents the hash distance between the initial mapping position of data table b and its nearest data table.

[0108] Finally, the mapping distance is adjusted to move the data table away from its nearest corresponding data table, completing the adjustment of the data table's mapping position. Multiple virtual network nodes distribute the responsibility of the physical network nodes across different locations on the hash ring, preventing data skew. Please refer to [link / reference]. Figure 2 This diagram illustrates the distribution of physical and virtual network nodes in the hash ring before and after adjustment. Data replicas are distributed across the hash ranges of multiple network nodes; if any network node fails, the replica data is taken over by other network nodes.

[0109] It should be noted that the order of the above embodiments of the present invention is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. The processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0110] The various embodiments in this specification are described in a progressive manner. For the same or similar parts between the various embodiments, please refer to each other. Each embodiment focuses on describing the differences from other embodiments.

Claims

1. A method for cloud-edge collaborative data synchronization transmission, characterized in that, The method comprises: According to the working log of the network node, the efficiency of the network node is analyzed, and the efficiency of the network node is analyzed, and the performance consumption index of each network node is obtained in real time; According to the performance consumption index, the network node is mapped to the hash ring, and the corresponding hash partition of the network node is obtained; On the hash ring, under the unit step, the load change of the network node is analyzed, and the urgency of each data table reported by the network node under each step is obtained; The step is accumulated, the change trend and the accumulation of the urgency are analyzed, and the damage projection weight of each data table is obtained; According to the damage projection weight, the data table is mapped to the corresponding hash partition, and the initial mapping position of the data table is obtained; In the hash partition, according to the initial mapping position, the position relationship between the data table and the nearest data table is analyzed, the mapping adjustment distance of the data table is obtained, and the mapping position of the data table is adjusted. 2.The cloud-edge collaboration based data synchronization transmission method according to claim 1, characterized in that, According to the working log of the network node, the efficiency of the network node is analyzed, and the efficiency of the network node is analyzed, and the performance consumption index of each network node is obtained in real time, comprising: From the working log of the network node, the real-time power of the processor of the network node at the current reporting time, the processor occupancy rate and the disk read-write delay time are obtained; The proportion of the real-time power of the processor to the rated power of the processor is calculated, and the real-time efficiency of the network node is obtained by combining the processor occupancy rate; According to the disk read-write delay time, the average read-write delay time in the prior performance index of the network node is combined to obtain the real-time efficiency of the network node; The performance consumption index of each network node in real time is obtained by combining the efficiency of the network node and the efficiency of the network node. 3.The cloud-edge collaboration based data synchronization transmission method of claim 1, wherein, On the hash ring, under the unit step, the load change of the network node is analyzed, and the urgency of each data table reported by the network node under each step is obtained, comprising: Obtain the data table reported by the network node; According to the data item quantity of the data table, the processing time of the data table and the time from the sending time of the previous data table to the current time, the load increase urgency of each data table is evaluated, and the load increase urgency factor of each data table is obtained; Under the unit step, the difference between the load increase urgency factors corresponding to the data table and its previous data table is analyzed, and the urgency increase degree of the data table is obtained; According to the urgency increase degree, the urgency increase degree corresponding to all data tables under each step is traversed, and the urgency of each data table reported by the network node under each step is obtained.

4. The cloud-edge collaboration based data synchronization transmission method according to claim 3, characterized in that, Under the unit step, the difference between the load increase urgency factors corresponding to the data table and its previous data table is analyzed, and the urgency increase degree of the data table is obtained, comprising: From the data table position, forward to the data table reported by the network node in sequence, and obtain the previous data table of the data table in unit step; The difference between the load increase urgency factor corresponding to the data table and its previous data table is calculated, and the average and standard deviation of the load increase urgency factors of all data tables are combined to obtain the urgency growth degree of the data table.

5. The cloud-edge collaboration based data synchronization transmission method according to claim 1, characterized in that, The step is accumulated, and the change trend of the urgency degree is analyzed, including: The step is accumulated, and the step with the increasing urgency degree is marked to obtain the first step number contained in the longest change direction of the urgency degree that is consistent; The second step number contained in the continuous same change direction of the urgency degree of the data table is obtained from the data table position; The change trend of the urgency degree is evaluated according to the first step number and the second step number.

6. The cloud-edge collaboration based data synchronization transmission method according to claim 5, characterized in that, The accumulated amount of the urgency degree is analyzed, including: The accumulated amount of the urgency degree of the data table at each step is obtained by accumulating the urgency degree of the data table within the second step number range.

7. The cloud-edge collaboration based data synchronization transmission method according to claim 6, characterized in that, The change trend and the accumulated amount of the urgency degree are analyzed by accumulating the step to obtain the damage projection weight of each data table, including: The output synchronization cost of the data table is obtained by weighting the accumulated amount according to the change trend; The maximum value of the output synchronization cost of the data table is obtained during the accumulation of the step, which is recorded as the damage projection weight of each data table. 8.The cloud-edge collaboration based data synchronization transmission method of claim 1, wherein, In the hash partition, the mapping adjustment distance of the data table is obtained by analyzing the positional relationship between the data table and its nearest data table according to the initial mapping position, and the mapping position of the data table is adjusted, including: In the hash partition, the mapping adjustment distance of the data table is obtained by calculating the hash distance between the initial mapping position of the data table and its nearest data table and calculating the mapping hash value ratio between the data table and its nearest data table according to the mapping hash value of the initial mapping position of the data table and the mapping hash value of its nearest data table; The mapping position of the data table is adjusted by adjusting the mapping adjustment distance in the direction away from the nearest data table corresponding to the data table. 9.The cloud-edge collaboration based data synchronization transmission method of claim 1, wherein, The method further includes: The central server broadcasts time to all network nodes in the network in real time through the same time unit to synchronize the network node timing. 10.The cloud-edge collaboration based data synchronization transmission method according to claim 1, characterized in that, According to the performance consumption index, the network node is mapped into a hash ring to obtain the hash partition corresponding to the network node, including: The performance consumption index of the network node is subjected to SHA-256 algorithm to obtain the mapping hash value of the network node; According to the mapping hash value of the network node, the network node is mapped into a hash ring to obtain the hash partition corresponding to the network node.

Citation Information

Patent Citations

  • Affinity dynamic load balancing method based on consistent hash algorithm

    CN107197035A

  • Load balancing method based on clusters and multiple processes

    CN111858033A