Data synchronization method based on edge side distributed cache

Through the occupational behavior entropy (OBE) model and dynamic sharding mechanism, the problems of low cache hit rate and network resource waste in edge-side cache are solved, and efficient and accurate data synchronization is achieved.

CN120743882AInactive Publication Date: 2025-10-03JILIN COMM POLYTECHNIC
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510949829.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-10
Publication Date
2025-10-03
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

When processing employee data, the existing edge-side distributed caching technology uses a fixed sharding strategy, resulting in low cache hit rates and a surge in back-to-origin requests. The full synchronization mechanism also wastes network resources and causes data inconsistency.

Method used

The occupational behavior entropy (OBE) quantitative model and dynamic sharding-incremental synchronization mechanism are adopted to calculate data activity through access frequency variance, timeliness decay rate and multi-source correlation, dynamically divide data into hot, warm and cold shards, and use reverse cache temporary storage and UDP broadcast retry mechanism to achieve incremental synchronization when the network is interrupted.

Benefits of technology

It improves the cache hit rate, reduces back-to-origin requests, reduces network resource consumption, ensures the efficiency and accuracy of data synchronization, and adapts to changes in data activity in different business scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120743882A_ABST
    Figure CN120743882A_ABST
Patent Text Reader

Abstract

The invention discloses a data synchronization method based on edge side distributed cache, and relates to the technical field of data management. The method comprises the following steps: S1, calculating the occupational behavior entropy OBE of data, wherein the OBE is calculated through three dimensions of access frequency variance, timeliness attenuation rate and multi-source correlation Cr; according to the method, a dynamic fragmentation mechanism is constructed by introducing occupational behavior entropy OBE, the problems of low cache hit rate and back-to-source request sharp increase caused by a traditional fixed fragmentation strategy are effectively solved, and the operation principle of the method is as follows: firstly, the data activity degree is quantized based on three dimensions of access frequency variance, timeliness attenuation rate and multi-source association degree Cr; wherein the data distribution imbalance is reflected through the access frequency variance of different edge nodes of a certain type of data in a dynamic time window, the data type is combined to set an attenuation coefficient to calculate the timeliness, and the Cr represents the data association degree through knowledge graph association edge number normalization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of data management technology, and in particular relates to a data synchronization method based on edge-side distributed cache. Background Art

[0002] Amidst the backdrop of digital transformation, companies are increasingly prioritizing employee data. This data encompasses not only employee training logs and skill assessment records, but also diverse behavioral information, playing a key role in helping companies assess employee capabilities and formulate development strategies. Edge-side distributed caching technology is a crucial means of storing and managing this data, and its performance directly impacts its availability and timeliness. However, existing edge-side distributed caching technologies, such as Redis and Memcached, exhibit numerous drawbacks when processing data.

[0003] Currently, traditional edge-side distributed caches employ fixed sharding rules, such as those based on time windows or regions, which are difficult to adapt to the dynamic nature of data. Data update frequencies for employees at different stages vary significantly. For example, probationary employees have extremely frequent data updates due to frequent training and job adaptation. Meanwhile, data for key employees, such as annual professional ethics files, is updated less frequently. In this scenario, a fixed sharding strategy results in a mixed storage of frequently updated "hot data" and less frequently updated "cold data," significantly reducing cache hit rates. Furthermore, cold data permanently occupies valuable edge node storage resources, significantly crowding out cache space for hot data. When edge nodes need to retrieve hot data, they are forced to frequently initiate cross-node back-to-source requests due to insufficient hot data in their local caches. This significantly increases network transmission pressure and leads to a surge in back-to-source requests. For data synchronization, existing technologies primarily rely on full pull or scheduled task mechanisms. Full synchronization means that a large amount of data must be transmitted each time synchronization is performed. This not only seriously consumes network bandwidth resources, but also, in a multi-node environment, due to the lack of an effective conflict detection mechanism, data inconsistencies are very likely to occur. For example, multi-node cache version conflicts lead to differences in data stored in different nodes, which in turn affects the company's accurate assessment and decision-making on employee professionalism.

[0004] To this end, the present invention provides a data synchronization method based on edge-side distributed cache to solve the above problems. Summary of the Invention

[0005] The purpose of the present invention is to provide a data synchronization method based on edge-side distributed cache. By combining the occupational behavior entropy (OBE) quantification model with the dynamic sharding-incremental synchronization mechanism, the present invention solves the problems of rigid cache resource allocation caused by the fixed sharding strategy in the prior art, low hit rate caused by mixed storage of hot and cold data, and waste of network resources and data inconsistency caused by the full synchronization mechanism.

[0006] To solve the above technical problems, the present invention is implemented through the following technical solutions.

[0007] The present invention is a data synchronization method based on edge-side distributed cache, comprising the following steps: S1: Calculate the OBE of the data The OBE accesses frequency variance , time-dependent decay rate , multi-source correlation Cr three-dimensional calculation, the calculation formula is OBE=α⋅ +β⋅ +γ⋅Cr, where α+β+γ=1, and the weights of α, β, and γ are dynamically adjusted according to the business scenario. OBE quantifies data activity through three dimensions. The weight coefficients α, β, and γ meet the normalization condition α+β+γ=1, ensuring that the calculation result is in the range [0,1]. The weights are configured according to the business scenario to reflect the different contributions of different dimensions to data activity; S2: Dynamically shard the data based on the OBE entropy The data is divided into hot shards, warm shards and cold shards, where OBE>preset threshold is hot shards, OBE within the preset threshold range is warm shards, and OBE≤preset threshold is cold shards. The hot shards are deployed to the edge nodes of the terminal devices, the warm shards are stored in the regional edge servers, and the cold shards are archived to the central cloud. After the shards are stored, the system continuously monitors the changes in the OBE values ​​of the data.

[0008] S3: Incremental synchronization trigger and failure retry When the OBE entropy change of the data exceeds the preset threshold, the shard migration mechanism is triggered. When the data is upgraded from a warm shard to a hot shard, the migration of the data to the terminal node is started, and only the fields related to the OBE change are transmitted to avoid full synchronization. When the network is interrupted, the data to be migrated is temporarily stored in the local cache and retried through UDP broadcast, triggering incremental synchronization between edge nodes and transmitting differential data through the gRPC lightweight protocol. If the synchronization fails, the reverse cache temporary storage strategy is adopted to temporarily store the data to be synchronized in the local temporary cache area and mark it as waiting for retransmission. After the network is restored, retry is performed through UDP broadcast.

[0009] The present invention is further configured such that the access frequency variance The calculation method is: statistically analyze the variance of the number of visits of a certain type of data at different edge nodes within the dynamic time window. The calculation formula is: ,in, is the total number of edge nodes, For the The number of visits to edge nodes, is the average number of visits, = ,The standard deviation is used to measure the discreteness of the access times, which reflects the distribution balance of data in edge nodes. The larger the value, the more concentrated the data access hotspot.

[0010] The present invention is further configured such that the time-dependent decay rate The calculation method is: set the attenuation coefficient α according to the data type, calculate the current time Data generation time , the calculation formula is: =α⋅( ), time difference ( ) linearly amplifies the attenuation effect, and α controls the attenuation speed.

[0011] The present invention is further configured such that the multi-source correlation degree Cr is calculated by constructing a professional behavior data association network through a knowledge graph, extracting the number of associated edges between the target data and other data, and normalizing them to the interval [0,1]. The calculation formula is: Cr= , where the number of associated edges is the number of associated edges between the target data and other data in the knowledge graph, and the maximum possible number of associated edges is the maximum number of edges that may exist between all nodes in the graph. Through normalization, the number of associated edges is compressed to the [0,1] interval. The higher the value, the stronger the data correlation.

[0012] The present invention is further configured such that the business scenario includes a manufacturing training scenario and a service industry assessment scenario, and the weight numerical ratios of α, β, and γ are different in the manufacturing training scenario and the service industry assessment scenario.

[0013] The present invention is further configured such that the edge node periodically recalculates the OBE entropy of the governed data, triggers a shard migration mechanism based on the calculation result, migrates hot shards from the regional server to the terminal node, or archives cold shards from the terminal node to the central cloud.

[0014] The present invention is further configured such that the difference data only includes fields related to OBE changes, and the synchronization packet structure includes a data unique identifier, old and new OBE values, a difference field list, and a CRC32 check code.

[0015] The present invention is further configured such that the local temporary buffer area is provided with a time to live TTL.

[0016] The present invention is further configured such that the regional edge server covers 3-5 surrounding branch nodes.

[0017] The present invention is further configured such that the hot shards are deployed to the edge nodes of the terminal devices, and local cache can be used to reduce cloud back-to-source; the warm shards are stored in the regional edge servers, which can balance storage costs and access efficiency; the cold shards are archived to the central cloud and are only pulled on demand when queried.

[0018] The present invention has the following beneficial effects.

[0019] 1. This invention introduces occupational behavior entropy (OBE) to build a dynamic sharding mechanism, which effectively solves the problems of low cache hit rate and surge in back-to-source requests caused by traditional fixed sharding strategies. Its operating principle is as follows: First, based on the access frequency variance , time-dependent decay rate , multi-source correlation Cr are three dimensions to quantify the data activity, among which The data distribution imbalance is reflected by the variance of the number of visits to different edge nodes of a certain type of data in the dynamic time window. The timeliness is calculated by setting the attenuation coefficient in combination with the data type. Cr represents the data association by normalizing the number of associated edges in the knowledge graph. The three are weighted by dynamic weights α, β, and γ to obtain the OBE value. According to the OBE value, the data is divided into three types of shards: hot, warm, and cold. They are deployed to the edge nodes of terminal devices, regional edge servers, and central clouds respectively, to achieve tiered storage of hot data locally, warm data regionally, and cold data in the cloud. This mechanism enables hot data with high frequency updates to occupy the cache space of terminal nodes first, avoiding mixed storage with low frequency cold data. The cache hit rate of hot data is greatly improved, and the cross-node back-to-source requests are greatly reduced, which completely changes the cache resource mismatch problem caused by the traditional fixed sharding strategy and fundamentally reduces the dependence of edge nodes on the cloud.

[0020] 2. To address the bandwidth waste and synchronization delays caused by traditional full synchronization, the present invention designs an OBE-driven incremental synchronization and reverse cache retry mechanism. When the data OBE value changes by more than a threshold, only the difference fields related to the OBE change are transmitted via the gRPC protocol. The synchronization packet structure contains the data's unique identifier, the old and new OBE values, a list of difference fields, and a CRC32 checksum. This significantly reduces invalid data transmission compared to full synchronization. If synchronization fails due to network fluctuations, a reverse cache temporary storage strategy is adopted to store the data to be synchronized in a temporary buffer and mark it for retransmission. After the network recovers, a retry is initiated via UDP broadcast, avoiding the long delays caused by single failures in traditional synchronization mechanisms. This mechanism triggers precise synchronization by dynamically sensing changes in data activity. Combined with a lightweight transmission protocol and an intelligent retry strategy, synchronization delay is reduced. Especially in low-bandwidth industrial scenarios, this mechanism ensures second-level synchronization of hot data such as real-time operation scores while avoiding redundant transmission of cold data such as monthly summaries. This completely solves the resource waste and timeliness issues caused by the one-size-fits-all approach of traditional full synchronization, providing efficient and reliable underlying support for real-time data evaluation. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments.

[0022] Figure 1 This is the flowchart for calculating occupational behavior entropy (OBE) in the data synchronization method based on edge-side distributed cache.

[0023] Figure 2 This is a flowchart of dynamic sharding and storage deployment in the data synchronization method based on edge-side distributed cache.

[0024] Figure 3 This is a flowchart of synchronization trigger determination in a data synchronization method based on edge-side distributed cache.

[0025] Figure 4 This is a flowchart for incremental synchronization execution in the data synchronization method based on edge-side distributed cache.

[0026] Figure 5 This is a flowchart of retrying synchronization failure in the data synchronization method based on edge-side distributed cache.

[0027] Figure 6 Schematic diagram of the knowledge graph in the data synchronization method based on edge-side distributed cache. DETAILED DESCRIPTION

[0028] The technical solutions in the embodiments of the present invention will be described below in conjunction with the drawings in the embodiments of the present invention. The described embodiments are only part of the embodiments of the present invention, rather than all the embodiments.

[0029] Example 1 See also Figure 1-6 The present invention provides a data synchronization method based on edge-side distributed cache, comprising the following steps: S1: Calculate the OBE of the data OBE by accessing frequency variance , time-dependent decay rate , multi-source correlation Cr three-dimensional calculation, the calculation formula is OBE=α⋅ +β⋅ +γ⋅Cr, where α+β+γ=1, and the weights of α, β, and γ are dynamically adjusted according to the business scenario. OBE quantifies data activity through three dimensions. The weight coefficients α, β, and γ meet the normalization condition α+β+γ=1, ensuring that the calculation result is in the range [0,1]. The weights are configured according to the business scenario to reflect the different contributions of different dimensions to data activity; S2: Dynamically shard data based on OBE entropy The data is divided into hot shards, warm shards and cold shards. Among them, OBE> preset threshold is hot shards, OBE within the preset threshold range is warm shards, OBE≤ preset threshold is cold shards. Hot shards are deployed to the edge nodes of terminal devices, warm shards are stored in regional edge servers, and cold shards are archived to the central cloud. After the shards are stored, the system continuously monitors the changes in the OBE value of the data.

[0030] S3: Incremental synchronization trigger and failure retry When the OBE entropy change of the data exceeds the preset threshold, the shard migration mechanism is triggered. When the data is upgraded from a warm shard to a hot shard, the migration of the data to the terminal node is started, and only the fields related to the OBE change are transmitted to avoid full synchronization. When the network is interrupted, the data to be migrated is temporarily stored in the local cache and retried through UDP broadcast, triggering incremental synchronization between edge nodes and transmitting differential data through the gRPC lightweight protocol. If the synchronization fails, the reverse cache temporary storage strategy is adopted to temporarily store the data to be synchronized in the local temporary cache area and mark it as waiting for retransmission. After the network is restored, retry is performed through UDP broadcast.

[0031] Specifically: When calculating the occupational behavior entropy OBE of data in step S1, it is necessary to first comprehensively collect the access logs of different edge nodes to various types of data. These logs contain detailed access time and access data type information. By statistically analyzing the number of accesses of a certain type of data at different edge nodes within the dynamic time window, the access frequency variance is calculated, which reflects the fluctuation of data access at different nodes and can intuitively show the degree of imbalance in data distribution. For the timeliness decay rate, the system sets different decay coefficients in advance according to the characteristics of the data type. For real-time operation scoring, a higher decay coefficient is set because of its extremely high real-time requirements. Then, the decay value is calculated by combining the current time with the data generation time to reflect the value change of data over time. The calculation of multi-source correlation uses the knowledge graph to construct an occupational behavior data association network. The algorithm accurately extracts the number of associated edges between the target data and other related data and normalizes them to [0,1] interval, so that the correlation of different types of data can be uniformly compared. Finally, the values ​​of these three dimensions are weighted and calculated according to the weights α, β, and γ that are dynamically adjusted according to different business scenarios to obtain the OBE value that accurately reflects the activity level of the data in the professional scenario. Step S2 dynamically shards the data according to the OBE value. The edge computing resource pool continuously monitors the changes in the OBE value. Once the latest OBE value is obtained, it will be immediately compared with the preset threshold. If OBE>preset threshold, it is determined to be hot sharding, and the system will quickly deploy such data to the edge node of the terminal device. Because the terminal device can respond to the user's access needs for high-frequency data most quickly in actual business scenarios, the use of local cache greatly reduces the frequency of access to cloud data. When OBE is in the preset threshold interval, it is determined to be a warm shard. At this time, it is stored in the regional edge server. The regional edge server has moderate storage and processing capabilities, which can effectively cover 3-5 surrounding branch nodes, while ensuring a certain access speed and reasonably balancing storage costs. To improve access efficiency, cold shards with OBE values ​​≤ the preset threshold are archived to the central cloud. The central cloud has powerful storage capabilities, and cold shards are only retrieved on demand during queries, which is consistent with their low-frequency access characteristics. In step S3, the incremental synchronization trigger and failure retry phase, edge nodes constantly monitor changes in the data's OBE value. Once the OBE value change exceeds the preset threshold, the incremental synchronization process between edge nodes is immediately initiated. Through the gRPC lightweight protocol, only fields related to the OBE change are transmitted. The synchronization package is carefully designed and includes a unique data identifier for accurate data location, the old and new OBE values ​​so that the recipient can understand the data changes, a difference field list that lists the specific changes, and a CRC32 checksum for data integrity verification. If a network failure occurs during the synchronization process, the system immediately adopts a reverse cache temporary storage strategy, storing the data to be synchronized in a set local temporary cache area and clearly marking it as pending retransmission. Once the network returns to normal, a retry is quickly initiated via UDP broadcast to ensure the successful completion of data synchronization.

[0032] Example 2 See also Figure 1-6 , based on the first embodiment, the access frequency variance The calculation method is: statistically analyze the variance of the number of visits of a certain type of data at different edge nodes within the dynamic time window. The calculation formula is: ,in, is the total number of edge nodes, For the The number of visits to edge nodes, is the average number of visits, = ,The standard deviation is used to measure the discreteness of the access times, reflecting the distribution balance of data in edge nodes. The larger the value, the more concentrated the data access hotspot is, and the timeliness attenuation rate is The calculation method is: set the attenuation coefficient α according to the data type, calculate the current time Data generation time , the calculation formula is: =α⋅( ), time difference ( ) linear amplification attenuation effect, α controls the attenuation speed, and the calculation method of multi-source correlation degree Cr is: construct the occupational behavior data association network through the knowledge graph, extract the number of associated edges between the target data and other data, and normalize them to the [0,1] interval. The calculation formula is: Cr= , where the number of associated edges is the number of associated edges between the target data and other data in the knowledge graph, and the maximum possible number of associated edges is the maximum number of edges that may exist between all nodes in the graph. Through normalization, the number of associated edges is compressed to the [0,1] interval. The higher the value, the stronger the data correlation. Business scenarios include manufacturing training scenarios and service industry assessment scenarios. The weight values ​​of α, β, and γ are different in the manufacturing training scenarios and the service industry assessment scenarios. The edge nodes regularly recalculate the OBE entropy of the jurisdiction data and trigger the shard migration mechanism based on the calculation results to move the hot shards from the regional service The server is migrated to the terminal node, or the cold shard is archived from the terminal node to the central cloud. The difference data only contains fields related to OBE changes. The synchronization package structure contains the data unique identifier, the old and new values ​​of OBE, the difference field list, and the CRC32 check code. The local temporary cache area is set with a survival time TTL. The regional edge server covers 3-5 branch nodes in the surrounding area. The hot shard is deployed to the edge node of the terminal device, and the local cache can be used to reduce the cloud back to the source; the warm shard is stored in the regional edge server to balance the storage cost and access efficiency; the cold shard is archived to the central cloud and is only pulled on demand when queried.

[0033] Specifically: Through a specific calculation method, the variance of access frequency is accurately counted, making the measurement of data access fluctuation more scientific, providing a more accurate data basis for subsequent sharding based on OBE value, and then optimizing the cache hit rate. Different attenuation coefficients are set according to the data type to calculate the timeliness attenuation rate, which can better meet the actual timeliness requirements of various types of data, ensure the timely update of hot data and the reasonable storage of cold data, improve the rationality of data management, build an association network with the help of knowledge graph to calculate the multi-source association, and comprehensively explore the potential connections between data, so that the OBE value can comprehensively reflect the activity of data in professional scenarios, optimize the accuracy of data sharding, set different weights for different business scenarios, and make the OBE value calculation adaptable to multiple Diversified business needs are met to improve the applicability of the method in different industries. By regularly recalculating the OBE value and triggering the shard migration mechanism, it can dynamically adapt to changes in data activity, continuously optimize the data storage structure, ensure cache hit rate and synchronization efficiency, limit the difference data to only include fields related to OBE changes, and design a reasonable synchronization package structure to reduce invalid data transmission, improve synchronization efficiency and data accuracy, set the local temporary cache area lifetime TTL to avoid long-term resource occupation, ensure the rational use of system resources, clarify the coverage of regional edge servers, ensure the balance of warm shard storage and access, explain the advantages of different shard deployment locations, give full play to the characteristics of each level of storage, and improve overall data processing performance.

[0034] Example 3: Mechanical Manufacturing Factory Training Scenario The mechanical manufacturing plant includes 12 edge nodes in the welding workshop, assembly workshop, and training and assessment area, including workshop industrial computers, training terminals, and quality inspection equipment. The data that needs to be synchronized includes welder pass rates, assembly specification compliance records, and the number of safety operation violations. The core requirement of this scenario is real-time feedback on training results to avoid incorrect skill assessments due to data delays.

[0035] When calculating occupational behavior entropy (OBE), the weights were adjusted based on scenario characteristics: since the real-time operation score of welders has a great impact on subsequent practical training, the weight of the timeliness decay rate β was set to 0.4; the frequency of access to assembly specification records in different workshops varies significantly, and the number of visits to welding workshops is three times that of warehouses, so the weight of the access frequency variance α was set to 0.5; the multi-source correlation degree Cr, that is, the weight of associating violation records with safety training files γ was set to 0.1.

[0036] During the dynamic sharding stage, welders are evaluated for real-time welding quality scores, updated every 10 minutes. When the OBE value stabilizes at 0.85, it is classified as a hot shard and stored in the edge node of the industrial computer in the welding workshop. Training instructors can directly retrieve data through local terminals without accessing the cloud, and the cache hit rate is increased to 92%. The monthly skill improvement report is updated once a week with an OBE value of 0.6. It is stored as a warm shard in the factory area edge server, covering the three core workshop nodes of welding, assembly, and quality inspection. It not only meets the weekly review needs of workshop directors, but also avoids occupying terminal storage resources. The annual welder qualification certification file with an OBE value of 0.3 is archived as a cold shard in the group center cloud and is only pulled on demand during the annual review.

[0037] In the synchronization mechanism, when a welder meets the welding pass rate standard for three consecutive times, the OBE value jumps from 0.78 to 0.83, exceeding the threshold of 0.05, triggering incremental synchronization: only two difference fields, the pass rate change value and the latest assessment time, are transmitted to the assembly workshop terminal and regional server through the gRPC protocol. The synchronization packet size is only 12% of the full synchronization. If the workshop causes network packet loss due to interference from arc welding equipment, the synchronization fails. The synchronized data is temporarily stored in the temporary cache area of ​​the industrial computer with a TTL of 5 minutes. After the network is restored, it is retried through UDP broadcast to ensure that the synchronization is completed within 30 seconds to prevent the training instructor from adjusting the training plan based on old data.

[0038] Example 4: Hotel chain service industry assessment scenario A hotel chain deployed edge nodes across 30 city locations, including front desk terminals, room service PADs, and regional management servers. Data to be synchronized included room service response speed, customer complaint handling timelines, and monthly service star award results. The core pain points of this scenario were unstable store networks, bandwidth congestion during peak tourist season, and significant variations in data access requirements across stores. The data access volume in first-tier cities was five times that of stores in third-tier cities.

[0039] When calculating OBE, the weights are adapted to the characteristics of the service industry: because customer complaint data from first-tier city stores is frequently accessed, the access frequency variance weight α is set to 0.4; service response speed data decays rapidly over time, and the reference value of response records from 2 hours ago decreases, so the timeliness decay rate weight β is set to 0.4; the multi-source correlation degree Cr, the weight γ for associating complaint handling results with employee training records is set to 0.2.

[0040] During the dynamic sharding phase, front desk staff provide real-time customer satisfaction scores, updated once per order. An OBE value of 0.9 is stored as a hot shard on the store front terminal, allowing store managers to instantly view it through the local system, reducing cross-regional back-source requests by 90%. Weekly service quality summaries with an OBE value of 0.65 are stored as warm shards on the regional edge server, covering stores in four neighboring cities. When regional managers access these shards through the management platform, the response time is reduced from 2.3 seconds to 0.5 seconds. Annual outstanding employee files with an OBE value of 0.2 are archived as cold shards on the group cloud and are only retrieved through a dedicated line during annual performance appraisals.

[0041] In the synchronization mechanism, when a store experiences an unexpected escalation of customer complaints, the OBE value plummets from 0.55 to 0.48, exceeding the threshold of 0.07. This triggers a cold shard migration synchronization: Only the complaint handling progress and the responsible employee ID fields are transmitted via the gRPC protocol, reducing the synchronization package size to 8% of a full synchronization. If a sudden drop in store network bandwidth during peak tourist season causes synchronization failure, the synchronized data is temporarily stored in the front-end terminal cache. Retries are then made via UDP broadcast during the early morning hours when bandwidth is idle, avoiding peak bandwidth usage during check-in hours. This increases the synchronization success rate from 76% to 98%.

[0042] Example 5: Vocational skills training institution certification scenario A vocational skills training institution, offering 10 trades including electrician, fitter, and forklift, deployed edge nodes across five campuses, including theoretical classroom terminals, practical assessment equipment, and campus management servers. The data to be synchronized included students' practical assessment scores, theoretical exam error rates, and certificate acquisition progress. The key requirement in this scenario was to ensure real-time synchronization of certification data to avoid certificate issuance errors due to data inconsistencies.

[0043] When calculating OBE, the weight is tilted towards correlation: because practical scores are directly related to certificate application qualifications, the multi-source correlation degree Cr correlates assessment scores with application conditions, and the weight γ is set to 0.3; during the students' pre-exam sprint phase, one week before the exam, the frequency of visits to the simulation exam error rate is four times that of the usual time, so the access frequency variance weight α is set to 0.4; the timeliness of theoretical exam scores decays slowly, and their reference value is stable within one week, so the timeliness decay rate weight β is set to 0.3.

[0044] During the dynamic sharding stage, students’ real-time practical assessment results and forklift reversing and storage scores are updated once per minute. The OBE value of 0.88 is stored as a hot shard on the edge node of the practical equipment. Examiners can confirm the scores instantly through the terminal to avoid errors in the arrangement of re-examinations due to data delays; the periodic learning report is updated once every three days, and the OBE value of 0.7 is stored as a warm shard on the campus server, covering the current campus and the two adjacent campuses. When counselors review cross-campus student data, they do not need to wait for a response from the cloud; historical student certificate files, with an OBE value of 0.4, are archived to the institution’s cloud as a cold shard and are only retrieved when students re-apply for certificates.

[0045] In the synchronization mechanism, when a student's practical score improves from failing to excellent, the OBE value rises from 0.52 to 0.83, exceeding the threshold of 0.3, triggering a hot shard synchronization. The score change value and the examiner signature time difference field are transmitted via the gRPC protocol to the campus management server and the certificate application system, reducing the synchronization time from 1.8 seconds for a full synchronization to 0.3 seconds. If synchronization fails on a practical device due to Bluetooth interference, the data is temporarily stored in the device cache and automatically retried after the network is restored within 3 minutes. This ensures that the certificate application system obtains the latest scores within 10 minutes, avoiding missing eligible students.

[0046] Example 6: Logistics Park Driver Assessment Scenario A large logistics park includes 200 freight vehicles, mobile edge nodes, and three fixed edge nodes in sorting centers. Data that needs to be synchronized includes driver fatigue warnings, on-time delivery rates, and customer satisfaction scores. The core challenges of this scenario are network instability caused by vehicle movement and weak signals on highways. Furthermore, the data must be real-time, provide fatigue warnings, and minimize bandwidth consumption.

[0047] When calculating OBE, the weights are adapted to the mobile scenario: because fatigue driving warning data needs to be synchronized to the dispatch center immediately, the timeliness decay rate weight β is set to 0.5; the access volume of driver data on different routes varies greatly, and the access volume of long-distance drivers is twice that of short-distance drivers, so the access frequency variance weight α is set to 0.4; the multi-source correlation degree Cr associates punctuality with road condition training records, and the weight γ is set to 0.1.

[0048] During the dynamic sharding stage, the driver's real-time fatigue driving time is updated every 30 seconds, with an OBE value of 0.92 stored as a hot shard in the vehicle terminal and the edge node of the sorting center. The dispatcher can receive early warnings in seconds through the local system to avoid accident risks; the monthly transportation performance ranking, with an OBE value of 0.6, is stored as a warm shard in the regional edge server, covering 5 surrounding logistics sites. The fleet captain does not need to transmit across regions when reviewing it, and bandwidth consumption is reduced by 65%; the annual driver safety award file, with an OBE value of 0.25, is archived to the headquarters cloud as a cold shard and is only retrieved during the annual commendation.

[0049] In the synchronization mechanism, when a driver's fatigue driving duration exceeds the threshold, increasing from 15 minutes to 20 minutes, the OBE value rises from 0.81 to 0.93, exceeding the threshold of 0.12, triggering incremental synchronization: the real-time duration and vehicle position difference fields are transmitted to the dispatch center via the gRPC protocol. The synchronization packet contains only 128 bytes, and the full synchronization is 2.1KB. If the vehicle enters a tunnel and causes a network interruption, the synchronized data is temporarily stored in the on-board terminal cache with a TTL of 5 minutes. After exiting the tunnel, a UDP broadcast is retried to ensure that the dispatch center receives an early warning and arranges a shift change within 2 minutes. The synchronization delay is reduced from 820ms to 190ms, and fatigue driving accidents caused by data delays have been eliminated.

[0050] The working principle of the present invention is: first, a quantitative analysis is performed on the three dimensions of data access frequency variance, timeliness decay rate, and multi-source correlation, and the occupational behavior entropy (OBE) is calculated in combination with dynamic weights. This is used as a measure of data activity, and then the data is divided into hot shards, warm shards, and cold shards according to the OBE value, and deployed accordingly to the edge nodes of terminal devices, regional edge servers, and central clouds. During the data synchronization process, incremental synchronization is triggered by monitoring changes in OBE values, and the gRPC protocol is used to transmit differential data. When synchronization fails, a reverse cache temporary storage and UDP broadcast retry mechanism are used to ensure efficient and accurate data synchronization, thereby achieving efficient management and synchronization of data.

[0051] The preferred embodiments of the present invention disclosed above are only used to help illustrate the present invention. The preferred embodiments do not describe all details in detail, nor do they limit the invention to only the specific implementation methods described. This specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the present invention, so that those skilled in the art can better understand and utilize the present invention.

Claims

1. A data synchronization method based on edge-side distributed cache, characterized by: The following steps are involved: S1: Calculate the OBE of the data The OBE accesses frequency variance , time-dependent decay rate , multi-source correlation Cr three-dimensional calculation, the calculation formula is OBE=α⋅ +β⋅ +γ⋅Cr, where α+β+γ=1, and the weights of α, β, and γ are dynamically adjusted according to the business scenario. OBE quantifies data activity through three dimensions. The weight coefficients α, β, and γ meet the normalization condition α+β+γ=1, ensuring that the calculation result is in the range [0,1]. The weights are configured according to the business scenario to reflect the different contributions of different dimensions to data activity; S2: Dynamically shard the data based on the OBE entropy The data is divided into hot shards, warm shards, and cold shards. A hot shard is defined as one with an OBE greater than a preset threshold, a warm shard is defined as one within a preset threshold range, and a cold shard is defined as one with an OBE less than or equal to a preset threshold. The hot shards are deployed to the edge nodes of the terminal devices, the warm shards are stored in the regional edge servers, and the cold shards are archived to the central cloud. After the shards are stored, the system continuously monitors changes in the OBE values ​​of the data. S3: Incremental synchronization trigger and failure retry When the OBE entropy change of the data exceeds the preset threshold, the shard migration mechanism is triggered. When the data is upgraded from a warm shard to a hot shard, the migration of the data to the terminal node is started, and only the fields related to the OBE change are transmitted to avoid full synchronization. When the network is interrupted, the data to be migrated is temporarily stored in the local cache and retried through UDP broadcast, triggering incremental synchronization between edge nodes and transmitting differential data through the gRPC lightweight protocol. If the synchronization fails, the reverse cache temporary storage strategy is adopted to temporarily store the data to be synchronized in the local temporary cache area and mark it as waiting for retransmission. After the network is restored, retry is performed through UDP broadcast.

2. The data synchronization method based on edge-side distributed cache according to claim 1, characterized in that: The access frequency variance The calculation method is: statistically analyze the variance of the number of visits of a certain type of data at different edge nodes within the dynamic time window. The calculation formula is: ,in, is the total number of edge nodes, For the The number of visits to edge nodes, is the average number of visits, = ,The standard deviation is used to measure the discreteness of the access times, which reflects the distribution balance of data in edge nodes. The larger the value, the more concentrated the data access hotspot.

3. The data synchronization method based on edge-side distributed cache according to claim 1, characterized in that: The time-dependent decay rate The calculation method is: set the attenuation coefficient α according to the data type, calculate the current time Data generation time , the calculation formula is: =α⋅( ), time difference ( ) linearly amplifies the attenuation effect, and α controls the attenuation speed.

4. The data synchronization method based on edge-side distributed cache according to claim 1, characterized in that: The multi-source correlation degree Cr is calculated as follows: construct a professional behavior data association network through the knowledge graph, extract the number of associated edges between the target data and other data, and normalize it to the [0,1] interval. The calculation formula is: Cr= , where the number of associated edges is the number of associated edges between the target data and other data in the knowledge graph, and the maximum possible number of associated edges is the maximum number of edges that may exist between all nodes in the graph. Through normalization, the number of associated edges is compressed to the [0,1] interval. The higher the value, the stronger the data correlation.

5. The data synchronization method based on edge-side distributed cache according to claim 1, characterized in that: The business scenarios include manufacturing training scenarios and service industry assessment scenarios. The weight numerical ratios of α, β, and γ are different in the manufacturing training scenarios and the service industry assessment scenarios.

6. The data synchronization method based on edge-side distributed cache according to claim 1, characterized in that: The edge node periodically recalculates the OBE entropy of the governed data, and triggers the shard migration mechanism based on the calculation results, migrating hot shards from the regional server to the terminal node, or archiving cold shards from the terminal node to the central cloud.

7. The data synchronization method based on edge-side distributed cache according to claim 1, characterized in that: The difference data only includes fields related to OBE changes, and the synchronization packet structure includes a data unique identifier, old and new OBE values, a difference field list, and a CRC32 check code.

8. The data synchronization method based on edge-side distributed cache according to claim 1, characterized in that: The local temporary buffer area is set with a survival time TTL.

9. The data synchronization method based on edge-side distributed cache according to claim 1, characterized in that: The regional edge server covers 3-5 surrounding branch nodes.

10. The data synchronization method based on edge-side distributed cache according to claim 1, characterized in that: The hot shards are deployed to the edge nodes of the terminal devices, and local cache can be used to reduce cloud back-to-source. The warm shards are stored in regional edge servers, which can balance storage costs and access efficiency; the cold shards are archived to the central cloud and pulled on demand only when queried.

Citation Information

Cited By

  • Distributed partition storage security method and system for new energy truck carbon data

    CN121502821A

  • A distributed partition storage security method and system for carbon data of new energy trucks

    CN121502821B