Hierarchical distribution model warehouse synchronization system and method
Through the hierarchical distribution of the model warehouse synchronization system, the synchronization strategy is generated using model priority and topological features, which solves the problems of low efficiency, large resource consumption and poor data security in the existing technology, and achieves efficient, reliable and secure model warehouse synchronization.
Patent Information
- Application Number
- CN202510469716.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-04-15
AI Technical Summary
The existing model warehouse synchronization technology is inefficient, resource consumption is high, and it is difficult to ensure data security, especially in large-scale and complex application scenarios.
Design a hierarchical distribution model warehouse synchronization system, through the model priority calculation module, topological feature acquisition module and synchronization module, generate synchronization policies, optimize synchronization processes, reduce resource consumption, and enhance data security.
It improves the efficiency of model synchronization, reduces resource consumption, enhances the reliability and security of data synchronization, and adapts to complex distributed environments and cloud edge computing scenarios.
Smart Images

Figure CN119988503A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of model synchronization, and in particular relates to a hierarchical distribution model warehouse synchronization system and method. Background Art
[0002] With the rapid development of machine learning and deep learning technologies, enterprises and research institutions need to manage and synchronize a large number of model files. These model files are usually large in size and need to be frequently synchronized between development, testing, and production environments, and even involve cross-enterprise collaboration. Therefore, an efficient and secure model repository synchronization mechanism is crucial to ensure model version consistency, improve development efficiency, and optimize resource utilization.
[0003] At present, mainstream model repository synchronization technologies rely on version control systems (such as Git) and their large file support (such as Git LFS), or technologies based on object storage services (such as Amazon S3, Alibaba Cloud OSS). These technologies provide basic data synchronization and version management functions, allowing model files to be shared between different environments. However, with the continuous expansion of model scale and the increasing complexity of application scenarios, traditional synchronization methods face many challenges in practical applications.
[0004] First, in terms of synchronization efficiency, existing methods often do not distinguish between factors such as file size and update frequency, but use full synchronization or simple incremental updates for all model data. This method will take up a lot of bandwidth when synchronizing model data on a large scale, resulting in a slow synchronization process and may even cause network congestion. Secondly, the problem of resource consumption is also very prominent. Full or simple incremental synchronization requires a large storage space and network bandwidth, especially when synchronizing large files. Resource consumption is particularly serious, increasing the operating costs of the enterprise. In addition, when the network environment is unstable or the server load is high, the existing synchronization mechanism may cause synchronization failure, and the consistency and integrity of the model data cannot be guaranteed, thus affecting the normal development and deployment of the model.
[0005] Security and permission control are also important issues. Most current synchronization solutions focus on data transmission and storage, but lack support for access control and permission management, making it difficult to meet security requirements in a multi-tenant environment. In addition, existing solutions have limitations in flexibility and scalability, making it difficult to adapt to rapidly changing business needs, especially in emerging scenarios such as cloud computing and edge computing. Traditional synchronization mechanisms lack dynamic adjustment capabilities and cannot efficiently support complex distributed environments.
[0006] Therefore, in response to the above problems, there is an urgent need for an efficient, reliable, scalable model warehouse synchronization solution with complete security controls to improve synchronization efficiency, reduce resource consumption, enhance synchronization stability, and meet the company's strict requirements for data security. Summary of the invention
[0007] The present invention provides a hierarchical distribution model warehouse synchronization system and method to solve the problems of low efficiency, large consumption of resources and difficulty in ensuring data security in existing synchronization methods. In order to solve the above technical problems, the embodiments of the present invention disclose the following technical solutions: One aspect of the present invention provides a hierarchical distribution model repository synchronization system, which is used to synchronize a model of a source model repository to a target model repository, including: A model priority calculation module is configured to calculate the priority of each model in the source end model warehouse according to model characteristics, wherein the model characteristics at least include update frequency, access frequency, importance and type; A topology feature acquisition module is configured to obtain the priority of each target node and a transmission quality parameter between connected target nodes, wherein the target node is a server for synchronizing the model; The synchronization module is configured to generate a synchronization strategy according to the model priority, the target node priority and the transmission quality parameters between the connected nodes, and synchronize the model to the target end model warehouse.
[0008] Optionally, the model priority calculation module includes: An update frequency submodule is configured to determine update frequency parameter values of all models based on the historical update data of each model in the source end model warehouse; The access frequency submodule is configured to determine access frequency parameter values of all models based on historical access data of each model in the source model warehouse; The importance quantification submodule is configured to quantify the importance parameters of each model in the source model warehouse according to the user's input data; The type quantification submodule is configured to quantify the type parameters of each model in the source model warehouse according to the type of the model.
[0009] Optionally, the model priority calculation module further includes a priority submodule configured to calculate the priority of the model according to the following formula: in, is the update frequency parameter value; is the access frequency parameter value; is the importance parameter value; is the type parameter value; is the weight of the update frequency parameter, which is obtained by searching a preset update frequency reference table, wherein the update frequency parameter value is proportional to the weight; is the weight of the access frequency parameter, which is obtained by searching a preset access frequency reference table, wherein the access frequency parameter value is proportional to the weight; is the weight of the importance parameter, which is obtained by looking up the preset importance reference table; is the weight of the type parameter, which is found from the preset type reference table.
[0010] Optionally, the topological feature acquisition module includes: A node priority submodule is configured to update the priority of each target node according to the real-time status parameters of the target node or the input data of the user, wherein the priority includes core, common and edge; A link weight submodule, configured to determine the weight of a link between any two connected target nodes according to a transmission quality parameter, wherein the transmission quality parameter includes at least real-time bandwidth, delay, and packet loss rate; The dynamic bandwidth allocation submodule is configured to update the bandwidth of each link according to the real-time traffic and weight of the link.
[0011] Optionally, the synchronization module includes: A model grading submodule is configured to grade each model in the source end model warehouse according to the priority of the model, wherein a model with a priority greater than a first threshold is set to a first level, a model with a priority less than the first threshold and greater than a second threshold is set to a second level, and a model with a priority less than the second threshold is set to a third level; A distribution submodule is configured to transmit the first level model using a link with a bandwidth exceeding a preset bandwidth threshold and a delay below a preset delay threshold, or to fragment the first level model and transmit it through multiple target nodes; The distribution submodule is further configured to simultaneously transmit a plurality of second level models within a preset time period; The distribution submodule is further configured to transmit the third level model when the network occupancy rate is lower than a set occupancy rate threshold, or to suspend transmission of the third level model.
[0012] Optionally, the synchronization module compares the same model in the source-end model warehouse and the target-end model warehouse to obtain the changed part in the model data, and only compresses the changed part and transmits it to the target end.
[0013] Optionally, the system further includes a cleaning module configured to delete models in edge nodes whose access frequency parameter values are less than a preset cleaning threshold after each preset time period.
[0014] Optionally, the system further includes: A network environment perception module is configured to detect network bandwidth and packet loss rate, and to predict network environment data at a future time using a machine learning algorithm; The server status awareness module is configured to detect the storage space, CPU load, and I / O performance of each target node.
[0015] Optionally, the synchronization module is in communication connection with the network environment perception module and the server status perception module, and is configured to set the transmission speed according to the network bandwidth, and determine whether to retransmit according to the packet loss rate; The synchronization module is also configured to dynamically adjust the synchronization strategy according to the network environment data and the server status data.
[0016] Optionally, the system further includes: A data security module is configured to encrypt model data during transmission and to perform integrity verification on synchronized data; The permission control module is configured to match the corresponding synchronization function according to the user's identity permission.
[0017] Optionally, the system further comprises an adaptive fault-tolerant module configured to automatically detect errors during synchronization and execute a preset recovery strategy when a synchronization error, network instability or server failure occurs; The adaptive fault-tolerant module is also configured to record synchronization history logs and error logs.
[0018] Optionally, the system further includes a data receiving and merging module, which is configured to decompress and decrypt the received synchronization data at the target end, and merge the data into the target end model warehouse.
[0019] Another aspect of the present invention provides a hierarchical distribution model warehouse synchronization method, the method is applied to a hierarchical distribution model warehouse synchronization system, comprising: Calculating the priority of each model in the model warehouse according to model characteristics, wherein the model characteristics include at least update frequency, access frequency, importance, and type; Obtaining the priority of each target node and the transmission quality parameter between the connected target nodes, wherein the target node is a server for synchronizing the model; Generate a synchronization strategy based on model priority, target node priority, and transmission quality parameters between connected nodes to synchronize the model to the target model repository.
[0020] The present invention provides a hierarchical distribution model warehouse synchronization system and method, which ensures fast synchronization of key and frequently accessed data through the design of multiple synchronization strategies and the implementation of dynamic scheduling, avoiding the inefficiency caused by indiscriminate synchronization; reduces the amount of data transmission through business classification strategy and on-demand synchronization strategy, effectively reduces the occupation of network bandwidth and the demand for storage resources, thereby reducing resource consumption and reducing operating costs; ensures efficient synchronization tasks through real-time adjustment of distribution paths and load balancing, avoiding overload due to single-point synchronization. At the same time, the adaptive fault-tolerant mechanism can automatically identify and handle abnormal situations in the synchronization process (such as network instability, server failure, etc.), improving the reliability of data synchronization.
[0021] In addition, the present invention proposes an end-to-end data encryption and integrity verification mechanism to protect the security of data during transmission and avoid the risk of data being illegally stolen or tampered with during synchronization. The role-based access control mechanism ensures that only authorized users and systems can access and synchronize data, enhancing data security management.
[0022] The present invention has good scalability and can flexibly adapt to synchronization requirements of different scales and support new synchronization strategies in the future to meet the needs of large-scale expansion of synchronization tasks. While ensuring the efficiency and reliability of model data synchronization, it reduces resource consumption and strengthens data security and authority management.
[0023] This Summary is provided to introduce a selection of concepts in a simplified form that are further described below in the Detailed Description. This Summary is not intended to identify key features or essential features of the disclosure, nor is it intended to limit the scope of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] The above and other objects, features and advantages of the present disclosure will become more apparent through a more detailed description of exemplary embodiments of the present disclosure in conjunction with the accompanying drawings, wherein like reference numerals generally represent like components throughout the exemplary embodiments of the present disclosure.
[0025] Figure 1 A schematic diagram of the structure of a hierarchical distribution model warehouse synchronization system provided by an embodiment of the present invention; Figure 2 A method provided by an embodiment of the present invention Figure 1 The structural diagram of the model priority calculation module; Figure 3 A method provided by an embodiment of the present invention Figure 1 The schematic diagram of the structure of the topological feature acquisition module; Figure 4 A method provided by an embodiment of the present invention Figure 1 The structural diagram of the synchronization module; Figure 5 A flowchart of a hierarchical distribution model warehouse synchronization method provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0026] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided to make the present disclosure more thorough and complete, and to fully convey the scope of the present disclosure to those skilled in the art.
[0027] As used herein, the term "including" and its variations mean open inclusion, i.e., "including but not limited to". Unless otherwise stated, the term "or" means "and / or". The term "based on" means "based at least in part on". The terms "an example embodiment" and "an embodiment" mean "at least one example embodiment". The term "another embodiment" means "at least one additional embodiment". The terms "first", "second", etc. may refer to different or the same objects. Other explicit and implicit definitions may also be included below.
[0028] Figure 1 The present invention provides a structural diagram of a hierarchical distribution model warehouse synchronization system, which is used to synchronize the model of the source model warehouse to the target model warehouse. Figure 1 As shown, the system includes the following modules: 1. Model priority calculation module The model priority calculation module 1 is used to calculate the priority of each model in the source model warehouse according to the model characteristics, wherein the model characteristics at least include update frequency, access frequency, importance and type.
[0029] In one embodiment disclosed in the present invention, Figure 2 As shown, the model priority calculation module 1 includes an update frequency submodule 11, an access frequency submodule 12, an importance quantization submodule 13 and a type quantization submodule 14, which are respectively used to obtain the quantized values of different model features, wherein: (1) The update frequency submodule 11 is used to determine the update frequency parameter values of all models based on the historical update data of each model in the source model warehouse.
[0030] After obtaining the historical update data of each model, the update cycle is determined according to the update frequency, and the update cycle is graded. For example, a model that is updated once a day is set as level one, and the update frequency parameter value of the level one model is set to 5; a model that is updated once a week is set as level two, and the update frequency parameter value of the level two model is set to 4; a model that is updated once a month is set as level three, and the update frequency parameter value of the level three model is set to 3; a model that is updated once a quarter is set as level four, and the update frequency parameter value of the level four model is set to 2; a model with a longer update cycle is set as level five, and the update frequency parameter value of the level five model is set to 1.
[0031] (2) The access frequency submodule 12 is used to determine the access frequency parameter values of all models based on the historical access data of each model in the source model warehouse.
[0032] Standardize the number of visits to the model and assign different weights to different visit periods. For example, the weight of visits in the last 7 days is 1.5, the weight of visits in the last 30 days is 1.2, the weight of visits in the last 90 days is 1, and the weight of visits before 90 days is 0.8.
[0033] The access frequency parameter of each model in a preset time period (e.g., half a year) is calculated as follows: Access frequency parameter = (number of visits in the past 7 days × 1.5) + (number of visits in the past 30 days × 1.2) + (number of visits in the past 90 days × 1.0) + (number of visits more than 90 days ago × 0.8).
[0034] (3) The importance quantification submodule 13 is used to quantify the importance parameters of each model in the source model warehouse based on the user's input data.
[0035] Based on business needs, the impact of the model on the business is evaluated and divided into core models, general models and backup models. The core model affects the core functions of the business, is updated frequently, and has a large number of visits; the general model has a smaller impact on the business, has a medium update frequency, and has a medium number of visits; the backup model serves as a backup of the core model, with a low update frequency and a low number of visits.
[0036] The user sets the importance parameter for each model according to the above division rules. For example, the importance parameter of the core model is 10, the importance parameter of the general model is 5, and the importance parameter of the backup model is 2.
[0037] (4) The type quantification submodule 14 is used to quantify the type parameters of each model in the source model warehouse according to the type of the model.
[0038] Pre-label different types of model data, such as training models, inference models, configuration files, etc., and set different quantization values for different types.
[0039] In one embodiment disclosed in the present invention, the model priority calculation module 1 further includes a priority submodule, which can calculate the priority of the model according to the following formula: in, is the update frequency parameter value; is the access frequency parameter value; is the importance parameter value; is the type parameter value; is the weight of the update frequency parameter, which is obtained by searching a preset update frequency reference table, wherein the update frequency parameter value is proportional to the weight; is the weight of the access frequency parameter, which is obtained by searching a preset access frequency reference table, wherein the access frequency parameter value is proportional to the weight; is the weight of the importance parameter, which is obtained by looking up the preset importance reference table; is the weight of the type parameter, which is found from the preset type reference table.
[0040] - It can be adjusted according to actual business needs. For example, the weight coefficient of the target node can be set according to the business needs of the target node. The target node is the server used to synchronize the model. It can be divided into core nodes, ordinary nodes and edge nodes according to priority. Edge nodes will pay more attention to access frequency and importance. and Compared with other types of nodes, core nodes are more concerned with update frequency and type. and Higher compared to other types of nodes.
[0041] 2. Topological feature acquisition module The topology feature acquisition module 2 is used to obtain the priority of each target node and the transmission quality parameters between connected target nodes.
[0042] In one embodiment disclosed in the present invention, Figure 3 As shown, the topology feature acquisition module 2 includes a node priority submodule 21 , a link weight submodule 22 and a dynamic bandwidth allocation submodule 23 .
[0043] (1) The node priority submodule 21 is used to update the priority of each target node according to the real-time status parameters of the target node or the input data of the user.
[0044] Priorities include core, ordinary and edge. Core nodes carry critical services and have priority in obtaining bandwidth resources. Ordinary nodes allocate remaining resources to meet daily synchronization needs. Edge nodes have limited resources and only obtain critical data.
[0045] In a specific embodiment disclosed in the present invention, the node priority can be calculated using the following formula: Priority = (bandwidth usage × 0.4) + (packet loss rate × -0.3) + (CPU usage × 0.2) + (task queue length × 0.1) + (user policy weight).
[0046] The priority division rules are as follows: core nodes, priority score ≥ 0.8; ordinary nodes, 0.4 ≤ priority score < 0.8; edge nodes, priority score < 0.4.
[0047] The system recalculates priorities and adjusts resource allocation strategies periodically or under specific trigger conditions (such as bandwidth surges, task backlogs, etc.).
[0048] (2) The link weight submodule 22 is used to determine the weight of the link between any two connected target nodes according to the transmission quality parameter.
[0049] A network topology graph is constructed based on all target nodes, and the weight of the link between every two connected nodes is calculated. The weight is determined by the transmission quality parameters, which include at least real-time bandwidth, delay, and packet loss rate.
[0050] The weight of a link represents the transmission capacity of the link. The larger the weight, the higher the link quality. This submodule calculates the weight of each link by integrating multiple transmission quality parameters.
[0051] In order to adapt to changes in network status, link weights need to be adjusted dynamically. For example, bandwidth, delay, and packet loss rate are remeasured every N seconds (such as 30 seconds) and weights are updated; or, when a link experiences a sudden decrease in bandwidth, increase in delay, or increase in packet loss rate, its weight is adjusted immediately and the transmission path is reselected.
[0052] (3) A dynamic bandwidth allocation submodule 23, which is used to update the bandwidth of each link according to the real-time traffic and weight of each link.
[0053] The higher the weight, the more important the link is. The lower the weight, the more likely the link is a backup link or a secondary data transmission path. This submodule monitors link traffic in real time and adjusts bandwidth allocation strategies to automatically increase bandwidth for high-load links to ensure service stability and reduce bandwidth for low-load links to avoid resource waste. In addition, when link weights change, bandwidth is automatically reallocated to ensure that links with higher weights get more resources.
[0054] (III) Synchronization module The synchronization module 3 is used to synchronize the model to the target end model warehouse according to the model priority, the target node priority and the transmission quality parameters between the connected nodes.
[0055] In one embodiment disclosed in the present invention, Figure 4 As shown, the synchronization module 3 includes the following submodules: (1) A model grading submodule 31, which is used to grade each model in the source end model warehouse according to the priority of the model. After obtaining the priority of each model, the model with a priority greater than a first threshold is set to the first level, the model with a priority less than the first threshold and greater than a second threshold is set to the second level, and the model with a priority less than the second threshold is set to the third level.
[0056] (2) A distribution submodule 32 is used to adopt a first-level model for transmission of links with bandwidth exceeding a preset bandwidth threshold and delay below a preset delay threshold, that is, to adopt a first-level model for transmission of links with high bandwidth and low delay.
[0057] Alternatively, the first-level model can be sharded and transmitted through multiple target nodes to improve transmission efficiency. This module can cut the model file into several data blocks (such as 100MB / block), and each node is responsible for transmitting part of the data blocks, and finally splicing them into a complete model. Each data block has a unique index to ensure correct splicing. Hash checks (such as MD5 / SHA-256) can also be used to prevent data corruption.
[0058] The distribution submodule 32 is also used to simultaneously transmit multiple second-level models within a preset time period to achieve timed batch synchronization and reduce network pressure.
[0059] The distribution submodule 32 is further used to transmit the third level model when the network occupancy rate is lower than a set occupancy rate threshold, or to suspend transmission of the third level model, that is, to transmit the third level model when the network is idle, or to suspend synchronization.
[0060] In one embodiment disclosed in the present invention, the synchronization module 3 compares the same model in the source model warehouse and the target model warehouse, obtains the changed part in the model data, and compresses only the changed part and transmits it to the target end.
[0061] This module can adopt a version management method to synchronize only the changed parts of the data, reduce the amount of transmission, and improve efficiency. For example, use version control tools (such as Git) to manage model versions and only transmit incremental data (such as file changes). This can be implemented from two levels: file-level differential and block-level differential. In addition, use an efficient compression algorithm to compress incremental data to further reduce the transmission volume.
[0062] In an embodiment disclosed in the present invention, the synchronization module 3 can also transmit only specific branches or files of the model according to user requirements.
[0063] In one embodiment disclosed by the present invention, the system further includes a cleaning module. In the case where the edge node storage resources are limited, after each preset period (e.g., 30 days), the model whose access frequency parameter value is less than the preset cleaning threshold in the edge node is deleted, that is, the model data with low activity is regularly cleaned to release storage space. When the cleaned model data is requested again, the model can be retrieved.
[0064] In one embodiment disclosed in the present invention, the system further includes the following modules: (1) Network environment perception module, which is used to detect network bandwidth, dynamically adjust synchronization speed, and detect packet loss rate, and take measures such as retransmission when the packet loss rate is high. This module can also use machine learning algorithms to predict network environment data at future times to adjust synchronization strategies in advance.
[0065] This module regularly samples the network bandwidth, identifies the current available bandwidth and dynamically adjusts the synchronization speed to ensure efficient transmission without causing network congestion. In addition, it continuously tracks the data packet transmission status and calculates the packet loss rate. When the packet loss rate exceeds the threshold, it automatically triggers retransmission and attempts to change the transmission path.
[0066] This module can also predict future network status based on historical network data using time series analysis or deep learning (such as LSTM, GRU) models, so as to reduce the synchronization rate or adjust the synchronization period in advance when network congestion is expected.
[0067] (2) Server status perception module, which is used to detect the storage space, CPU load, and I / O performance of each target node to avoid insufficient storage space or resource exhaustion, thereby ensuring synchronization efficiency.
[0068] The server status awareness module is used to monitor the resource status of the target node to ensure that data synchronization does not affect the normal operation of the server. The module regularly checks the storage space to prevent synchronization failures due to insufficient space. When the storage space is close to full, it cleans up low-frequency access data or temporary files. At the same time, it monitors the CPU load to avoid resource contention caused by synchronization tasks, lowers the synchronization priority when the load is high, and continues synchronization after resources are restored. The module can also analyze the disk read and write performance in real time to prevent I / O bottlenecks, and use block asynchronous writing for large file transfers to reduce I / O waiting time.
[0069] In one embodiment disclosed in the present invention, the synchronization module 3 is in communication connection with the network environment perception module and the server status perception module, and is configured to set the transmission speed according to the network bandwidth, and determine whether to retransmit according to the packet loss rate.
[0070] The synchronization module 3 is also configured to dynamically adjust the synchronization strategy according to the network environment data and the server status data. The module can dynamically adjust the transmission speed according to the real-time network bandwidth to ensure that the synchronization is accelerated when the bandwidth is sufficient and the rate is reduced when the bandwidth is limited to reduce the impact on other services. At the same time, the module monitors the packet loss rate. When the packet loss rate is high, measures such as retransmission, error correction or link switching can be taken to ensure data integrity.
[0071] In terms of server resource management, the server status perception module detects the storage space, CPU load and I / O resources of the target node in real time to avoid synchronization failure caused by resource exhaustion. The synchronization module 3 adjusts the data synchronization strategy accordingly, such as synchronizing key data first when storage space is insufficient, or reducing the priority of synchronization tasks when the CPU load is too high, to ensure stable operation of the system.
[0072] In addition, by combining machine learning algorithms, the future network environment and server resource status are predicted, and the synchronization strategy is optimized in advance. For example, if the bandwidth is predicted to drop, the system can complete the large file transfer in advance; if the server load is expected to increase, non-urgent synchronization tasks can be postponed. Through this intelligent regulation, the synchronization module 3 can improve the efficiency, stability and reliability of data transmission.
[0073] In one embodiment disclosed in the present invention, the system further includes the following modules: (1) Data security module, which is used to encrypt model data during transmission and to perform integrity verification on synchronized data. At the same time, a fine-grained permission control system is designed to ensure the legality and security of data access and synchronization operations.
[0074] The data security module can achieve end-to-end encryption of data, ensure the security of the synchronization process, and prevent data from being illegally intercepted or tampered with during transmission. For example, by using symmetric encryption (such as AES) or asymmetric encryption (such as RSA) technology, it can ensure that only the legitimate recipient can decrypt the data, thereby improving data security.
[0075] To further ensure data integrity, the data security module can also integrate hash verification (such as SHA-256) and digital signature technology to perform integrity verification after data synchronization is completed to ensure that the data has not been tampered with. If data anomalies are found during the verification process, the module can automatically trigger a retransmission mechanism or notify the administrator to intervene manually to prevent erroneous data from affecting system operation.
[0076] (2) The permission control module is configured to match the corresponding synchronization function according to the user's identity permissions. This module uses role-based access control (RBAC) to ensure that only authorized users can perform specific synchronization tasks.
[0077] This module has designed a fine-grained permission control system to strictly limit the permissions for data access and synchronization operations. Through the role-based access control (RBAC) mechanism, it ensures that different users can only perform synchronization operations that are within their permission range, thereby improving the security and manageability of the system. This module strictly limits access to data and system resources based on user identity, role level and permission configuration to prevent unauthorized operations or malicious tampering of data.
[0078] Under the RBAC mechanism, users are assigned to different roles, each of which corresponds to a specific set of permissions. For example, the system administrator has the highest permissions and can manage all synchronization tasks, adjust synchronization policies, and assign permissions; ordinary users can only access and synchronize model data that has been granted permissions, and guest users may only have limited read-only permissions. In addition, the module supports a permission inheritance mechanism based on the organizational structure, for example, department heads can manage the synchronization permissions of their team members, while ordinary employees can only access the model data they are responsible for.
[0079] To further enhance security, the permission control module also combines multi-factor authentication (MFA) and behavior analysis technology. MFA requires users to provide additional authentication (such as SMS verification code, dynamic password) when performing sensitive operations (such as modifying synchronization rules and adjusting node priorities). Behavior analysis technology can be used to detect abnormal operations, such as frequently triggering large-scale synchronization in a short period of time or attempting to access unauthorized data. Once an abnormality is found, the module can automatically restrict access or trigger a security alarm.
[0080] In addition, the module provides detailed permission logs and audit functions to record all synchronization requests and permission changes, making it easier for administrators to conduct security audits and problem tracking. Combined with RBAC mechanism, multi-factor authentication and log audit, the permission control module can effectively protect the security and compliance of the synchronization system and ensure that data flows efficiently and securely within a controlled range.
[0081] In one embodiment disclosed in the present invention, the system also includes an adaptive fault-tolerant module configured to automatically detect errors in the synchronization process and execute a preset recovery strategy when a synchronization error, network instability or server failure occurs.
[0082] This module has an adaptive error recovery mechanism, including automatic retry, failure fallback and other strategies, to deal with abnormal situations such as network instability and server failure. For example, it automatically detects errors in the synchronization process and adopts corresponding recovery strategies, such as retry, switching to backup channels, etc.
[0083] This module can monitor in real time various types of failures that may occur during the synchronization process, such as data transmission interruption, network fluctuations, server downtime or storage anomalies, and dynamically select the optimal fault-tolerant strategy based on the type and severity of the failure.
[0084] When a synchronization error is detected, the module first performs fault diagnosis and analyzes the source of the error. For example, a high network packet loss rate may indicate an unstable link, while insufficient storage space on the target node may cause data writing failure. The module uses a layered recovery mechanism for different types of failures. For example, when a short network jitter causes packet loss, an automatic retransmission mechanism can be used for rapid recovery; if the synchronization process is interrupted for a long time, you can switch to a backup link or renegotiate the synchronization strategy to reduce the scope of impact. For server failures, the module supports a master-slave switching mechanism. When a target node is detected to be unavailable, the synchronization task is automatically migrated to an available backup node to ensure that the task is not interrupted.
[0085] The adaptive fault-tolerant module is also configured to record synchronization history logs and error logs, and provide fault diagnosis and early warning functions to reduce manual intervention.
[0086] During the data synchronization process, the module will record the detailed historical information of synchronization in real time, including synchronization time, data volume, target node, network status and other parameters, and archive them for storage. These historical data can be used for subsequent analysis and optimization, providing a reliable reference for the system. In addition, for all synchronization errors that occur, the module will generate detailed error logs, recording the time of the failure, the scope of impact, the specific cause, and the recovery measures. These logs can not only be used by technicians for troubleshooting, but also as training data for machine learning models to help the system continuously optimize fault prediction and recovery strategies.
[0087] In order to reduce manual intervention, the module also integrates intelligent fault diagnosis functions. When an exception occurs during the synchronization process, the system automatically analyzes the error log and combines historical data to determine the root cause of the failure. For example, if the disk space of a target node continues to approach the upper limit, the module will issue an early warning to prompt the administrator to clean up the storage or expand the capacity; if the packet loss rate of a link increases abnormally, it will be inferred that there may be network congestion or hardware failure, and the synchronization path will be automatically adjusted.
[0088] In addition, the module supports a real-time early warning mechanism. Based on preset alarm thresholds, such as high packet loss rate, long unfinished synchronization tasks, and excessive server load, the system will automatically send notifications to operation and maintenance personnel to remind them of potential risks and provide possible solutions. For example, when the server CPU load is close to the limit, the module can suggest reducing the synchronization rate or adjusting the synchronization task allocation to avoid system crashes.
[0089] In an embodiment disclosed in the present invention, the system also includes a data receiving and merging module, which is configured to decompress and decrypt the received synchronization data at the target end and merge it into the target end model warehouse.
[0090] First, the module receives the synchronization data from the source end at the target end and performs data decompression. During the data transmission process, in order to improve bandwidth utilization and transmission efficiency, the system may compress the data. Therefore, the data receiving and merging module needs to decompress the received compressed data to restore the original data structure and ensure the accuracy of subsequent processing.
[0091] Secondly, the module performs decryption processing on the data. During the data transmission process, the system uses an end-to-end encryption mechanism to ensure data security and prevent unauthorized access or tampering. Therefore, the target end needs to use a security key to decrypt the data to restore the readability and integrity of the synchronized data. This process is usually combined with an access control mechanism to ensure that only nodes or users with corresponding permissions can successfully decrypt the data.
[0092] After decompression and decryption, this module is responsible for data merging and integrating the received model data with the existing model warehouse on the target end. This process involves data version management, duplicate data checking, and consistency verification to ensure the correctness of the data. For example, the system can use a version control mechanism to decide whether to overwrite, append, or create a new version based on the model version information; at the same time, hash verification and other methods can be used to detect data integrity to prevent damage or tampering during transmission.
[0093] In addition, the module can also support parallel merging and distributed processing, especially in large-scale model synchronization scenarios, to improve data integration efficiency. For fragmented data transmission, the module will reassemble multiple data fragments into a complete model according to the preset merging strategy to ensure the correctness and availability of the data.
[0094] Through the processing of the data receiving and merging module, the target end can seamlessly connect and synchronize data to ensure the continuous updating of the model warehouse, while maintaining the consistency and security of the data, thereby improving the stability and reliability of the system.
[0095] In one embodiment disclosed in the present invention, the system can be structured according to the following aspects: (1) User interface layer Users initiate model synchronization requests and manage synchronization tasks through this interface.
[0096] (2) Application service layer Permission management service: handles user permission verification to ensure that only authorized users can initiate synchronization tasks.
[0097] Synchronization scheduling service: responsible for receiving synchronization requests and scheduling synchronization tasks according to the grading strategy and network conditions.
[0098] Synchronization policy service: implements data synchronization strategies, dynamically calculates synchronization priorities and server resource allocation.
[0099] (3) Data processing layer Data compression and encryption service: compress and encrypt differential data in preparation for transmission.
[0100] Data decryption and decompression service: After receiving the data, it is decrypted and decompressed back to its original state.
[0101] Model data merging service: merges the received data into the target warehouse.
[0102] (4) Network transport layer Transport service: responsible for the secure transmission of data, including the implementation of retry and failure fallback mechanisms.
[0103] Monitoring service: realize server and network environment status monitoring Log service: stores logs generated during the synchronization process, including task status, error records, etc.
[0104] In one embodiment disclosed in the present invention, the synchronization system can complete the task according to the following process: (1) User initiates synchronization request: The user initiates a model data synchronization request through the user interface layer.
[0105] (2) Permission verification: The permission management service in the application service layer verifies the user's requested permissions.
[0106] (3) Request scheduling: The scheduling service determines the priority of the synchronization task and the allocated network bandwidth based on the classification strategy and the current network conditions.
[0107] (4) Confirm the synchronization data: Compare the source warehouse with the target warehouse to determine the incremental data that needs to be synchronized.
[0108] (5) Data processing: Data compression and encryption services process differential data and perform compression and encryption.
[0109] (6) Secure transmission: Encrypted data is securely transmitted to the target warehouse through the network transport layer.
[0110] (7) Data reception and merging: The data decryption and decompression service of the synchronization target layer processes the received data, and the model data merging service merges it into the target model repository.
[0111] (8) Status feedback: After synchronization is completed, the system will feedback the synchronization results to the user.
[0112] Figure 5 A flowchart of a hierarchical distribution model warehouse synchronization method provided by the present invention is shown as follows: Figure 5 As shown, the method comprises the following steps: Step S100: Calculate the priority of each model in the model warehouse according to the model characteristics, where the model characteristics at least include update frequency, access frequency, importance and type.
[0113] Step S200: obtaining the priority of each target node and the transmission quality parameters between connected target nodes, where the target node is a server for synchronizing the model.
[0114] Step S300: Generate a synchronization strategy according to the model priority, the target node priority and the transmission quality parameters between the connected nodes, and synchronize the model to the target end model warehouse.
[0115] The embodiments of the present disclosure have been described above, and the above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and changes will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The selection of terms used herein is intended to best explain the principles of the embodiments, practical applications, or technical improvements in the market, or to enable other persons of ordinary skill in the art to understand the embodiments disclosed herein.
Claims
1. A hierarchical distribution model warehouse synchronization system, used to synchronize the model of the source end model warehouse to the target end model warehouse, characterized in that: include: A model priority calculation module is configured to calculate the priority of each model in the source end model warehouse according to model characteristics, wherein the model characteristics at least include update frequency, access frequency, importance and type; A topology feature acquisition module is configured to obtain the priority of each target node and a transmission quality parameter between connected target nodes, wherein the target node is a server for synchronizing the model; The topological feature acquisition module includes: A node priority submodule is configured to update the priority of each target node according to the real-time status parameters of the target node or the input data of the user, wherein the priority includes core, common and edge; A link weight submodule, configured to determine the weight of a link between any two connected target nodes according to a transmission quality parameter, wherein the transmission quality parameter includes at least real-time bandwidth, delay, and packet loss rate; A dynamic bandwidth allocation submodule, configured to update the bandwidth of each link according to the real-time traffic and weight of the link; The synchronization module is configured to generate a synchronization strategy according to the model priority, the target node priority and the transmission quality parameters between the connected nodes, and synchronize the model to the target end model warehouse.
2. The synchronization system according to claim 1, characterized in that: The model priority calculation module includes: An update frequency submodule is configured to determine update frequency parameter values of all models based on the historical update data of each model in the source end model warehouse; The access frequency submodule is configured to determine access frequency parameter values of all models based on historical access data of each model in the source model warehouse; The importance quantification submodule is configured to quantify the importance parameters of each model in the source model warehouse according to the user's input data; The type quantification submodule is configured to quantify the type parameters of each model in the source model warehouse according to the type of the model.
3. The synchronization system according to claim 2, characterized in that: The model priority calculation module also includes a priority submodule, which is configured to calculate the priority of the model according to the following formula: in, is the update frequency parameter value; is the access frequency parameter value; is the importance parameter value; is the type parameter value; is the weight of the update frequency parameter, which is obtained by searching a preset update frequency reference table, wherein the update frequency parameter value is proportional to the weight; is the weight of the access frequency parameter, which is obtained by searching a preset access frequency reference table, wherein the access frequency parameter value is proportional to the weight; is the weight of the importance parameter, which is obtained by looking up the preset importance reference table; is the weight of the type parameter, which is found from the preset type reference table.
4. The synchronization system according to claim 1, characterized in that: The synchronization module comprises: A model grading submodule is configured to grade each model in the source end model warehouse according to the priority of the model, wherein a model with a priority greater than a first threshold is set to a first level, a model with a priority less than the first threshold and greater than a second threshold is set to a second level, and a model with a priority less than the second threshold is set to a third level; A distribution submodule is configured to transmit the first level model using a link with a bandwidth exceeding a preset bandwidth threshold and a delay below a preset delay threshold, or to fragment the first level model and transmit it through multiple target nodes; The distribution submodule is further configured to simultaneously transmit a plurality of second level models within a preset time period; The distribution submodule is further configured to transmit the third level model when the network occupancy rate is lower than a set occupancy rate threshold, or to suspend transmission of the third level model.
5. The synchronization system according to claim 1, characterized in that: The synchronization module compares the same model in the source model warehouse and the target model warehouse, obtains the changed part in the model data, and compresses only the changed part and transmits it to the target end.
6. The synchronization system according to claim 4, characterized in that: The system also includes a cleaning module configured to delete models in edge nodes whose access frequency parameter values are less than a preset cleaning threshold after each preset time period.
7. The synchronization system according to claim 1, characterized in that: The system further comprises: A network environment perception module is configured to detect network bandwidth and packet loss rate, and to predict network environment data at a future time using a machine learning algorithm; The server status awareness module is configured to detect the storage space, CPU load, and I / O performance of each target node.
8. The synchronization system according to claim 7, characterized in that: The synchronization module is in communication with the network environment perception module and the server status perception module, and is configured to set the transmission speed according to the network bandwidth, and determine whether to retransmit according to the packet loss rate; The synchronization module is also configured to dynamically adjust the synchronization strategy according to the network environment data and the server status data.
9. The synchronization system according to claim 1, characterized in that: The system further comprises: A data security module is configured to encrypt model data during transmission and to perform integrity verification on synchronized data; The permission control module is configured to match the corresponding synchronization function according to the user's identity permission.
10. The synchronization system according to claim 1, characterized in that: The system also includes an adaptive fault-tolerant module configured to automatically detect errors during synchronization and execute a preset recovery strategy when a synchronization error, network instability, or server failure occurs; The adaptive fault-tolerant module is also configured to record synchronization history logs and error logs.
11. The synchronization system according to claim 1, characterized in that: The system also includes a data receiving and merging module, which is configured to decompress and decrypt the received synchronization data at the target end and merge it into the target end model warehouse.
12. A hierarchical distribution model warehouse synchronization method, characterized in that: The method uses the system described in any one of claims 1 to 11, including: Calculating the priority of each model in the model warehouse according to model characteristics, wherein the model characteristics include at least update frequency, access frequency, importance, and type; Obtain the priority of each target node and the transmission quality parameters between connected target nodes, including: According to the real-time status parameters of the target node or the input data of the user, the priority of each target node is updated, and the priority includes core, common and edge; the target node is a server for synchronizing the model; Determine the weight of a link between any two connected target nodes according to transmission quality parameters, wherein the transmission quality parameters include at least real-time bandwidth, delay, and packet loss rate; Updating the bandwidth of each link according to the real-time traffic and weight of the link; Generate a synchronization strategy based on model priority, target node priority, and transmission quality parameters between connected nodes to synchronize the model to the target model repository.
Citation Information
Patent Citations
Distributed data encryption transmission system
CN117955749A
Cross-platform game data synchronization system
CN118041933A
Informatization machine room monitoring and management system
CN118400314A
Data transmission and updating system and method of cloud mobile phone
CN119402886A
Dynamic switching distributed file storage system based on elastic data network
CN119788692A
Cited By
Data security disaster recovery method and system of data center
CN120596869A
A data security disaster recovery method and system for a data center
CN120596869B
Large file distribution method and device, equipment, medium and product
CN120935160A
AI asset block-level storage multiplexing method, system and equipment based on Git protocol
CN122363617A