Enterprise multi-source data real-time synchronization system based on data center
Through the enterprise multi-source data real-time synchronization system based on the data middle platform, the containerized deployment and hierarchical architecture is used to solve the problems of insufficient guarantee of data source heterogeneity, real-time guarantee and difficult to ensure data consistency, and efficient and reliable data synchronization and consistency guarantee are achieved.
Patent Information
- Application Number
- CN202411936041.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-26
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2044-12-26
AI Technical Summary
The existing data synchronization solutions have problems such as data source heterogeneity, insufficient real-time guarantee, difficult to guarantee data consistency, and low system coordination, resulting in large data synchronization delays, prominent performance bottlenecks, and high operation and maintenance costs.
The enterprise multi-source data real-time synchronization system based on the data middle platform is adopted, and through containerized deployment and hierarchical architecture, including the data access layer, data processing layer, data distribution layer and data control layer, the data access adapter module, data format processing module, data routing distribution module and consistency guarantee module are used to realize unified access, real-time data routing and consistency guarantee of multi-source heterogeneous data.
It realizes the reduction of data synchronization delay, improves system throughput, and effectively guarantees data consistency, which is suitable for data synchronization and cross-system data integration scenarios of enterprise core business systems.
Smart Images

Figure CN120045619A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of data processing and distributed systems, and in particular, relates to a real-time synchronization system for enterprise multi-source data based on a data middle platform. Background Art
[0002] As enterprises’ digital transformation deepens, the number of internal systems continues to increase, and data synchronization requirements become increasingly complex. Existing data synchronization solutions have the following main problems:
[0003] 1. Heterogeneous data sources:
[0004] There are many different types of data source systems within the enterprise, the data format and interface standards are not unified, the data synchronization adaptation development cost is high, and the system expansion and maintenance are difficult.
[0005] 2. Real-time guarantee issues:
[0006] The traditional batch synchronization method has large delays, the real-time synchronization mechanism is not perfect, the performance bottleneck problem is prominent, and it is difficult to meet the real-time requirements of the business.
[0007] 3. Data consistency issues:
[0008] It is difficult to ensure data consistency in a distributed environment. Data is prone to inconsistency under abnormal circumstances. The data verification mechanism is imperfect, and the repair cost is high and the efficiency is low.
[0009] 4. System coordination issues:
[0010] There is insufficient integration with the data center, lack of unified management and control, imperfect security mechanisms, and high operation and maintenance costs. Summary of the invention
[0011] The purpose of the present invention is to provide an enterprise multi-source data real-time synchronization system based on a data middle platform to solve the real-time synchronization problem of multi-source heterogeneous data within an enterprise.
[0012] The present invention provides an enterprise multi-source data real-time synchronization system based on a data middle platform, which adopts a containerized deployment method and a layered architecture, including a data access layer, a data processing layer, a data distribution layer and a data control layer; the data access layer communicates with the data middle platform and obtains original data from the data middle platform;
[0013] The data access layer is provided with a data access adapter module, the data processing layer is provided with a data format processing module, the data distribution layer is provided with a data routing distribution module, and the data control layer is provided with a consistency guarantee module; each layer communicates through a standardized interface to ensure the scalability and maintainability of the system;
[0014] The data access adapter module is used for unified access of multi-source heterogeneous data, including a relational database adapter unit, a distributed database adapter unit, a data warehouse adapter unit and a search engine adapter unit, which are respectively used to access a relational database, a distributed database, a data warehouse and a search engine;
[0015] The data format processing module is used to perform data standardization processing on the accessed multi-source heterogeneous data, and send the processed standardized data to the data routing distribution module;
[0016] The data routing distribution module is used for:
[0017] During the data routing process, the load status of each node is monitored in real time, and the ARIMA model is used to perform dynamic load forecasting in combination with historical data;
[0018] Based on the load prediction results, an adaptive routing strategy is adopted to dynamically distribute the data flow to the optimal node, and a flow control mechanism based on message middleware is adopted to prevent system overload and ensure the stability of data processing; wherein, the adaptive routing strategy considers node load, network status, and data affinity factors, and performs optimal route selection based on the ant colony algorithm to ensure the balance and efficiency of data distribution;
[0019] The consistency assurance module is used for:
[0020] Capture change information of source data in real time, adopt a multi-level consistency verification method based on hash verification, and regularly verify the data on the source and target ends based on the established data consistency verification points. If data anomalies are found, data repair is automatically performed to ensure the consistency of data synchronization, and through the distributed transaction coordination mechanism, ensure the atomicity and consistency of data operations in a distributed environment.
[0021] Furthermore, the data access adapter module adopts a plug-in structure to support dynamic loading and management, as well as personalized extension of specific data sources; the data access adapter module includes an adapter management unit for unified management of the life cycle of various adapters to achieve dynamic deployment and upgrading of adapters.
[0022] Furthermore, the data format processing module includes:
[0023] A data format identification unit, used for identifying a data format type;
[0024] Data cleaning unit, used to clean abnormal data;
[0025] A data conversion unit, used for data format conversion;
[0026] Data standardization unit, used to unify data formats;
[0027] Quality control unit for data quality checking.
[0028] Furthermore, the steps of performing optimal route selection based on the ant colony algorithm are as follows:
[0029] Step 1: Initialization:
[0030] The initial pheromone concentration is set to a constant value (τ ij (0) = τ0); set to the initial value (τ 0 ); Initialize system status: collect the current load of each node (L j ) and path bandwidth (B ij ); Step 2: Calculate path weight:
[0031] Comprehensive weight formula:
[0032]
[0033] in:
[0034] (L j ): the load of node (j), the lower the better;
[0035] (B ij ): Path (P ij )’s remaining bandwidth, the higher the better;
[0036] (α, β): weight factors that control the impact of load and bandwidth on the decision;
[0037] Step 3: Path selection:
[0038] Selection probability formula:
[0039]
[0040] in:
[0041] τ ij (t): path (P ij )'s pheromone concentration;
[0042] (W ij ): Path (P ij )’s comprehensive weight;
[0043] (η, γ): parameters that control the importance of pheromone and path weight, respectively;
[0044] Step 4: Pheromone Update:
[0045] Update formula:
[0046] τ ij (t+1)=(1-ρ)τij (t)+Δτ ij ;
[0047] in:
[0048] (1-ρ): pheromone volatilization factor, which prevents pheromone from growing indefinitely;
[0049] (Δτ ij ): Newly added pheromone, indicating the path quality.
[0050] Furthermore, the multi-level consistency verification method based on hash verification is as follows:
[0051] 1) Generate a hash value for the transmitted data;
[0052] 2) The hash value is calculated again after the target end receives the data;
[0053] 3) Compare the hash values of the source and target ends;
[0054] 4) If the data is inconsistent, the repair process is triggered to accurately locate and repair the abnormal data by tracing back the historical change records.
[0055] Furthermore, the synchronization system also includes a performance optimization module, and the performance optimization module includes:
[0056] An incremental recognition unit, used to recognize incremental changes in data;
[0057] Parallel processing unit, used for distributing data processing tasks to multiple nodes for parallel execution based on the parallel processing mechanism of the distributed computing framework;
[0058] Data compression unit, used for data transmission compression optimization;
[0059] Resource scheduling unit, which implements resource scheduling based on container technology;
[0060] Performance monitoring unit, used for system performance monitoring.
[0061] Furthermore, the incremental identification unit is specifically used for:
[0062] The Bloom filter records the synchronized data and only identifies the newly added or changed data to reduce the amount of data transmission. The specific process is as follows:
[0063] 1) Initialize the Bloom filter and load the source data hash;
[0064] 2) Calculate the hash value for the new data and check whether the Bloom filter already exists;
[0065] 3) Add the new data to the synchronization task.
[0066] Furthermore, the data compression unit is specifically used for:
[0067] Based on the efficient compression algorithm of LZ4, an intelligent data compression strategy is implemented. The specific process is as follows:
[0068] 1) Process the input data stream in blocks to extract repeated patterns;
[0069] 2) Match data blocks through sliding windows and replace them with pointers or symbols;
[0070] 3) Transmit the compressed data and decompress it on the target end.
[0071] Furthermore, the resource scheduling unit is specifically used for:
[0072] Based on the containerized resource scheduling mechanism, the synchronization system can dynamically adjust the allocation of computing resources according to the load conditions to ensure the optimization of processing performance.
[0073] Furthermore, the synchronization system further includes a safety control module, and the safety control module includes:
[0074] Identity authentication unit, used for unified identity authentication;
[0075] Access control unit, used to implement role-based access control;
[0076] Data encryption unit, used to implement transport layer security encryption;
[0077] Data desensitization unit, used to process sensitive data;
[0078] Audit log unit, used to record data access logs.
[0079] Through the above solution, through the enterprise multi-source data real-time synchronization system based on the data middle platform, the data synchronization delay is reduced, the system throughput is improved, and the data consistency is effectively guaranteed. It is particularly suitable for data synchronization and cross-system data integration scenarios of enterprise core business systems. The specific technical effects are as follows:
[0080] 1. The system adopts containerized deployment mode, supports elastic scaling, and realizes reliable operation of the system through a unified service governance framework. Each layer communicates through standardized interfaces to ensure the scalability and maintainability of the system. A layered design architecture is adopted. At the access layer, unified access to multi-source heterogeneous data is achieved through enterprise-level data access adapters. At the processing layer, unified data processing is achieved through data format standardization processing modules. At the distribution layer, an intelligent routing mechanism based on a stream processing framework is adopted to achieve efficient data distribution. At the control layer, the reliability of real-time data synchronization is ensured through a distributed consistency guarantee mechanism.
[0081] 2. The system uses a data access adaptation mechanism to implement a unified data access adaptation framework, supporting the access of multiple heterogeneous data sources. The system provides standardized adapter interface specifications, supporting the access of multiple types of data sources such as relational databases (such as MySQL), distributed databases (such as HBase), data warehouses (such as Hive), search engines (such as Elasticsearch), etc. The adapter adopts a plug-in design and supports dynamic loading and management. Each adapter implements standard functions such as data reading, format conversion, and state maintenance, and supports personalized extensions for specific data sources. The system manages the life cycle of various adapters in a unified manner through the adapter management unit to achieve dynamic deployment and upgrade of adapters.
[0082] 3. The system uses an intelligent data routing mechanism based on a stream processing framework to implement dynamic load prediction and adaptive routing strategies. In the data routing process, understanding the real-time load of each node helps to distribute tasks more efficiently, thereby avoiding overload. The system monitors the load status of each node in real time, combines historical data analysis, and implements dynamic load prediction based on the ARIMA model. Based on the prediction results, the system adopts an adaptive routing strategy to dynamically distribute data streams to the optimal node. The routing strategy takes into account multiple factors such as node load, network conditions, and data affinity to ensure the balance and efficiency of data distribution. The system adopts the optimal routing selection based on the ant colony algorithm to solve the problem of decreased efficiency caused by node load, network bandwidth, and other issues in data routing, and improves the overall performance of the system. The system adopts a flow control mechanism based on message middleware to prevent system overload and ensure the stability of data processing.
[0083] 4. The system has established a complete distributed consistency guarantee mechanism, using database change capture (CDC) technology for multi-level data verification to capture source data change information in real time. The system uses multi-level consistency verification based on hash verification to quickly discover data anomalies. When data inconsistencies are found, the system automatically triggers the data repair process. At the same time, the system uses a distributed transaction coordination mechanism to ensure the atomicity and consistency of data operations in a distributed environment.
[0084] 5. The system implements a high-performance data processing mechanism to ensure the real-time performance of the system. At the data collection level, the system implements incremental data recognition based on Bloom filters, and only synchronizes the changed data, which greatly reduces the amount of data transmission, solves the problem of high cost of full data synchronization, and significantly reduces the computing and transmission costs. At the data processing level, the system adopts a parallel processing mechanism based on a distributed computing framework to distribute data processing tasks to multiple nodes for parallel execution, providing high-performance data processing capabilities. The system implements an intelligent data compression strategy based on the efficient compression algorithm of LZ4, which reduces network bandwidth usage while ensuring performance. Through the containerized resource scheduling mechanism, the system can dynamically adjust the allocation of computing resources according to the load conditions to ensure the optimization of processing performance.
[0085] 6. The system has established a complete security control mechanism to achieve secure protection of data transmission and access. The system implements refined management of user permissions through unified identity authentication (SSO) and role-based access control (RBAC). At the data transmission level, the Transport Layer Security Protocol (TLS) is used for encryption to ensure the security of data transmission. The system supports flexible data desensitization strategies, and different desensitization rules can be configured for different data fields to protect the security of sensitive data. At the same time, the system implements a complete audit log mechanism to record all data access and operation behaviors, and support subsequent security audits and tracing.
[0086] The above description is only an overview of the technical solution of the present invention. In order to more clearly understand the technical means of the present invention and implement it according to the contents of the specification, the following is a detailed description of the preferred embodiments of the present invention in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0087] Figure 1 This is the overall architecture diagram of the enterprise multi-source data real-time synchronization system of the present invention;
[0088] Figure 2 It is a structural and functional schematic diagram of the data access adapter module of the present invention;
[0089] Figure 3 The flowchart of the data routing distribution module of the present invention;
[0090] Figure 4 This is a schematic diagram of the structure and function of the consistency assurance module of the present invention;
[0091] Figure 5 It is a structural and functional schematic diagram of the performance optimization module of the present invention;
[0092] Figure 6 It is a schematic diagram of the structure and function of the safety control module of the present invention. DETAILED DESCRIPTION
[0093] The specific implementation of the present invention is further described in detail below in conjunction with the accompanying drawings and examples. The following examples are used to illustrate the present invention, but are not intended to limit the scope of the present invention.
[0094] This embodiment provides an enterprise multi-source data real-time synchronization system based on a data middle platform, which adopts a containerized deployment method and a layered architecture, including a data access layer, a data processing layer, a data distribution layer and a data control layer; the data access layer communicates with the data middle platform and obtains original data from the data middle platform; the data access layer is provided with a data access adapter module, the data processing layer is provided with a data format processing module, the data distribution layer is provided with a data routing distribution module, and the data control layer is provided with a consistency assurance module; the layers communicate through standardized interfaces to ensure the scalability and maintainability of the system.
[0095] The data access adapter module is used for unified access of multi-source heterogeneous data, including a relational database adapter unit, a distributed database adapter unit, a data warehouse adapter unit and a search engine adapter unit, which are respectively used to access relational databases, distributed databases, data warehouses and search engines. The data access adapter module adopts a plug-in structure, supports dynamic loading and management, and personalized extension of specific data sources; the data access adapter module includes an adapter management unit, which is used to uniformly manage the life cycle of various adapters and realize dynamic deployment and upgrade of adapters.
[0096] The data format processing module is used to perform data standardization processing on the accessed multi-source heterogeneous data, and send the processed standardized data to the data routing distribution module. Specifically, the data format processing module includes:
[0097] A data format identification unit, used for identifying a data format type;
[0098] Data cleaning unit, used to clean abnormal data;
[0099] A data conversion unit, used for data format conversion;
[0100] Data standardization unit, used to unify data formats;
[0101] Quality control unit for data quality checking.
[0102] The data routing distribution module is used for:
[0103] During the data routing process, the load status of each node is monitored in real time, and the ARIMA model is used to perform dynamic load forecasting in combination with historical data. The full name of the ARIMA model is the Auto-Regressive Integrated Moving Average model. The ARIMA model consists of three parts, namely autoregression (AR), difference (I) and moving average (MA). Generally, the ARIMA model is mainly used to predict time series data and is suitable for data with the following characteristics:
[0104] 1. The data has certain trends or periodicities.
[0105] 2. There is time dependency between data.
[0106] 3. The data can become stationary after difference processing.
[0107] It is widely used in load forecasting, economic indicator analysis and other fields. Its parameters can reflect historical data, CPU usage, memory utilization and other related data. The prediction model is as follows:
[0108] y t =c+φ 1 y t-1 +φ 2 y t-2 +…+φ p y t-p +∈ t ;
[0109] Where yt is the predicted value at the current time t; c is a constant term, which represents the long-term mean trend of the predicted data; φ1, φ2, …, φp: are autoregressive coefficients, which are used to measure the impact of the values at the past p time points on the current value; ∈t is a random disturbance term, or error. In this embodiment, there are several possible scenarios: sudden events in system operation (for example, high concurrent requests caused by users submitting a large number of queries at the same time), random interference caused by the external environment (such as network, hardware); p is the autoregressive order, which represents the relationship between the current value yt and the data values yt-1, yt-2, …, yt-p at the previous p time points.
[0110] In load forecasting, indicators such as CPU usage and memory utilization can be used as time series input data (i.e., the time series of yt). Through ARIMA modeling, the data trends, fluctuation patterns, and periodicity of these indicators can be captured, thereby realizing the prediction of load at future time points. Among them:
[0111] Historical data (yt-1, yt-2, …): includes past CPU usage and memory utilization data, which is used to build time series models.
[0112] Autoregressive coefficients (φ1, φ2, ...): reflect the short-term trend of each indicator. For example, the current value of CPU usage may be strongly affected by the value at some point in the past.
[0113] Random error (∈t): It refers to the part of the load forecasting process that cannot be explained by historical data, such as instantaneous irregular burst load.
[0114] Based on the load prediction results, an adaptive routing strategy is adopted to dynamically distribute the data flow to the optimal node, and a flow control mechanism based on message middleware is adopted to prevent system overload and ensure the stability of data processing; wherein, the adaptive routing strategy takes into account node load, network conditions, and data affinity factors, and performs optimal route selection based on the ant colony algorithm to ensure the balance and efficiency of data distribution.
[0115] The ant colony algorithm simulates the foraging behavior of ants and uses pheromone concentration to guide path selection. The specific steps are as follows:
[0116] Step 1: Initialization:
[0117] The initial pheromone concentration is set to a constant value (τ ij (0) = τ0). Set to the initial value (τ 0 ).
[0118] Initialize system status: collect the current load (L j ) and path bandwidth (B ij ).
[0119] Step 2: Calculate path weight:
[0120] Comprehensive weight formula:
[0121]
[0122] in:
[0123] (L j ): The load of node (j) (the lower the better).
[0124] (B ij ): Path (P ij )’s remaining bandwidth (the higher the better).
[0125] (α, β): weight factors that control the impact of load and bandwidth on the decision.
[0126] Step 3: Path selection:
[0127] Selection probability formula:
[0128]
[0129] in:
[0130] τ ij (t): path (P ij )'s pheromone concentration.
[0131] (W ij ): Path (P ij )’s comprehensive weight.
[0132] (η, γ): parameters that control the importance of pheromone and path weight respectively.
[0133] Step 4: Pheromone Update:
[0134] Update formula:
[0135] τ ij (t+1)=(1-ρ)τ ij (t)+Δτ ij ;
[0136] in:
[0137] (1-ρ): Pheromone volatilization factor, which prevents pheromones from growing infinitely.
[0138] (Δτ ij ): Newly added pheromone, indicating the path quality.
[0139] In a specific example, the data routing distribution module includes:
[0140] A load prediction unit, used to implement dynamic load prediction based on a stream processing framework;
[0141] A routing strategy unit, used to implement adaptive routing strategies;
[0142] Traffic control unit, used to implement traffic control based on message middleware;
[0143] Task scheduling unit, used for synchronous task scheduling;
[0144] Resource management unit, used for system resource management.
[0145] The consistency assurance module is used for:
[0146] Capture change information of source data in real time, adopt a multi-level consistency verification method based on hash verification, and regularly verify the data on the source and target ends based on the established data consistency verification points, including data integrity verification, format consistency verification, business rule verification and other dimensions. If data anomalies are found, data repair is automatically performed to ensure the consistency of data synchronization, and through the distributed transaction coordination mechanism, ensure the atomicity and consistency of data operations in a distributed environment.
[0147] In a distributed environment, data consistency is a core issue. The system uses a multi-level consistency check based on hash verification to quickly detect data anomalies. Use hash functions to encode data uniquely. Calculate and compare hash values at different verification points (field level, table level, global level). The steps are as follows:
[0148] 1) Generate a hash value for the transmitted data;
[0149] 2) The hash value is calculated again after the target end receives the data;
[0150] 3) Compare the hash values of the source and target ends;
[0151] 4) If the data is inconsistent, the repair process is triggered to accurately locate and repair the abnormal data by tracing back the historical change records.
[0152] In a specific example, the specific process of the above steps is as follows:
[0153] Step 1: Transfer data to generate hash value
[0154] (1.1) Generate data blocks:
[0155] The data to be transmitted is divided into multiple parts (such as field level, table level, and global level), and a hash value is generated for each part.
[0156] Assume that the data block is (D i ), where (i=1, 2, ..., n).
[0157] (1.2) Calculate the source hash value:
[0158] For data blocks (D i ) Use the hash function (H) to generate the corresponding hash value (h i ):
[0159] h i =H(D i );
[0160] Generate a global level comprehensive hash value (Hglobal) by combining the hash values of all data blocks:
[0161] Hglobal=H(h 1 ||h 2 ||...||h n );
[0162] Among them, (||) represents the string concatenation operation.
[0163] (1.3)Additional hash value:
[0164] The calculated hash value (hi ) and (Hglobal) are attached to the data packet and transmitted to the target end along with the data.
[0165] Step 2: The target receives the data and recalculates the hash value
[0166] (2.1) Extracting data blocks: After receiving data at the target end, extract the transmitted data (D i ′) and the additional hash value (h i ) and (Hglobal).
[0167] (2.2) Calculate the target end hash value:
[0168] For the received data block (D i ′) Recalculate the hash value (h′) using the same hash function (H) i ):
[0169] h′ i =H(D′ i );
[0170] Calculate the global hash value (Hglobal′) of the target:
[0171] Hglobal′=H(h′ 1 ||h′ 2 ||…||h′ n ).
[0172] Step 3: Compare the hash values of the source and target
[0173] (3.1) Level by level comparison:
[0174] Field level validation:
[0175] If h i ≠h′ i , trigger field-level repair
[0176] Table level validation:
[0177] If Htable≠Htable′, trigger table-level repair.
[0178] Global level validation:
[0179] If Hglobal≠Hglobal′, trigger global-level repair.
[0180] (3.2) Verification logic:
[0181] Compare layer by layer from field level to global level, and stop further comparison when the smallest granularity difference is found.
[0182] Step 4: Trigger the repair process
[0183] (4.1) Retrospective change history: Use the change record system to find the changes that occurred between the source and target during the transmission process:
[0184] Find i ) in the transfer path.
[0185] Determines the location of differences in data transfer.
[0186] (4.2) Accurately locate anomalies:
[0187] According to the change record, determine the affected data blocks (D i ) and their locations.
[0188] (4.3) Repair data:
[0189] Resend the correct data from the source to the target, overwriting the abnormal data block. Recalculate the hash value of the repaired data and verify it again to ensure data consistency.
[0190] Related formula:
[0191] Field level validation:
[0192] hfield,i=H(Dfield,i);
[0193] Table level validation:
[0194] htable=H(hfield, 1||hfield, 2||...||hfield, n);
[0195] Global verification:
[0196] hglobal=H(htable, 1||htable, 2||...||htable, m);
[0197] Parameter explanation:
[0198]
[0199] In a specific example, the consistency assurance module includes:
[0200] Change capture unit, used to capture data changes based on CDC technology;
[0201] A data verification unit, used to implement multi-level data verification;
[0202] An anomaly detection unit, used to detect data anomalies;
[0203] A data repair unit, used to automatically repair abnormal data;
[0204] Transaction coordination unit, used for distributed transaction coordination.
[0205] The enterprise multi-source data real-time synchronization system also includes a performance optimization module, which is used for system performance management and optimization, and specifically includes:
[0206] An incremental recognition unit, used to recognize incremental changes in data;
[0207] Parallel processing unit, used for distributing data processing tasks to multiple nodes for parallel execution based on the parallel processing mechanism of the distributed computing framework;
[0208] Data compression unit, used for data transmission compression optimization;
[0209] Resource scheduling unit, which implements resource scheduling based on container technology;
[0210] Performance monitoring unit, used for system performance monitoring.
[0211] In this embodiment, the increment recognition unit
[0212] The Bloom filter records the synchronized data and only identifies the newly added or changed data to reduce the amount of data transmission. The Bloom filter is an efficient probabilistic data structure used to detect whether an element exists in a set. The steps are as follows:
[0213] 1) Initialize the Bloom filter and load the source data hash;
[0214] 2) Calculate the hash value for the new data and check whether the Bloom filter already exists;
[0215] 3) Add the new data to the synchronization task.
[0216] In a specific example, the specific process of the above steps is as follows:
[0217] Step 1: Initialize Bloom filter and load source data hash
[0218] (1.1) Bloom filter initialization:
[0219] A Bloom filter is a bit array of length (m) with all bits initially set to 0.
[0220] Define a set of (k) independent hash functions (H = H 1 , H 2 , ..., H k), each function maps an input to a position in a Bloom filter.
[0221] (1.2) Load source data:
[0222] For each element (x) in the source data set, calculate its (k) hash values:
[0223] H i (x)(1≤i≤k), i is the number of the hash function currently used, ranging from 1, 2, ...k; according to the hash value, the bit at the corresponding position is set to 1:
[0224] B[H i (x)]=1 for all i
[0225] Result: The hash values of all elements in the source data are marked as 1 in the Bloom filter.
[0226] Step 2: Hash the new data and check the Bloom filter
[0227] (2.1) New data detection:
[0228] For each new data element (y), the hash value is calculated using the same (k) hash functions:
[0229] H i (y)(1≤i≤k)(2.2) Bloom filter check:
[0230] Check if the corresponding bits in the Bloom filter are all 1:
[0231]
[0232] in:
[0233] (F(y)): Indicates whether (y) already exists in the Bloom filter.
[0234] If (F(y)=True), then (y) is considered to exist.
[0235] If (F(y)=False), (y) is considered to be newly added or changed data.
[0236] Step 3: Add new data to the synchronization task
[0237] (3.1) Data classification:
[0238] If (F(y)=True), (y) is marked as new data and added to the synchronization task queue.
[0239] If (F(y)=False), skip the data to avoid repeated synchronization.
[0240] (3.2) Bloom filter update:
[0241] For the newly added data (y), set the Bloom filter bit corresponding to its hash value to 1:
[0242] [B[H i (y)]=1 for all i.
[0243] In this embodiment, the data compression unit
[0244] Based on the efficient compression algorithm of LZ4, an intelligent data compression strategy is implemented. In a distributed environment, transmitting a large amount of data will increase bandwidth usage, and the compression algorithm can be used to reduce transmission overhead. The LZ4 algorithm is a fast compression algorithm that matches duplicate data through a sliding window mechanism to reduce storage requirements. The steps are as follows:
[0245] 1) Process the input data stream in blocks to extract repeated patterns;
[0246] 2) Match data blocks through sliding windows and replace them with pointers or symbols;
[0247] 3) Transmit the compressed data and decompress it on the target end.
[0248] In a specific example, the specific process of the above steps is as follows:
[0249] Step 1: Process the input data stream in blocks to extract repeated patterns;
[0250] (1.1) Divide the input data stream S into multiple data blocks B i , each data block size is L(B i ).
[0251] (1.2) Total amount of original data:
[0252]
[0253] (1.3) In each data block, the repeated pattern is extracted through the sliding window mechanism, and the starting position and length of the repeated segment are marked.
[0254] Step 2: Match the data block through the sliding window and replace it with a pointer or symbol;
[0255] (2.1) Use a sliding window to match repeated segments in a data block, replace the matched segments with pointers (p, l), and retain the unmatched data as literals.
[0256] (2.2) Calculate the compressed size of each data block:
[0257] C(B i )=L(Bi )-R(B i );
[0258] in:
[0259] R(B i ): The number of bytes that are reduced after the matching data segment is replaced with a pointer.
[0260] Step 3: Transmit the compressed data and decompress it on the target end
[0261] (3.1) Transmit all compressed data blocks, the total amount of compressed data:
[0262]
[0263] (3.2) After receiving the compressed data, the target end reconstructs the original data based on the pointer (p, l)) and the literal value and verifies:
[0264] L(S)=C(S)+R(S):
[0265] In the formula, L(S) represents the original data volume, C(S) represents the compressed data volume, and R(S) is the compression ratio.
[0266] In this embodiment, the resource scheduling unit is based on a containerized resource scheduling mechanism, so that the synchronization system can dynamically adjust the allocation of computing resources according to load conditions to ensure the optimization of processing performance.
[0267] The enterprise multi-source data real-time synchronization system also includes a security control module, which is used for data access control and security protection, including:
[0268] Identity authentication unit, used for unified identity authentication;
[0269] Access control unit, used to implement role-based access control;
[0270] Data encryption unit, used to implement transport layer security encryption;
[0271] Data desensitization unit, used to process sensitive data;
[0272] Audit log unit, used to record data access logs.
[0273] Compared with the prior art, the present invention has the following advantages:
[0274] 1. Access capability advantages: support unified access to multiple heterogeneous data sources; standardized adapter framework reduces development costs; plug-in design improves system scalability; unified management improves operation and maintenance efficiency; dynamic loading mechanism enhances system flexibility.
[0275] 2. Performance advantages: data synchronization delay is reduced; system throughput is improved; system availability is improved; data consistency is effectively guaranteed; and large-scale concurrent synchronization tasks are supported.
[0276] 3. Reliability assurance advantages: multi-level data verification ensures data accuracy; automated repair improves system reliability; distributed transactions ensure data consistency; complete exception handling mechanism; data operations are traceable.
[0277] The overall architecture of the present invention and the structure and function of each module are further described in detail below.
[0278] 1. Specific implementation of the overall system architecture
[0279] like Figure 1 As shown, the enterprise multi-source data real-time synchronization system of the present invention adopts a layered architecture design, specifically including the following parts:
[0280] 1. Hardware environment configuration:
[0281] Server Configuration:
[0282] CPU: Intel Xeon E5-2680 v4 and above
[0283] Memory: 256GB DDR4
[0284] Storage: NVMe SSD 2TB
[0285] Network card: 10GbE network card.
[0286] 2. Cluster size:
[0287] Management nodes: 3 nodes
[0288] Computing nodes: 10-100 nodes
[0289] Storage nodes: elastically configured according to the amount of data.
[0290] 3. Infrastructure layer:
[0291] Container orchestration platform: Kubernetes cluster
[0292] Distributed file system: HDFS cluster
[0293] Distributed Cache: Redis Cluster
[0294] Message middleware: Kafka cluster
[0295] Load balancing: high availability solution based on HAProxy
[0296] 4. Service component deployment:
[0297] Service Registry: Unified Service Registration and Discovery
[0298] Configuration Center: Unified Configuration Management
[0299] Monitoring Center: System Operation Status Monitoring
[0300] Log Center: Centralized log management
[0301] Alarm center: Real-time alarm for abnormal situations.
[0302] 2. Specific implementation of data access adaptation module
[0303] like Figure 2 As shown, the data access adapter module realizes the unified access of multi-source heterogeneous data, including the adapter framework 201, the data source connection pool 202, the data reading engine 203, the first state manager 204, and the configuration manager 205; wherein, the adapter framework 201 includes a standard interface and a plug-in management module; the data source connection pool 202 is used for connection management and resource reuse; the data reading engine 203 is used for data extraction and data reading; the first state manager 204 is used to display the synchronization status and monitoring records; the configuration manager 205 is used for configuration information management and parameter management. The configuration manager 205 performs configuration management on the adapter framework 201, the data source connection pool 202, and the data reading engine 203; the first state manager 204 performs status monitoring on the data source connection pool 202 and the data reading engine 203. Specifically including:
[0304] 1. Adapter framework implementation:
[0305] Standard interface definition:
[0306] Data reading interface
[0307] State Management API
[0308] Configuration Management Interface
[0309] Lifecycle management interface.
[0310] Plug-in implementation:
[0311] Dynamic class loading mechanism
[0312] Plugin registration mechanism
[0313] Version management mechanism hot loading support.
[0314] 2. Data source access implementation:
[0315] Relational database access:
[0316] JDBC-based universal adapter Incremental collection based on log analysis
[0317] Connection pool optimization
[0318] Concurrency control.
[0319] Distributed database access:
[0320] Dedicated API connection
[0321] Sharding data processing
[0322] State synchronization mechanism
[0323] Abnormal recovery.
[0324] 3. Adapter management implementation:
[0325] Adapter life cycle:
[0326] Registration and initialization
[0327] Start and run
[0328] Stop and uninstall.
[0329] Condition Monitoring:
[0330] Running status detection
[0331] Performance indicator collection
[0332] Automatic failover with abnormal situation alarm.
[0333] 3. Specific implementation of data routing distribution module
[0334] like Figure 3 As shown, the data routing and distribution module implements an intelligent data distribution mechanism, and realizes data distribution through load monitoring 301, policy calculation 302, data distribution 303, flow control 304, and state feedback 305. During the distribution process, the state of load monitoring 301 is updated through state feedback 305, and the policy calculation 302 is optimized and fed back; the flow of load monitoring 301 is monitored through flow control 304. Specifically, load monitoring 301 includes system load and resource usage monitoring; policy calculation 302 includes routing computer optimization selection; data distribution 303 includes data routing and data forwarding; flow control 304 includes rate limiting and flow balancing; state feedback 305 includes result collection and status update. Specifically including:
[0335] 1. Load prediction implementation:
[0336] Index collection:
[0337] CPU usage
[0338] Memory usage
[0339] Network bandwidth utilization
[0340] Disk I / O status
[0341] Prediction Algorithm:
[0342] Sliding Window Analysis
[0343] Trend Forecast
[0344] Threshold Dynamic Adjustment
[0345] Early warning mechanism
[0346] 2. Routing strategy implementation:
[0347] Strategy calculation:
[0348] Node load weight
[0349] Network delay factor
[0350] Data Affinity
[0351] Resource utilization.
[0352] Dynamic Adjustment:
[0353] Real-time load balancing
[0354] Automatic failover
[0355] Performance optimization
[0356] Configure dynamic updates.
[0357] 4. Specific Implementation of the Consistency Assurance Module
[0358] like Figure 4 As shown, the consistency assurance module implements data consistency assurance in a distributed environment, including a CDC engine 401, a checker 402, a repair processor 403, a transaction coordinator 404, and a second state manager 405. Among them, the CDC engine 401 is used for change capture and data analysis; the checker 402 is used for data verification and consistency check; the repair processor 403 is used for difference analysis and data repair; the transaction coordinator 404 is used for transaction management and commit control; the second state manager 405 is used for state tracking and state maintenance; the second state manager 405 monitors the state of the CDC engine 401, the checker 402, the repair processor 403, and the transaction coordinator 404; the transaction coordinator 404 performs transaction control on the checker 402 and the repair processor 403. Specifically including:
[0359] 1. Change capture implementation:
[0360] CDC engine configuration:
[0361] Data source connection configuration
[0362] Table-level filtering rules
[0363] Field-level filtering rules
[0364] Change event definition.
[0365] Data collection and processing:
[0366] Log parsing and processing
[0367] Transaction integrity assurance
[0368] Breakpoint resume mechanism
[0369] Concurrency control management.
[0370] 2. Data verification implementation:
[0371] Checkpoint settings:
[0372] Time point verification
[0373] Transaction Checkpointing
[0374] Data volume verification
[0375] Business rule verification.
[0376] Verification process:
[0377] Full data verification
[0378] Incremental data verification
[0379] Real-time consistency check
[0380] Regular consistency checks.
[0381] 3. Data repair implementation:
[0382] Repair strategy:
[0383] Automatic repair rules
[0384] Human intervention mechanism
[0385] Priority Management
[0386] Impact surface analysis.
[0387] Repair process:
[0388] Anomaly Detection
[0389] Differences
[0390] Data backfill
[0391] Result verification.
[0392] 5. Specific implementation of performance optimization module
[0393] like Figure 5 As shown, the performance optimization module realizes the comprehensive optimization of system performance, including an incremental processor 501, a parallel scheduler 502, a resource manager 503, a performance monitor 504, and an optimization decision maker 505. Among them, the incremental processor 501 is used for incremental identification and incremental synchronization; the parallel scheduler 502 is used for task allocation and concurrency control; the resource manager 503 is used for resource allocation and load balancing; the performance monitor 504 is used for performance collection and indicator analysis; the optimization decision maker 505 is used for strategy generation and optimization execution. The optimization decision maker 505 provides processing strategies, scheduling strategies and optimization strategies to the incremental processor 501, the parallel scheduler 502 and the resource manager 503 respectively; the incremental processor 501 and the parallel scheduler 502 provide performance data to the performance monitor 504; the resource manager 503 provides resource status information to the performance monitor 504. Specifically including:
[0394] 1. Incremental processing implementation:
[0395] Incremental recognition:
[0396] Timestamp
[0397] Version number tag
[0398] Change flag incremental log analysis. Data processing:
[0399] Incremental data extraction
[0400] Data merging process
[0401] Conflict resolution state maintenance.
[0402] 2. Parallel processing implementation: Task decomposition:
[0403] Data Sharding Strategy
[0404] Task prioritization Dependency analysis Resource estimation.
[0405] Parallel Scheduling:
[0406] Work thread pool management tasks dynamically allocate progress monitoring exception handling.
[0407] 3. Resource scheduling implementation: container management:
[0408] Container resource configuration
[0409] Elastic Scaling Strategy
[0410] Resource Limitation Policy
[0411] Scheduling optimization.
[0412] Performance Monitoring:
[0413] Real-time performance data collection Performance bottleneck analysis
[0414] Trend Forecast
[0415] The alarm is triggered.
[0416] 6. Specific implementation of the security control module
[0417] like Figure 6 As shown, the security control module implements all-round security protection, including the authentication center 601, the authority manager 602, the encryption processor 603, the desensitizing processor 604, and the audit log manager 605. Among them, the authentication center 601 is used for identity authentication and credential management; the authority manager 602 is used for access control and policy execution; the encryption processor 603 is used for encryption operations and key management; the desensitizing processor 604 is used for data desensitization and rule application; the audit log manager 605 is used for log recording and log analysis. The audit log manager 605 audits the authentication center 601, the authority manager 602, the encryption processor 603, and the desensitizing processor 604; the authority manager 602 controls the encryption processor 603 and the desensitizing processor 604. Specifically including:
[0418] 1. Identity authentication implementation:
[0419] Authentication mechanism:
[0420] SSO Integration Configuration
[0421] Multi-factor authentication
[0422] Token Management
[0423] Session control.
[0424] Permission management:
[0425] Role Definition
[0426] Permission Assignment
[0427] Permission inheritance
[0428] Dynamic authorization.
[0429] 2. Data security implementation:
[0430] Transmission encryption:
[0431] TLS configuration management
[0432] Certificate Management
[0433] Key Update
[0434] Encryption policy.
[0435] Data desensitization:
[0436] Desensitization rule configuration
[0437] Field-level desensitization
[0438] Dynamic desensitization
[0439] Verification of desensitization effect.
[0440] 3. Audit log implementation:
[0441] Logging:
[0442] Operation log collection
[0443] Access Logging
[0444] System log management
[0445] Audit log storage.
[0446] Log analysis:
[0447] Real-time log analysis
[0448] Abnormal behavior detection
[0449] Audit report generation
[0450] Compliance check.
[0451] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. It should be pointed out that a person skilled in the art can make several improvements and modifications without departing from the technical principles of the present invention, and these improvements and modifications should also be regarded as within the scope of protection of the present invention.
Claims
1. A real-time synchronization system for enterprise multi-source data based on data middle platform, characterized in that: Adopt containerized deployment mode and layered architecture, including data access layer, data processing layer, data distribution layer and data control layer; the data access layer communicates with the data middle platform and obtains the original data from the data middle platform; The data access layer is provided with a data access adapter module, the data processing layer is provided with a data format processing module, the data distribution layer is provided with a data routing distribution module, and the data control layer is provided with a consistency guarantee module; each layer communicates through a standardized interface to ensure the scalability and maintainability of the system; The data access adapter module is used for unified access of multi-source heterogeneous data, including a relational database adapter unit, a distributed database adapter unit, a data warehouse adapter unit and a search engine adapter unit, which are respectively used to access a relational database, a distributed database, a data warehouse and a search engine; The data format processing module is used to perform data standardization processing on the accessed multi-source heterogeneous data, and send the processed standardized data to the data routing distribution module; The data routing distribution module is used for: During the data routing process, the load status of each node is monitored in real time, and the ARIMA model is used to perform dynamic load forecasting in combination with historical data; Based on the load prediction results, an adaptive routing strategy is adopted to dynamically distribute the data flow to the optimal node, and a flow control mechanism based on message middleware is adopted to prevent system overload and ensure the stability of data processing; wherein, the adaptive routing strategy considers node load, network status, and data affinity factors, and performs optimal route selection based on the ant colony algorithm to ensure the balance and efficiency of data distribution; The consistency assurance module is used for: Capture change information of source data in real time, adopt a multi-level consistency verification method based on hash verification, and regularly verify the data on the source and target ends based on the established data consistency verification points. If data anomalies are found, data repair is automatically performed to ensure the consistency of data synchronization, and through the distributed transaction coordination mechanism, ensure the atomicity and consistency of data operations in a distributed environment.
2. According to claim 1, the enterprise multi-source data real-time synchronization system based on the data middle platform is characterized in that: The data access adapter module adopts a plug-in structure, supports dynamic loading and management, and personalized expansion of data sources; The data access adapter module includes an adapter management unit, which is used to uniformly manage the life cycles of various adapters and realize dynamic deployment and upgrade of adapters.
3. According to claim 1, the enterprise multi-source data real-time synchronization system based on the data middle platform is characterized in that: The data format processing module includes: A data format identification unit, used for identifying a data format type; Data cleaning unit, used to clean abnormal data; A data conversion unit, used for data format conversion; Data standardization unit, used to unify data formats; Quality control unit for data quality checking.
4. According to claim 1, the enterprise multi-source data real-time synchronization system based on the data middle platform is characterized in that: The steps of performing optimal route selection based on ant colony algorithm are as follows: Step 1: Initialization: The initial pheromone concentration is set to a constant value (τ ij (0) = τ0); set to initial value (τ0); Initialize system status: collect the current load (L j ) and path bandwidth (Bij); Step 2: Calculate path weight: Comprehensive weight formula: in: (L j ): the load of node (j), the lower the better; (Bij): the remaining bandwidth of the path (Pij), the higher the better; (α, β): weight factors that control the impact of load and bandwidth on the decision; Step 3: Path selection: Selection probability formula: in: τij(t): pheromone concentration of path (Pij); (Wij): comprehensive weight of path (Pij); (η, γ): parameters that control the importance of pheromone and path weight, respectively; Step 4: Pheromone Update: Update formula: τij(t+1)=(1-ρ)τij(t)+Δτij; in: (1-ρ): pheromone volatilization factor, which prevents pheromone from growing indefinitely; (Δτij): Newly added pheromone, indicating the path quality.
5. According to claim 1, the enterprise multi-source data real-time synchronization system based on the data middle platform is characterized in that: The multi-level consistency verification method based on hash verification is as follows: 1) Generate a hash value for the transmitted data; 2) The hash value is calculated again after the target end receives the data; 3) Compare the hash values of the source and target ends; 4) If the data is inconsistent, the repair process is triggered to accurately locate and repair the abnormal data by tracing back the historical change records.
6. The enterprise multi-source data real-time synchronization system based on data middle platform according to claim 1 is characterized in that: It also includes a performance optimization module, which includes: An incremental recognition unit, used to recognize incremental changes in data; Parallel processing unit, used for distributing data processing tasks to multiple nodes for parallel execution based on the parallel processing mechanism of the distributed computing framework; Data compression unit, used for data transmission compression optimization; Resource scheduling unit, which implements resource scheduling based on container technology; Performance monitoring unit, used for system performance monitoring.
7. The enterprise multi-source data real-time synchronization system based on data middle platform according to claim 6 is characterized in that: The incremental identification unit is specifically used for: The Bloom filter records the synchronized data and only identifies the newly added or changed data to reduce the amount of data transmission. The specific process is as follows: 1) Initialize the Bloom filter and load the source data hash; 2) Calculate the hash value for the new data and check whether the Bloom filter already exists; 3) Add the new data to the synchronization task.
8. The enterprise multi-source data real-time synchronization system based on data middle platform according to claim 7 is characterized in that: The data compression unit is specifically used for: Based on the efficient compression algorithm of LZ4, an intelligent data compression strategy is implemented. The specific process is as follows: 1) Process the input data stream in blocks to extract repeated patterns; 2) Match data blocks through sliding windows and replace them with pointers or symbols; 3) Transmit the compressed data and decompress it on the target end.
9. The enterprise multi-source data real-time synchronization system based on data middle platform according to claim 8 is characterized in that: The resource scheduling unit is specifically used for: Based on the containerized resource scheduling mechanism, the synchronization system can dynamically adjust the allocation of computing resources according to the load conditions to ensure the optimization of processing performance.
10. The enterprise multi-source data real-time synchronization system based on data middle platform according to claim 1 is characterized in that: It also includes a safety control module, which includes: Identity authentication unit, used for unified identity authentication; Access control unit, used to implement role-based access control; Data encryption unit, used to implement transport layer security encryption; Data desensitization unit, used to process sensitive data; Audit log unit, used to record data access logs.
Citation Information
Patent Citations
Cross-platform equipment information acquisition system
CN117640748A
Order synchronization method of business management platform
CN117725122A
Data interaction system for archive management
CN117827743A
Multi-source computing power data integration and intelligent scheduling system and method
CN118916147A
Distributed Database System Providing Data and Space Management Methodology
US20060101081A1
Cited By
Data consistency verification method and system based on cursor propulsion and electronic equipment
CN120892436A
A data consistency checking method and system based on cursor pushing and an electronic device
CN120892436B