Data management method for real-time incremental data based on dynamic scheduling and elastic rollback
By adopting a real-time incremental data management method of dynamic scheduling and elastic rollback in the big data environment, the problems of low efficiency of full updates and consistency challenges are solved, and efficient and reliable data updates and consistency guarantees are achieved.
Patent Information
- Application Number
- CN202510279894.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-06-27
AI Technical Summary
In the era of big data, the traditional full-scale update method has problems such as low efficiency, high cost and inability to meet real-time requirements in scenarios where data scale is large and real-time requirements are high. Incremental update technology faces the challenges of data consistency and error recovery.
The real-time incremental data management method based on dynamic scheduling and elastic rollback is adopted. The data change capture unit is monitored and captured in real time. The incremental data processing unit uses AI model for dynamic optimization and scheduling. The data sharding technology and consistency detection module ensure data consistency. The elastic rollback mechanism automatically rolls back when the update fails.
It significantly improves the update efficiency and data consistency guarantee capabilities in large-scale data environments, improves the system's self-healing ability and business continuity, and reduces manual intervention and system resource utilization.
Smart Images

Figure CN120216516A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of information data management, and particularly relates to a data management method for real-time incremental data based on dynamic scheduling and elastic rollback. Background Art
[0002] In scenarios with relatively small data volumes, full update (i.e., reloading the entire data set every time an update is made) used to be the mainstream method. However, with the advent of the big data era, the data scale has grown exponentially, and transmitting and processing the complete data set will consume a large amount of network bandwidth and computing resources. Especially in distributed systems, the cost of synchronizing full data across nodes is extremely high, and the full update cycle is long, which cannot meet the requirements of scenarios with high real-time requirements such as financial transactions and Internet of Things monitoring. Large-scale data refreshing may cause the system to be unavailable for a long time, affecting business continuity.
[0003] To address the deficiencies of full update, incremental update technology has gradually become the core solution. Its core idea is to only process data changes (additions, modifications, deletions) since the last update, thereby significantly improving efficiency. Especially in fields such as financial transactions, e-commerce inventory management, and logistics tracking, where data synchronization in seconds is required to ensure the accuracy of decision-making and operations. Incremental update technology reduces the dependence on database locks and improves concurrent processing capabilities through methods such as change data capture (CDC), timestamps, or version control.
[0004] Although incremental update is efficient, it faces challenges in data consistency and error recovery. For example, node failures or network interruptions in a distributed system may cause some updates to fail, leading to data inconsistency. Therefore, by adopting an elastic rollback mechanism to enhance system reliability, in case of update failure, the rollback operation can restore the data to a consistent state, avoiding the risk of "half-successful" updates; combined with log records (recording operation steps), redundant backups (snapshots before updates), and distributed locks (preventing concurrent conflicts) to ensure data integrity in complex environments; through predefined rollback strategies (such as version rollback or reverse log operations), manual intervention is reduced and the system's self-healing ability is improved.
[0005] When implementing elastic rollback real-time incremental updates in a distributed architecture, since a distributed system needs to balance consistency, availability, and partition tolerance, real-time incremental updates require the optimization of the eventual consistency model and the implementation of algorithms for efficient fault recovery and data consistency verification. For example, two-phase commit (2PC) or distributed lock coordination of multi-node operations, stream processing frameworks (such as Apache Flink) and message queues (such as Kafka) perform low-latency processing on the capture, transmission, and application of incremental data in high-concurrency scenarios, avoiding performance bottlenecks introduced by the rollback mechanism;
[0006] However, for complex distributed systems, simply relying on elastic rollback and incremental update technology may still face many challenges. In extreme cases, such as large-scale data migration or system reconstruction, the accumulation of incremental data may make the rollback operation extremely complex and time-consuming. In addition, when there are a large number of concurrent operations in the system, how to ensure that the capture and rollback of incremental data will not introduce new data conflicts and consistency issues is also a technical problem that needs to be solved urgently. Summary of the invention
[0007] In order to solve the above problems existing in the prior art, the present invention provides a data management method for real-time incremental data based on dynamic scheduling and elastic rollback;
[0008] The purpose of the present invention can be achieved by the following technical solutions:
[0009] A data management method for real-time incremental data based on dynamic scheduling and elastic rollback, comprising:
[0010] S1: Set up a data change capture unit to monitor and capture data changes in the system in real time, identify data insertion, update, and deletion operations, and record the change information to provide a basis for subsequent incremental updates and elastic rollbacks;
[0011] The specific components of the data change capture unit include: a change listener, a change extractor, a change storage, and a change distributor;
[0012] The change listener monitors data source changes through the database log of the data source monitoring layer and generates data change events;
[0013] The change extractor is used to extract key features of the data change event, where the key features include the change type and the changed data content;
[0014] The change storage is used to extract the key features and store them in a change message queue;
[0015] The change distributor is used to distribute the change information stored in the change storage to corresponding processing modules, such as an incremental update module or a rollback module, according to a set rule;
[0016] S2: receiving the original incremental data sent by the data change capture unit through the incremental data processing unit, wherein the specific system components of the incremental data processing unit include a parsing and conversion module, a dynamic optimization module, a data sharding module, and a consistency detection module;
[0017] The parsing and conversion module is used to identify the data source format of the original incremental data, convert the identified data source format into an interactive data packet format, add a priority identifier, and send the interactive data packet to the dynamic optimization module;
[0018] The dynamic optimization module constructs a priority scheduling model based on the embedded lightweight AI model. The priority scheduling model generates a priority score corresponding to the interactive data packet by analyzing the multi-dimensional features of the interactive data packet in real time. Specifically, as the AI-driven decision-making core, the priority scheduling model adopts a neural network architecture or a reinforcement learning algorithm to jointly model and assign weights to the input features (including business criticality classification, data change hierarchy, transmission time limit margin, protocol data unit size, link load index), and finally outputs a priority score. The feature extraction layer and the scheduling decision layer are co-trained in an end-to-end manner to ensure that the model adapts to the dynamic network environment.
[0019] The business criticality classification classifies data according to the importance of the data to business operations, including core data that directly affects business continuity, data that affects key business processes but is fault-tolerant, and auxiliary business data.
[0020] The data change hierarchy is divided into an architecture level, an application level, and a user level according to the source of incremental data changes under the information governance framework; the architecture level represents cross-business unit integrated change control (such as database Schema changes), the application level represents data asset modifications within a single business domain (such as API interface parameter adjustments), and the user level represents data updates triggered by end users (such as customer information editing in the CRM system);
[0021] The transmission time limit margin represents the estimated time required for data to reach the target system, which is dynamically calculated based on historical data transmission speeds and the current network condition to ensure that critical data can be processed as soon as possible and reduce the impact of data latency on business continuity; it can be calculated through the average historical transmission rate and the current network quality factor. The specific calculation formula is:
[0022]
[0023] Among them, TTL is the transmission time limit margin, D current is the real-time delay of the current data transmission, n represents the number of levels of data changes, which is used to trace the transmission path of data from the source to the target system; D i represents the estimated delay of data transmission between each level, T i represents the actual time of data transmission between each level, α is the weight coefficient of the average historical transmission rate for calculating the transmission time limit margin, β is the weight coefficient of the current network quality factor for calculating the transmission time limit margin, B avail represents the available amount of the current network bandwidth, P loss is the packet loss rate, which reflects the stability of the current network, T elapsed represents the time that has elapsed since the data departed from the source, which is used to dynamically adjust the transmission strategy;
[0024] The protocol data unit size is the packet size; the link load index is quantitatively graded through bandwidth utilization, queue depth ratio, and packet loss rate weight;
[0025] A dynamic priority queue is established based on the priority score output by the priority scheduling model. Based on the time synchronization mechanism and traffic shaping of the TSN protocol, a fixed time window is allocated to the high-priority queue. When high-priority data arrives, the transmission of low-priority data is immediately interrupted, and microsecond-level preemption is achieved through the frame preemption function of TSN;
[0026] The data sharding module, based on the horizontal partitioning technology, divides a large-scale data set into multiple smaller subsets (shards), each shard is independently stored on different nodes or servers, and efficient data management is achieved through a distributed architecture; the shards are divided according to the range of data keys (such as timestamps, numerical intervals). For example, the incremental data within a certain time period is allocated to the same shard. After the original incremental data is sharded, each shard can independently perform operations such as parsing, compression, and verification, and utilize multi-node parallelism to accelerate the processing flow;
[0027] The consistency detection module depends on the replica distribution strategy of the data sharding module, ensures data integrity through cross-shard verification, and connects to the elastic rollback unit. Methods such as version control mechanism, timestamp marking, and data integrity constraints are used to ensure data consistency. When data inconsistency is detected, the corresponding recovery mechanism will be triggered;
[0028] S3: The real-time update execution unit receives the processed incremental data and updates the target data set according to the data change event; during the update process, the module uses an optimistic lock or pessimistic lock mechanism to control concurrent access to prevent data conflicts; in addition, the module records the operation log of each update for rollback operations when necessary.
[0029] S4: The elastic rollback mechanism unit is automatically triggered when an update failure or data inconsistency is detected. The module restores the data to a consistent state based on the operation log and redundant backup data. During the rollback process, the module uses predefined rollback strategies, such as version rollback or log reverse operations, to reduce manual intervention and improve the system's self-healing ability.
[0030] S5: The system monitoring and alarm unit monitors the running status of the data management method and system in real time, including data update success rate, rollback operation frequency, system resource occupancy, etc. When an anomaly or potential risk is detected, the module automatically triggers the alarm mechanism to notify the management personnel for intervention and handling.
[0031] The beneficial effects of the present invention are:
[0032] By implementing the data management method and system provided by the present invention, the update efficiency and data consistency guarantee ability in a large-scale data environment are significantly improved. Specifically, the efficient listening and parsing mechanism of the data change capture unit ensures the real-time capture and accurate recording of data changes, providing a solid foundation for subsequent incremental updates and rollback operations. The introduction of the incremental data processing unit, especially the priority scheduling strategy of the dynamic optimization module based on the AI model, effectively balances the real-time nature of data updates and the utilization rate of network resources, significantly improving the processing speed of key data and the overall response ability of the system.
[0033] The application of data sharding technology not only reduces the processing pressure on a single node, but also accelerates the data update process through parallel processing, while maintaining data consistency and integrity. The close cooperation between the consistency detection module and the elastic rollback mechanism further enhances the fault tolerance and self-healing ability of the system. Even in the face of a complex distributed environment and sudden failures, it can quickly restore data consistency and ensure the continuity and accuracy of business operations.
[0034] In addition, the real-time monitoring and early warning functions of the system monitoring and alarm unit provide timely and accurate system operation status information for management personnel, helping to promptly discover and handle potential problems, and avoiding the risks of business interruption and data loss caused by system anomalies. In summary, the present invention provides an efficient, reliable, and intelligent data management solution, which is applicable to various application scenarios with extremely high requirements for real-time nature and data consistency. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] For the convenience of those skilled in the art to understand, the present invention will be further described below with reference to the accompanying drawings.
[0036] Figure 1 It is a schematic flowchart of the data management method for real-time incremental data based on dynamic scheduling and elastic rollback of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0037] To further elaborate on the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the following will describe in detail the specific implementation manners, structures, features, and their effects of the present invention with reference to the accompanying drawings and preferred embodiments.
[0038] Please refer to Figure 1 , a data management method for real-time incremental data based on dynamic scheduling and elastic rollback, including:
[0039] S1: Set up a data change capture unit to monitor and capture data changes in the system in real time, used to identify data insertion, update, and deletion operations, and record the change information, providing a basis for subsequent incremental updates and elastic rollbacks;
[0040] The specific components of the data change capture unit include: a change listener, a change extractor, a change memory, and a change distributor;
[0041] The change listener monitors the data source changes through the database log of the data source monitoring layer and generates data change events;
[0042] The change extractor is used to extract the key features of the data change events, and the key features include the change type and the changed data content;
[0043] The change memory is used to extract the key features and store them into the change message queue;
[0044] The change distributor is used to distribute the change information stored in the change memory to the corresponding processing modules according to the set rules, such as the incremental update module or the rollback module;
[0045] In this embodiment, an efficient listening mechanism is adopted, such as technologies based on database triggers, log parsing, or file system inotify, etc., to ensure that relevant information can be captured at the first time when data changes occur. Parse the captured original change data to obtain the specific content and type of the change. For example, parse information such as table names, field names, old values, and new values from the database log. Filter out unnecessary change data according to preset rules, such as changes to system tables and changes to temporary data, etc., to reduce the amount of data for subsequent processing and improve processing efficiency. Store the parsed change data in a structured manner. Common data structures include JSON, XML, or binary formats, etc. The storage medium can be memory, local files, distributed databases, etc., to meet different performance and durability requirements. Push the captured change data to the incremental update module to update the data in other systems or databases in real time and maintain data consistency. When needed, roll back to the specified historical state according to the stored change data to quickly restore the accuracy of the system data.
[0046] S2: Receive the original incremental data sent by the data change capture unit through the incremental data processing unit. The specific system components of the incremental data processing unit include a parsing and conversion module, a dynamic optimization module, a data sharding module, and a consistency detection module;
[0047] The parsing and conversion module is used to identify the data source format of the original incremental data, uniformly convert the identified data source format into an interactive data packet format, add a priority identifier, and send the interactive data packet to the dynamic optimization module;
[0048] The dynamic optimization module constructs a priority scheduling model based on an embedded lightweight AI model. The priority scheduling model generates a priority score corresponding to an interactive data packet by analyzing the multi-dimensional features of the interactive data packet in real time. Specifically, as the AI-driven decision-making core, the priority scheduling model adopts a neural network architecture or a reinforcement learning algorithm to jointly model and assign weights to input features (including business criticality classification, data change hierarchy, transmission time limit margin, protocol data unit size, link load index), and finally outputs a priority score. The feature extraction layer and the scheduling decision layer are co-trained in an end-to-end manner to ensure that the model adapts to the dynamic network environment.
[0049] The business criticality classification classifies data according to its importance to business operations, including core data that directly affects business continuity, data that affects key business processes but is fault-tolerant, and auxiliary business data.
[0050] The data change hierarchy is divided into an architecture level, an application level, and a user level according to the source of incremental data changes under the information governance framework; the architecture level represents cross-business unit integrated change control (such as database Schema changes), the application level represents data asset modifications within a single business domain (such as API interface parameter adjustments), and the user level represents data updates triggered by end-users (such as customer information editing in a CRM system);
[0051] The transmission time limit margin represents the estimated time required for data to reach the target system, which is dynamically calculated based on historical data transmission speeds and the current network condition to ensure that critical data can be processed as soon as possible and reduce the impact of data latency on business continuity; it can be calculated through the mean historical transmission rate and the current network quality factor. The specific calculation formula is:
[0052]
[0053] where TTL is the transmission time limit margin, D current is the real-time delay of the current data transmission, n represents the number of levels of data changes, which is used to trace the transmission path of data from the source to the target system; D i represents the estimated delay of data transmission between each level, T i represents the actual time of data transmission between each level, α is the weight coefficient of the mean historical transmission rate for calculating the transmission time limit margin, β is the weight coefficient of the current network quality factor for calculating the transmission time limit margin, B avail represents the available amount of the current network bandwidth, P loss is the packet loss rate, which reflects the stability of the current network, T elapsed is the time elapsed since the data started from the source, which is used to dynamically adjust the transmission strategy;
[0054] The size of the protocol data unit is the size of the data packet; the link load index is quantitatively graded through bandwidth utilization rate, queue depth ratio, and packet loss rate weight;
[0055] In this embodiment, the objective function of the priority scoring is a linear combination of input features. The training method of the priority scheduling model adopts an end-to-end collaborative training architecture. Through the backpropagation algorithm, the parameters of the feature extraction layer (such as the normalization module) and the scheduling decision layer (weight allocation module) are updated synchronously to ensure the robustness of the model to feature changes. Incremental training is carried out using real-time interactive data packet flows. For example, the weights are dynamically updated using Stochastic Gradient Descent (SGD) to avoid model lag caused by sudden changes in network status; the regularization strength is dynamically adjusted according to the link load index, and the constraint is strengthened during high load to prevent overfitting; if the Transmission Time Limit (TTL) margin is tight, the learning rate is increased to accelerate model convergence. The formula is:
[0056]
[0057] where η is the baseline learning rate and ∈ is the smoothing term.
[0058] A dynamic priority queue is established based on the priority score output by the priority scheduling model. Based on the time synchronization mechanism and traffic shaping of the TSN protocol, a fixed time window is allocated to the high-priority queue. When high-priority data arrives, the transmission of low-priority data is immediately interrupted, and microsecond-level preemption is achieved through the frame preemption function of TSN;
[0059] The data sharding module, based on the horizontal partitioning technology, divides a large-scale data set into multiple smaller subsets (shards). Each shard is independently stored on different nodes or servers, and efficient data management is achieved through a distributed architecture; the shards are divided according to the range of data keys (such as timestamps, numerical intervals). For example, the incremental data within a certain time period is assigned to the same shard. After the original incremental data is sharded, each shard can independently perform operations such as parsing, compression, and verification, and multi-node parallelism is used to accelerate the processing flow;
[0060] The consistency detection module depends on the replica distribution strategy of the data sharding module, ensures data integrity through cross-shard verification, and connects to the elastic rollback unit. Methods such as version control mechanism, timestamp marking, and data integrity constraints are used to ensure data consistency. When data inconsistency is found, the corresponding recovery mechanism will be triggered;
[0061] S3: The real-time update execution unit receives the processed incremental data and updates the target data set according to the data change event; during the update process, the module uses an optimistic lock or pessimistic lock mechanism to control concurrent access to prevent data conflicts; in addition, the module records the operation log of each update for rollback operations when necessary.
[0062] In this embodiment, the real-time update execution unit also has a data conflict detection and resolution mechanism. When concurrent update conflicts are detected, it can automatically perform conflict arbitration to ensure the ultimate consistency of data. During the data update process, the real-time update execution unit will update the data in an orderly manner based on the business logic and dependencies of the data, avoiding inconsistency problems during the data update process. At the same time, the real-time update execution unit also has the function of resume breakpoint transfer. When the data update is interrupted due to network failure or other reasons, it can continue to transfer data from the breakpoint, ensuring the integrity and continuity of the data update. After the data update is completed, the real-time update execution unit will perform data verification and validation to ensure that the updated data is consistent with the data in the target dataset, improving the accuracy and reliability of the data. In addition, the real-time update execution unit also has the functions of data backup and recovery, and can quickly restore data when the data update fails or data is lost, ensuring the continuity of the business and the integrity of the data.
[0063] S4: The elastic rollback mechanism unit is automatically triggered when an update failure or data inconsistency is detected. The module restores the data to a consistent state based on the operation log and redundant backup data. During the rollback process, the module adopts predefined rollback strategies, such as version rollback or log reverse operation, to reduce manual intervention and improve the self-healing ability of the system.
[0064] In this embodiment, the liveness probe and readiness probe of Kubernetes are used to monitor the Pod status in real time. If the new version of the Pod fails to start or fails consecutive health checks (such as HTTP response timeout, TCP port unreachable), the system will mark it as "update failed"; the distributed log (such as a hash table based on in-memory log) is used to compare the data version differences between the primary node and the backup node. When it is detected that the transaction log sequence number (such as Log Sequence Number, LSN) is discontinuous, the data repair process is triggered; a rolling update triggers the execution of kubectl set image deployment / nginx nginx=nginx:1.25 to start the update, and the Deployment creates a new ReplicaSet and gradually replaces the old Pod. Exception detection: If the new Pod fails to pass the readinessProbe within 300 seconds, the CrashLoopBackOff state is triggered. The monitoring system (such as Prometheus) detects that the error rate exceeds the threshold (such as 5%). Automated rollback: Query the.spec.rollbackTo field to obtain the target version. Call the Kubernetes API to resize the old ReplicaSet retained by the revisionHistoryLimit of the Deployment. Synchronously clean up the abnormal version resources and generate audit logs through Events. Progressive rollback: Control the rollback speed through maxSurge=25% and maxUnavailable=10% to avoid service interruption. Cross-service coordination: In a microservices architecture, synchronously roll back associated services through a service mesh (such as Istio) to prevent version incompatibility.
[0065] S5: The system monitoring and alarm unit monitors the running status of the data management method and system in real time, including the data update success rate, rollback operation frequency, system resource occupancy, etc. When an anomaly or potential risk is detected, the module automatically triggers the alarm mechanism to notify the management personnel for intervention and handling.
[0066] In this embodiment, the system monitoring and alarm unit collects and analyzes key performance indicators (KPIs) in real time through integrated monitoring tools such as Prometheus, Grafana, etc. These KPIs include but are not limited to data update success rate, rollback operation frequency, CPU usage rate, memory occupancy rate, disk I / O performance, network bandwidth utilization, etc. By setting reasonable thresholds and alarm rules, the system can automatically trigger the alarm mechanism when detecting anomalies or potential risks. The alarm methods are diverse, including but not limited to email notifications, SMS alerts, phone reminders, and Webhook notifications, ensuring that management personnel can receive alarm information in a timely manner and respond and handle it quickly. In addition, the system monitoring and alarm unit also provides historical data analysis and visualization report functions to help management personnel deeply understand the system operation status, optimize resource allocation, and improve the stability and reliability of the system.
[0067] The computer storage medium of the embodiment of the present invention can adopt any combination of one or more computer-readable media. The computer-readable media can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the computer-readable storage medium include: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0068] The computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries the computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0069] The program code contained on a computer-readable medium can be transmitted using any suitable medium, including but not limited to wireless, wire, optical fiber cable, RF, etc., or any suitable combination of the foregoing. The computer program code for performing the operations of the present invention can be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (e.g., by using an Internet service provider to connect through the Internet).
[0070] As described above, the above are only the preferred embodiments of the present invention, and there is no limitation to the present invention in any form. Although the present invention has been disclosed above with the preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or variations equivalent to the equivalent embodiments within the scope of the technical solution of the present invention without departing from the technical solution of the present invention. However, any simple modification, equivalent change, and modification made to the above embodiments based on the technical essence of the present invention still fall within the scope of the technical solution of the present invention.
Claims
1. A data management method for real-time incremental data based on dynamic scheduling and elastic rollback, characterized in that: The method includes: Set up a data change capture unit to monitor and capture data changes in the system in real time, identify data insertion, update, and deletion operations, and record change information to provide a basis for subsequent incremental updates and elastic rollbacks; Receiving the original incremental data sent by the data change capture unit through the incremental data processing unit, and processing the original incremental data; Receive the processed incremental data through the real-time update execution unit and update the target data set according to the data change event; When an update failure or data inconsistency is detected, the elastic rollback mechanism unit automatically triggers a rollback operation to restore the data to a consistent state; The system monitoring and alarm unit monitors the data management method and the operating status of the system in real time, and automatically triggers the alarm mechanism when anomalies or potential risks are detected.
2. According to claim 1, a data management method for real-time incremental data based on dynamic scheduling and elastic rollback is characterized in that: The data change capture unit includes a change listener, a change extractor, a change storage, and a change distributor. The change listener is used to monitor changes in the data source and generate data change events. The change extractor is used to extract key features of the data change events. The change storage is used to store the key features in a change message queue. The change distributor is used to distribute the change information to corresponding processing modules according to set rules.
3. According to claim 1, a data management method for real-time incremental data based on dynamic scheduling and elastic rollback is characterized in that: The incremental data processing unit includes a parsing and conversion module, a dynamic optimization module, a data sharding module, and a consistency detection module. The parsing and conversion module is used to uniformly convert the original incremental data into an interactive data packet format and add a priority identifier. The dynamic optimization module builds a priority scheduling model based on a lightweight AI model to generate a priority score corresponding to the interactive data packet. The data sharding module divides a large-scale data set into multiple smaller subsets based on horizontal partitioning technology. The consistency detection module ensures data integrity through cross-shard verification and is connected to the elastic rollback unit to ensure data consistency.
4. According to claim 3, a data management method for real-time incremental data based on dynamic scheduling and elastic rollback is characterized in that: The priority scheduling model adopts a neural network architecture or a reinforcement learning algorithm to jointly model and weight input features including business criticality classification, data change hierarchy, transmission time limit margin, protocol data unit size, and link load index, and finally outputs a priority score, wherein the business criticality classification is classified according to the importance of the data to the business operation, the data change hierarchy is divided into the architecture level, application level, and user level according to the source of the incremental data change, and the transmission time limit margin represents the estimated time required for the data to reach the target system and is dynamically calculated through historical data transmission speed and current network conditions.
5. According to claim 1, a data management method for real-time incremental data based on dynamic scheduling and elastic rollback is characterized in that: The real-time update execution unit uses an optimistic lock or pessimistic lock mechanism to control concurrent access to prevent data conflicts, and records the operation log of each update so that a rollback operation can be performed when necessary. It also has a data conflict detection and resolution mechanism, breakpoint resumption function, data verification and validation function, and data backup and recovery function.
6. The data management method for real-time incremental data based on dynamic scheduling and elastic rollback according to claim 1 is characterized in that: The elastic rollback mechanism unit includes an operation log, redundant backup data and a predefined rollback strategy.
7. The data management method for real-time incremental data based on dynamic scheduling and elastic rollback according to claim 1 is characterized in that: The system monitoring and alarm unit collects and analyzes key performance indicators including data update success rate, rollback operation frequency, CPU usage, memory occupancy, disk I / O performance, and network bandwidth utilization in real time through integrated monitoring tools.
8. The data management method for real-time incremental data based on dynamic scheduling and elastic rollback according to claim 1 is characterized in that: The data change capture unit adopts an efficient listening mechanism to capture data change information, parses the captured original change data to obtain the specific content and type of the change, filters out unnecessary change data according to preset rules, stores the parsed change data in a structured manner, and then pushes the captured change data to the incremental update module to update the data in other systems or databases in real time, maintain data consistency, and roll back to the specified historical state according to the stored change data when necessary.
9. The data management method for real-time incremental data based on dynamic scheduling and elastic rollback according to claim 3 is characterized in that: The data sharding module divides the shards according to the range of data keys such as timestamps and numerical intervals, and allocates incremental data within a certain time period to the same shard. Each shard can be stored independently on different nodes or servers, and can independently perform parsing, compression, verification and other operations, using multi-node parallel acceleration of the processing flow, and the sharded data can be efficiently managed through a distributed architecture.
10. The data management method for real-time incremental data based on dynamic scheduling and elastic rollback according to claim 3, characterized in that: The priority scoring model of the dynamic optimization module generates a priority score through multi-dimensional features (business criticality classification, data change hierarchy, transmission time limit margin, protocol data unit size, link load index); the transmission time limit margin calculation formula is:. Among them, TTL is the transmission time limit margin, D current is the real-time delay of the current data transmission, n represents the number of levels of data change, which is used to track the transmission path of data from the source to the target system; D i represents the estimated delay of data transmission between each layer, T i represents the actual time of data transmission between each layer, α is the weight coefficient of the historical transmission rate mean for the calculation of the transmission time limit margin, β is the weight coefficient of the current network quality factor for the calculation of the transmission time limit margin, B avail Represents the current available network bandwidth, P loss is the packet loss rate, reflecting the stability of the current network, T elapsed It represents the time that the data has spent from the source to the present, and is used to dynamically adjust the transmission strategy.
Citation Information
Cited By
Industrial control data processing system and method based on embedded real-time operating system
CN120762376A
Data synchronization method, data synchronization device, access control system and storage medium
CN121193758A
Automatic processing method of experimental process state data based on artificial intelligence
CN122065004A