Double-channel dirty data processing system and method based on dynamic priority
By introducing a dual-channel dirty data processing system based on dynamic priority in the database, dynamically evaluate data priority and allocate processing channels, the problem of inefficient dirty data processing in traditional methods is solved, efficient and flexible dirty data processing is achieved, and the usability and compliance of the system are improved.
Patent Information
- Application Number
- CN202510565556.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-30
AI Technical Summary
Traditional database dirty data processing methods cannot flexibly process according to the urgency and importance of dirty data, resulting in inefficient processing and may affect the normal operation of the system.
A dual-channel dirty data processing system based on dynamic priority is proposed, including a dirty data monitoring module, a dynamic priority evaluation module, a hot repair channel module, a cold processing channel module, a channel collaboration mechanism and a logging and monitoring module. By dynamically evaluating data priority and assigning processing channels, efficient dirty data processing is achieved.
Through the dynamic priority scheduling mechanism, high-risk data is processed in a timely manner, low-priority tasks are delayed, business interruption time is reduced, repair efficiency is improved, sensitive data risks are reduced, and system availability and compliance are improved through hierarchical processing and resource isolation architecture.
Smart Images

Figure CN120066752A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of database management, and in particular relates to a dual-channel dirty data processing system and method based on dynamic priority. Background Art
[0002] In a database system, the consistency and integrity of data are of vital importance. However, due to various unforeseen factors, such as hardware failures, software errors, network problems, etc., it may lead to the situation of dirty data in some data. The existence of dirty data will affect the performance, accuracy of the database and the reliability of business decisions. Traditional dirty data processing methods often adopt a single processing method and cannot flexibly process according to the urgency and importance of dirty data, resulting in low processing efficiency and even possibly affecting the normal operation of the system. Therefore, there is an urgent need for a system and method that can efficiently process according to the dynamic priority of dirty data.
[0003] In the prior art, dirty data cleaning usually adopts an offline batch processing mode, which has the following defects: 1. Great business impact: Full-scale scanning or table locking causes online services to be interrupted (such as frequent table locking of the order table during large-scale e-commerce promotions). 2. Low repair efficiency: Relying on manually formulated rules, it cannot dynamically adapt to changes in data distribution (such as new abnormal data types not being processed in time). 3. Sensitive data risk: Without distinguishing data priorities, high-risk data (such as user payment information) and low-frequency data are processed together, increasing the risk of leakage. Summary of the Invention
[0004] In view of this, the present invention aims to propose a dual-channel dirty data processing system and method based on dynamic priority to solve the problem that traditional database dirty data processing methods cannot flexibly process according to the urgency and importance of dirty data.
[0005] To achieve the above object, the technical solution of the present invention is realized as follows: In a first aspect, the present invention proposes a dual-channel dirty data processing system based on dynamic priority, including a dirty data monitoring module, a dynamic priority evaluation module, a hot repair channel module, a cold processing channel module, a channel cooperation mechanism, and a log recording and monitoring module. The dirty data monitoring module outputs dirty data information to the dynamic priority evaluation module, the dynamic priority evaluation module outputs processed data information to the hot repair channel module or the cold processing channel module, the channel cooperation mechanism exchanges data information with the hot repair channel module and the cold processing channel module respectively, and the dirty data monitoring module, the dynamic priority evaluation module, the hot repair channel module, the cold processing channel module, and the channel cooperation mechanism all output log information to the log recording and monitoring module.
[0006] Second aspect, based on the same inventive concept, the present invention also proposes a dual-channel dirty data processing method based on dynamic priority, including the following steps: S1. The dirty data monitoring module monitors the data status in the database in real time: captures dirty data information based on preset rules and algorithms, and sends it to the dynamic priority evaluation module; S2. After receiving the dirty data information, the dynamic priority evaluation module automatically divides the data priority based on the data access frequency and field sensitivity, and sends the data information to the hot fix channel module or the cold processing channel module according to the data priority; S3. After receiving the high-priority data information, the hot fix channel module repairs the data in real time; S4. After receiving the low-priority data information, the cold processing channel module performs asynchronous batch processing on the data; S5. The channel coordination mechanism isolates the resources for hot fix and controls the data versions of hot fix and cold processing in real time; S6. The log recording and monitoring module is responsible for recording and monitoring the relevant information of dirty data and its processing process in real time.
[0007] Further, in step S1, the dirty data monitoring module monitors the data status in the database in real time, including: S11. The data acquisition layer captures change events through the database log; S12. Verifies the changed data through preset rules; S13. If dirty data information is captured, it is sent to the log recording and monitoring module and the dynamic priority evaluation module.
[0008] Further, in step S2, automatically dividing the data priority based on the data access frequency and field sensitivity includes: S21. Receive the dirty data information sent by the dirty data monitoring module; S22. The dynamic weight adjustment mechanism makes a priority judgment, and the formula of the priority expression is as follows: Priority value = (α × sensitivity coefficient) + (β × access frequency) + (γ × recent modification time decay factor) + (δ × data volume); In the formula, α is the sensitivity coefficient weight, β is the access frequency weight, γ is the time decay factor weight, δ is the data volume factor weight, the sensitivity coefficient is automatically mapped based on the field type, the access frequency statistical period is the number of queries within 24 hours, and the formula of the recent modification time decay factor expression is as follows: 1 / (1 + e^(-λ(t_current - t_last_modified))); Where t_current is the current time for data evaluation, t_last_modified is the time point of the most recent change, and λ is the adjustment coefficient; The formula for the data magnitude expression is as follows: Data magnitude = log10(V) / log10(1TB); Where V is the actual data volume of the field or table; S23. If the priority is P0 - P1, send the data information to the hotfix channel module; S24. If the priority is P2, send the data information to the cold processing channel module; S25. Send the log information to the log record and monitoring module.
[0009] Further, in step S3, after the hotfix channel module receives the high-priority data information, it performs real-time repair on the data, including: S31. Receive the data information sent by the dynamic priority evaluation module; S32. If the priority is P0, store the data in an independent memory area and replace the legal value; S33. If the priority is P1, check if there is any data with priority P0 that has not been repaired. If so, perform the operation in step S32. Otherwise, cache the data with priority P1 in the corresponding memory area and replace the legal value; S34. Transaction log double writing: Use a lock-free hash table, implement memory data replacement using compare-and-swap operations, and pre-allocate a memory pool; Adopt a double-write log consistency protocol to write to the database transaction log and the independent repair log simultaneously, and complete the atomic commit of the two through a log position identifier, and send the log information to the log record and monitoring module.
[0010] Further, in step S4, after the cold processing channel module receives the low-priority data information, it performs asynchronous batch processing on the data, including: S41. Receive the data information sent by the dynamic priority evaluation module; S42. Distributed task sharding: Divide the data information with priority P2 into multiple shards and distribute them to multiple nodes for parallel processing through the Kafka message queue; S43. Batch mode optimization: Combine multiple updates into a single batch SQL, and combine single updates into batch operations; S44. Asynchronous verification: Generate a data snapshot before the update. If the repair fails, roll back to the snapshot version. After the cold processing is completed, verify the repair result through sampling. Failed tasks are automatically retried, and an alarm notification is triggered. If the retry is 3 times, mark it as a "manual review" task.
[0011] Further, in step S5, the channel collaboration mechanism isolates resources for hotfix and controls the data versions of hotfix and cold processing in real time, including: S51. Resource isolation: The hotfix thread pool is independent of the business thread pool, and the maximum concurrency is restricted through the Hystrix-style isolation strategy; S52. Memory area division: Hotfix uses off-heap memory of the Java Virtual Machine; S53. Data version control: Multi-version concurrent control is adopted. Hotfix modifies the current active version, and cold processing repairs based on the historical snapshot version. Conflicts are prevented through the global version number.
[0012] Further, in step S6, the log recording and monitoring module is responsible for real-time recording and monitoring of relevant information about dirty data and its processing process, including: S61. Receive the data information sent by the dirty data monitoring module; S62. Receive the repair information from the hotfix channel module and the cold processing channel module, and mark the repaired fields and processing results; S63. Receive the version information of the channel collaboration mechanism and mark the data version number; S64. If the number of backlogged data in the cold processing queue exceeds 100,000, perform capacity expansion processing.
[0013] Compared with the prior art, the dual-channel dirty data processing system and method based on dynamic priority according to the present invention have the following advantages: Advantage 1: Dynamic priority scheduling mechanism Based on the business impact weight (such as the amount involved in transactions, the user scale) and sensitivity level (such as payment information, privacy fields) of dirty data, the processing priority is dynamically allocated, and the task scheduling is driven by a real-time weight calculation model. Business impact weight: By real-time statistics of business metrics associated with dirty data (such as order amount, user activity), the degree of its impact on the core business is quantified. Sensitivity level: Based on the data classification and grading strategy, high-risk fields (such as payment card numbers) are marked and given high priority.
[0014] The corresponding beneficial effects are as follows: (1) Solve the defect of "great business impact": High-risk data (such as payment information) is processed first to avoid lock tables or long transactions blocking online services; low-priority tasks (such as historical order error correction) are postponed for processing to reduce the business interruption time.
[0015] (2) Solve the defect of "low repair efficiency": Dynamically adjust according to the weight, dynamically adapt to the change of data distribution, and timely process new abnormal data types.
[0016] (3) Defects in solving "sensitive data risks": Distinguish data priorities to avoid mixed operations with ordinary data.
[0017] Advantage 2: Hierarchical processing and resource isolation architecture Adopt a hierarchical processing channel to isolate resources according to data sensitivity and business importance, and optimize storage access through hot and cold data diversion technology. Channel isolation design: High-risk data is processed in real time through a cache unit (such as Redis), and low-frequency data is processed in batches through object storage. Resource allocation strategy: High-priority tasks exclusively occupy computing nodes, and low-priority tasks share idle resources to improve resource utilization.
[0018] The corresponding beneficial effects are as follows: (1) Solve the defect of "great impact on business": Physically isolate online services from low-priority repair tasks, reduce lock table operations by 100%, and the database availability reaches 99.99%.
[0019] (2) Solve the defect of "sensitive data risks": Encrypt sensitive data such as payment information separately, control access rights throughout the link, and reduce compliance risks. Brief Description of the Drawings
[0020] The drawings forming a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings: Figure 1 Schematic diagram of the execution process between functional modules described in the embodiments of the present invention; Figure 2 Schematic diagram of the execution process of the dirty data monitoring module described in the embodiments of the present invention; Figure 3 Schematic diagram of the execution process of the dynamic priority evaluation module described in the embodiments of the present invention; Figure 4 Schematic diagram of the execution process of the hot fix channel module described in the embodiments of the present invention; Figure 5 Schematic diagram of the execution process of the cold processing channel module described in the embodiments of the present invention; Figure 6 Schematic diagram of the execution process of the channel cooperation mechanism described in the embodiments of the present invention. Detailed Embodiments
[0021] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.
[0022] In the description of the present invention, it should be understood that the orientation or positional relationships indicated by the terms "center", "longitudinal", "transverse", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. are based on the orientation or positional relationships shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation to the present invention. In addition, the terms "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first", "second", etc. may explicitly or implicitly include one or more of such features. In the description of the present invention, unless otherwise specified, the meaning of "a plurality" is two or more.
[0023] In the description of the present invention, it should be noted that unless otherwise clearly defined and limited, the terms "installed", "connected", "coupled" should be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or an integral connection; it may be a mechanical connection or an electrical connection; it may be directly connected or indirectly connected through an intermediate medium, and it may be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood through specific circumstances.
[0024] The present invention will be described in detail below with reference to the drawings and in conjunction with embodiments.
[0025] As Figure 1 shown, a dual-channel dirty data processing system and method based on dynamic priority includes a dirty data monitoring module, a dynamic priority evaluation module, a hot repair channel module, a cold processing channel module, a channel cooperation mechanism, and a log recording and monitoring module. The dirty data monitoring module outputs dirty data information to the dynamic priority evaluation module. The dynamic priority evaluation module outputs processed data information to the hot repair channel module or the cold processing channel module. The channel cooperation mechanism exchanges data information with the hot repair channel module and the cold processing channel module respectively. The dirty data monitoring module, the dynamic priority evaluation module, the hot repair channel module, the cold processing channel module, and the channel cooperation mechanism all output log information to the log recording and monitoring module.
[0026] A dual-channel dirty data processing method based on dynamic priority includes the following steps: S1. The dirty data monitoring module monitors the data status in the database in real time: captures dirty data information based on preset rules and algorithms, and sends it to the dynamic priority evaluation module; S2. After receiving the dirty data information, the dynamic priority evaluation module automatically divides the data priority based on the data access frequency and field sensitivity, and sends the data information to the hot fix channel module or the cold processing channel module according to the data priority; S3. After receiving the high-priority data information, the hot fix channel module repairs the data in real time; S4. After receiving the low-priority data information, the cold processing channel module processes the data asynchronously in batches; S5. The channel coordination mechanism isolates the resources for hot fix and controls the data versions of hot fix and cold processing in real time; S6. The log recording and monitoring module is responsible for recording and monitoring the relevant information of dirty data and its processing process in real time.
[0027] In this embodiment, as Figure 1 shown, the dirty data monitoring module enters the dynamic weight adjustment mechanism through the dynamic priority evaluation module, and judges whether to enter the hot fix channel module or the cold processing channel module through the dynamic weight adjustment mechanism. Both the hot fix channel module and the cold processing channel module are regulated by the channel coordination mechanism.
[0028] The specific implementation method is as follows: I. Dirty data monitoring module The dirty data monitoring module monitors the data status in the database system in real time, and detects the possible dirty data by matching with the preset rules and algorithms. Once dirty data is found, the relevant information is immediately sent to the priority evaluation module. As Figure 2 shown, the data acquisition layer judges whether the data is normal data through the preset rules. If it is normal data, the monitoring ends directly. If it is abnormal data, it is sent to the dynamic priority evaluation module and the log recording and monitoring module respectively. The specific process is as follows: (1) The data acquisition layer captures the changed events through the database log; (2) Verify the changed data through preset rules such as the mobile phone number is not 11 digits and the amount value exceeds 2 digits after the decimal point; (3) If abnormal data is captured, send it to the log recording and monitoring module and the dynamic priority evaluation module.
[0029] II. Dynamic priority evaluation module After receiving the dirty data information, the dynamic priority evaluation module automatically divides the data level (P0 - P2) based on the data access frequency and field sensitivity. For example, the order amount field is marked as P0 level, and the mobile phone number field in the user login table is marked as P1 level. By comprehensively considering various attributes of the dirty data and business requirements, a priority value is assigned to each dirty data.
[0030] A dynamic weight adjustment mechanism is introduced in priority calculation: Priority value = (α × Sensitivity coefficient) + (β × Access frequency) + (γ × Decay factor of last modification time) + (δ × Data volume); where: (1)The'sensitivity coefficient' is automatically mapped based on field type (e.g., payment information = 1.5, user ID = 1.2, ordinary field = 1.0); (2)The 'access frequency' is the number of queries within a 24-hour statistical period; (3)The formula for the 'decay factor of last modification time' is: 1 / (1+e^(-λ(t_current - t_last_modified))); t_current: The current time, serving as the reference point for calculating the time difference, representing the current moment of data evaluation.
[0031] t_last_modified: The last modification time, recording the freshness of the data and reflecting the time point of its last change.
[0032] λ is the adjustment coefficient, controlling the decay rate of the time difference.
[0033] Typical value: 0.01 ≤ λ ≤ 0.5; The larger λ is: The S-shaped curve is steeper, and a small change in the time difference (t_current - t_last_modified) will cause a drastic fluctuation in the decay factor.
[0034] Example: When λ = 0.5, if the data is not modified within 1 hour, the decay factor ≈ 1 / (1+e^(-0.5*1)) ≈ 0.73; if not modified for 2 hours, the decay factor ≈ 1 / (1+e^(-0.5*2)) ≈ 0.59.
[0035] The smaller λ is: The curve is flatter, and the impact of the time difference is more gradual. Example: When λ = 0.1, the change in the decay factor for the same time difference is smaller (1 hour ≈ 0.95, 2 hours ≈ 0.91).
[0036] (4)Data volume Normalization of data volume: Map the original data volume to a unified numerical range (e.g., 0~1) for addition with other parameters (sensitivity coefficient, access frequency): Data volume = log10(V) / log10(1TB); where, V: The actual data volume of the field or table (unit: byte, Byte).
[0037] Normalize the data volume to the 0~1 interval, and truncate it to 1 when it is greater than 1.
[0038] (5) α is the weight of the sensitivity coefficient Function: Controls the contribution degree of field sensitivity to the priority.
[0039] Typical value: 0.5 ≤ α ≤ 1.5; Example: If the system pays more attention to sensitive data (such as payment information), the value of α can be increased (such as α = 1.2), so that highly sensitive fields are more likely to be marked as high priority.
[0040] If the business has lower requirements for privacy, the value of α can be decreased (such as α = 0.8) to weaken the impact of sensitivity.
[0041] (6) β is the weight of the access frequency Function: Measures the contribution of data access popularity to the priority.
[0042] Typical value: 0.3 ≤ β ≤ 1.0; Example: For a hot table with high-frequency access (such as an order table), increasing the value of β (such as β = 1.0) can give priority to repairing frequently accessed fields and reduce business latency.
[0043] For a cold data table (such as a log table), decreasing the value of β (such as β = 0.3) can avoid frequent updates from interfering with low-frequency services.
[0044] (7) γ is the weight of the time decay factor Function: Adjusts the impact of data "freshness" on the priority.
[0045] Typical value: 0.1 ≤ γ ≤ 0.5; Example: If it is necessary to give priority to processing recently modified data (such as real-time transaction flow), increase the value of γ (such as γ = 0.4) so that the just modified fields can trigger repairs faster.
[0046] If the repair priority of historical data is relatively low, the value of γ can be decreased (such as γ = 0.1).
[0047] (8) δ is the weight of the data magnitude factor Function: Used to adjust the impact of data scale on the priority Typical value: 0 ≤ δ ≤ 0.5 If δ = 0.5 and the data magnitude is 1, the contribution to the priority is 0.5 × 1 = 0.5; If δ = 0, the impact of the data magnitude is completely ignored.
[0048] Business adaptation: High-concurrency write scenarios (such as e-commerce order tables): Increase the δ value (e.g., δ = 0.4) to prioritize the processing of large-scale data to avoid write bottlenecks; Low-frequency analysis scenarios (such as historical log tables): Decrease the δ value (e.g., δ = 0.1) to avoid excessive resource consumption by large table repairs.
[0049] Through the dynamic weight adjustment mechanism, the system can flexibly adapt to different scenario requirements (such as real-time vs cost sensitivity) to achieve fine-grained control of dynamic priorities.
[0050] As Figure 3 shown, the specific processing flow of the dynamic priority evaluation module is as follows: (1) Receive data sent by the dirty data monitoring module; (2) The dynamic weight adjustment mechanism makes a priority judgment; (3) If the priority is P0 - P1, send it to the hot fix channel module; (4) If the priority is P2, send it to the cold processing channel module; (5) Send the data and priority records to the log recording and monitoring module.
[0051] III. Hot Fix Channel Module Perform real-time repairs on high-priority data (such as high-frequency access tables, critical business fields), and reduce the impact on the business through in-memory computing and lock-free operations. The hot fix channel has the following two characteristics: In-memory real-time repair: Cache P0 - P1 level data in high-speed cache units such as Redis, Memcached, or local cache, and achieve fast reading and writing through a lock-free hash table. For example, when detecting "incorrect mobile phone number format", directly replace the legal value in memory to avoid frequent disk reads and writes. Among them, the P0 level has a higher priority, is more involved in highly sensitive data, uses an independent memory area, and is preferentially processed.
[0052] Transaction log dual writing: Record the repair operation to both the database transaction log and an independent "repair log" to ensure data consistency and traceability.
[0053] Implementation method: Lock-free hash table implementation: Use the Compare And Swap (CAS) operation to implement in-memory data replacement, and combine with a pre-allocated memory pool to reduce the pressure of Garbage Collection (GC); Double-Write Log Consistency Protocol: The repair operation needs to write to both the database transaction log and the independent repair log simultaneously, and ensure the atomic commit of both through the Log Position Identifier.
[0054] As Figure 4 shown, the specific implementation process is as follows: (1) Receive the data and priority level sent by the dynamic priority evaluation module; (2) Judge the priority. When the priority is P0, store the data in the independent memory area for P0-level processing, replace the legal value in the memory, and record the repair operation in the log record and monitoring module; (3) When the priority is P1, check if there is any un-repaired P0-level data. If there is, perform the second step operation. If not, cache the P1-level data in the memory area for P1-level processing, replace the legal value in the memory, and record the repair operation in the log record and monitoring module.
[0055] IV. Cold Processing Channel Module Asynchronous batch processing for low-priority data (such as log tables and historical data), and use distributed task queues and batch commits to optimize resource consumption. The cold processing channel has the following three characteristics: Distributed task sharding: Divide the data with priority level P2 into multiple shards, and distribute them to multiple nodes for parallel processing through the Kafka message queue.
[0056] Distributed task sharding strategy: Use the consistent hashing algorithm to shard the P2-level data into Kafka partitions, ensure that all shards of the same data are processed by the same consumer, and avoid cross-node data competition. The message queue can be Kafka, RabbitMQ, or RocketMQ, which does not affect the core logic of the system architecture.
[0057] Batch mode optimization: Merge multiple updates into a single batch SQL (such as 'UPDATE table SET field = CASE id WHEN... THEN... END'), reducing the database connection overhead.
[0058] Batch SQL optimization example: UPDATE order_table SET amount = CASE id WHEN 1001 THEN 299.9 WHEN 1002 THEN 199.9 ELSE amount END WHERE id IN (1001, 1002); Consolidate single updates into batch operations to reduce the consumption of database connection counts.
[0059] Asynchronous verification mechanism: Generate a data snapshot before the update. When the repair fails, roll back to the snapshot version. After the cold processing is completed, verify the repair result by sampling (such as randomly sampling 5% of the data). Failed tasks are automatically retried, and an alarm notification is triggered. After retrying 3 times, the task is marked as a "manual review" task.
[0060] As Figure 5 shown, the specific process is as follows: (1) Receive the data and priority level sent by the dynamic priority evaluation module; (2) Distribute the data to multiple node partitions; (3) Generate snapshot version data before the update; (4) Perform batch processing to update the data; (5) If the update fails, roll back to the snapshot version, perform the update process again. After retrying three times, trigger an alarm notification and mark it as "manual review" in the log recording and monitoring module; (6) If the update is successful, sample 5% of the data to verify the repair result. If the verification passes, record it in the log recording and monitoring module; if the verification fails, perform the fifth step.
[0061] V. Channel coordination mechanism Resource isolation: Allocate an independent thread pool and memory area for hotfix to avoid competing for resources with business queries.
[0062] Data version control: Hotfix only modifies the current active data version, and cold processing is based on snapshots for repair to prevent interference with each other.
[0063] Implementation method: Resource isolation strategy: The hotfix thread pool is independent of the business thread pool, and the maximum concurrency is restricted through the Hystrix-style isolation strategy; Memory area division: The hotfix uses off-heap memory (DirectByteBuffer) of the Java Virtual Machine (abbreviation: JVM) to avoid the impact of GC pauses on the business.
[0064] Data version control: Adopt Multiversion Concurrency Control (full English name: Multiversion Concurrency Control, abbreviation: MVCC). Hotfix modifies the current active version, and cold processing is based on the historical snapshot version for repair. Prevent conflicts through a global version number (such as'version = timestamp + unique identifier, that is, version = timestamp + UUID.hashCode()') to ensure global uniqueness.
[0065] As shown below Figure 6 The specific process is as follows: For the hot fix channel: (1) Generate the current global version number; (2) Read the current global version number as the baseline version for this hot fix; (3) When the hot fix channel receives P0 - P1 level data repair tasks, allocate independent memory; (4) After the repair is completed, generate a new global version number, update the memory version number to the new global version number and synchronize it to the log record and monitoring module.
[0066] For the cold processing channel: (1) When the cold processing queue receives P2 level data repair tasks, read the current version number and mark the snapshot version as the current version number; (2) After the cold processing task is completed, generate a new global version number; (3) During the execution of the cold processing task, the hot fix may have updated the global version number. Through version number comparison, the data is finally consistent, update the global version number and synchronize it to the log record and monitoring module.
[0067] VI. Log Record and Monitoring The log management module is responsible for recording relevant information about dirty data and the entire processing process, including the discovery time, priority, processing method, processing result, etc. of dirty data. At the same time, the log management module will also monitor and audit the processing process in real time to promptly discover and solve potential problems. If abnormal situations or processing results do not meet expectations are found, the scheduling control module can adjust the processing strategy or re - allocate tasks based on the log information.
[0068] Audit log content: Record the field - level repair details.
[0069] Alarm mechanism: Set a threshold to trigger capacity expansion (for example, when the backlog in the cold processing queue exceeds 100,000 items, automatically increase the number of Kafka consumer instances).
[0070] The specific process is as follows: (1) Receive the data information sent by the dirty data monitoring module; (2) Receive the repair information from the hot fix channel module and the cold processing channel module, and mark the repaired fields and processing results; (3) Receive the version information of the channel coordination mechanism and mark the data version number; (4) When the backlog of data in the cold processing queue exceeds 100,000 items, perform capacity expansion processing.
[0071] The advantages and beneficial effects of the present invention are as follows: Advantage 1: Dynamic Priority Scheduling Mechanism Technical Features: Based on the business impact weight of dirty data (such as transaction amount, user scale) and sensitivity level (such as payment information, privacy fields), dynamically allocate processing priorities, and drive task scheduling through a real-time weight calculation model. Analysis of Technical Solutions: Business Impact Weight: Quantify the impact on core business by real-time statistics of business metrics associated with dirty data (such as order amount, user activity).
[0072] Sensitivity Level: Based on the data classification and grading strategy, mark high-risk fields (such as payment card numbers) and assign high priorities. Beneficial Effects: 1. Solve the defect of "great business impact": High-risk data (such as payment information) is processed first, avoiding table locking or long transactions from blocking online services.
[0073] Low-priority tasks (such as historical order correction) are postponed, reducing service interruption time.
[0074] 2. Solve the defect of "low repair efficiency": Dynamically adjust according to weights, adapt to changes in data distribution, and process newly added abnormal data types in a timely manner.
[0075] 3. Solve the defect of "sensitive data risk": Distinguish data priorities to avoid mixed operations with ordinary data.
[0076] Advantage 2: Hierarchical Processing and Resource Isolation Architecture Technical Features: Adopt hierarchical processing channels, isolate resources according to data sensitivity and business importance, and optimize storage access through hot and cold data diversion technology. Analysis of Technical Solutions: Channel Isolation Design: High-risk data is processed in real time through a cache unit (such as Redis), and low-frequency data is processed in batches through object storage.
[0077] Resource Allocation Strategy: High-priority tasks occupy computing nodes exclusively, and low-priority tasks share idle resources, improving resource utilization. Beneficial Effects: 1. Solve the defect of "great business impact": Online services are physically isolated from low-priority repair tasks, eliminating 100% of table locking operations, and the database availability reaches 99.99%. 2. Solve the defect of "sensitive data risk": Sensitive data such as payment information is encrypted separately, and full-link access rights are controlled, reducing compliance risks.
[0078] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A dual-channel dirty data processing method based on dynamic priority, implemented by a dual-channel dirty data processing system based on dynamic priority, characterized in that: The dual-channel dirty data processing system based on dynamic priority includes a dirty data monitoring module, a dynamic priority evaluation module, a hot repair channel module, a cold processing channel module, a channel coordination mechanism and a log recording and monitoring module, wherein the dirty data monitoring module outputs dirty data information to the dynamic priority evaluation module, the dynamic priority evaluation module outputs processed data information to the hot repair channel module or the cold processing channel module, the channel coordination mechanism exchanges data information with the hot repair channel module and the cold processing channel module respectively, and the dirty data monitoring module, the dynamic priority evaluation module, the hot repair channel module, the cold processing channel module and the channel coordination mechanism all output log information to the log recording and monitoring module; The method comprises the following steps: S1. The dirty data monitoring module monitors the data status in the database in real time: it captures dirty data information based on preset rules and algorithms and sends it to the dynamic priority evaluation module; S2. After receiving the dirty data information, the dynamic priority assessment module automatically prioritizes the data based on the data access frequency and field sensitivity, and sends the data information to the hot repair channel module or the cold processing channel module according to the data priority; S3, after receiving the high-priority data information, the hot repair channel module repairs the data in real time; S4, after receiving the low-priority data information, the cold processing channel module performs asynchronous batch processing on the data; S5. The channel coordination mechanism isolates resources for hot fixes and controls the data versions of hot fixes and cold treatments in real time; S6. The logging and monitoring module is responsible for real-time recording and monitoring of the relevant information of dirty data and its processing process.
2. The dual-channel dirty data processing method based on dynamic priority according to claim 1, characterized in that: In step S1, the dirty data monitoring module monitors the data status in the database in real time, including: S11, the data collection layer captures change events through database logs; S12. Verify the changed data using preset rules; S13. If dirty data information is captured, it is sent to the logging and monitoring module and the dynamic priority evaluation module.
3. The dual-channel dirty data processing method based on dynamic priority according to claim 2, characterized in that: In step S2, data priorities are automatically divided based on data access frequency and field sensitivity, including: S21, receiving dirty data information sent by the dirty data monitoring module; S22. The dynamic weight adjustment mechanism performs priority judgment. The formula of the priority expression is as follows: Priority value = (α×sensitivity coefficient) + (β×access frequency) + (γ×last modification time attenuation factor) + (δ×data magnitude); In the formula, α is the sensitivity coefficient weight, β is the access frequency weight, γ is the time decay factor weight, δ is the data magnitude factor weight, the sensitivity coefficient is automatically mapped based on the field type, the access frequency statistical period is the number of queries within 24 hours, and the formula for the most recently modified time decay factor expression is as follows: 1 / (1+e^(-λ(t_current-t_last_modified))); In the formula, t_current is the current time of data evaluation, t_last_modified is the time point of the most recent change, and λ is the adjustment coefficient; The formula for the data magnitude expression is as follows: Data volume = log10(V) / log10(1TB); Where V is the actual data volume of the field or table; S23, if the priority is P0-P1, sending data information to the hot repair channel module; S24, if the priority is P2, sending the data information to the cold processing channel module; S25. Send the log information to the log recording and monitoring module.
4. The dual-channel dirty data processing method based on dynamic priority according to claim 3 is characterized in that: In step S3, after receiving the high-priority data information, the hot repair channel module performs real-time repair on the data, including: S31, receiving data information sent by the dynamic priority evaluation module; S32, if the priority is P0, the data is stored in an independent memory area and replaced with a legal value; S33, if the priority is P1, check whether there is unrepaired data with a priority of P0, if so, proceed to step S32, otherwise, cache the data with a priority of P1 to the corresponding memory area and replace the legal value; S34, double writing of transaction log: using lock-free hash table, adopting compare and exchange operation to realize memory data replacement, and pre-allocating memory pool; using double writing log consistency protocol, writing database transaction log and independent repair log at the same time, and completing atomic commit of both through log location identifier, and sending log information to log recording and monitoring module.
5. The dual-channel dirty data processing method based on dynamic priority according to claim 4 is characterized in that: In step S4, after receiving the low-priority data information, the cold processing channel module performs asynchronous batch processing on the data, including: S41, receiving data information sent by the dynamic priority evaluation module; S42, distributed task sharding: divide the data information with priority P2 into multiple shards, and distribute them to multiple nodes for parallel processing through the Kafka message queue; S43, batch mode optimization: merge multiple updates into a single batch SQL, merge single updates into batch operations; S44, asynchronous verification: Generate a data snapshot before updating. If the repair fails, roll back to the snapshot version. After the cold treatment is completed, verify the repair result through sampling. Failed tasks are automatically retried and an alarm notification is triggered. If they are retried three times, they are marked as "manual review" tasks.
6. The dual-channel dirty data processing method based on dynamic priority according to claim 5, characterized in that: In step S5, the channel coordination mechanism isolates resources for hot repair and controls the data versions of hot repair and cold processing in real time, including: S51, Resource isolation: The hot repair thread pool is independent of the business thread pool, and the maximum number of concurrency is limited through the Hystrix-style isolation strategy; S52, Memory area division: Hot fix uses Java Virtual Machine off-heap memory; S53. Data version control: Use multi-version concurrent control, hot fixes to modify the current active version, cold fixes based on historical snapshot versions, and prevent conflicts through global version numbers.
7. The dual-channel dirty data processing method based on dynamic priority according to claim 6, characterized in that: In step S6, the logging and monitoring module is responsible for real-time recording and monitoring of the relevant information of dirty data and its processing process, including: S61, receiving data information sent by the dirty data monitoring module; S62, receiving the repair information of the hot repair channel module and the cold processing channel module, and marking the repair field and the processing result; S63, receiving the channel coordination mechanism version information, and marking the data version number; S64. If the backlog of data in the cold processing queue exceeds 100,000, capacity expansion is performed.
Citation Information
Patent Citations
Rapid dirty data detection and processing method for log collection
CN114356908A
Big data life cycle management method
CN117331931A
Data storage method, system and equipment for territorial resource planning and medium
CN119127897A
CDC synchronization method and system based on OracIeRAC
CN119577038A
Intelligent resource scheduling and dynamic data priority management method and system
CN119739498A