Dynamic configurable data cleaning method and system based on big data
By combining simulated antigen recognition and three-dimensional storage topological space, and dynamically optimized cleaning rules and task scheduling, the problem of identifying and managing hot and cold states of multi-source heterogeneous data in the big data environment is solved, the accuracy and efficiency of data cleaning are improved, and the adaptability and intelligence of the system are enhanced.
Patent Information
- Application Number
- CN202510733091.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-04
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-06-04
AI Technical Summary
The existing dynamic configurable data cleaning methods are difficult to accurately identify the changes in the hot and cold state of multi-source heterogeneous data in big data environments, resulting in confusion of hot and cold data, affecting the priority scheduling and resource allocation of cleaning tasks. In addition, traditional methods fail to respond dynamically to changes in data value density, resulting in insufficiency of cleaning.
The simulated antigen recognition and memory mechanism is used to identify abnormal patterns in heterogeneous data sources, and the data value density scoring function is constructed. Combined with the three-dimensional storage topology space and lattice migration algorithm, the thermal diffusion path is simulated for data migration, and a cleaning flow chart is constructed through drag and drop, the cleaning rules are dynamically optimized, and the heterogeneous processing architecture is constructed for collaborative cleaning.
It realizes accurate identification and dynamic management of hot and cold states of multi-source heterogeneous data, improves the adaptability and execution efficiency of cleaning tasks, reduces manual maintenance costs, and enhances the robustness and intelligence of the system.
Smart Images

Figure CN120256834A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and more specifically, to a dynamic configurable data cleaning method and system based on big data. Background Art
[0002] The patent with the publication number CN105930523A discloses a data cleaning framework based on dynamically configurable rules for a big data environment. The inventive method is a new cross-domain, reusable, configurable method that combines data conversion, data checking, and data repair into one, thereby improving the description ability and execution efficiency of the cleaning process. Experimental results on multiple real datasets show that the system can seamlessly integrate dynamically configurable rules into multiple data sources and various different application fields and be implemented in multiple projects, further verifying the effective role of the method in real scenarios.
[0003] The existing configurable data cleaning methods and systems have the following main problems: In the process of dynamic configurable data cleaning based on big data, the number of heterogeneous data sources is huge, the structure is complex, and the data access patterns are variable. Traditional hot and cold data management methods mainly rely on static thresholds to divide the hot and cold states of data and perform data migration and hierarchical management through manual or periodically triggered methods. Multi-source heterogeneous data shows highly complex and dynamic behavior patterns in actual cleaning applications. Traditional threshold-based hot and cold division is difficult to accurately capture the critical changes in the hot and cold states of data, resulting in confusion between hot and cold data and affecting the priority scheduling and resource allocation of cleaning tasks.
[0004] The intrinsic value of data continuously changes during the big data cleaning process. Traditional methods fail to fully consider the phase transition characteristics of value density changing over time and are difficult to intelligently adjust the cleaning strategy, affecting the cleaning effect and efficiency. In large-scale data storage, there are obvious spatial local aggregation characteristics in the access hotspots of data, but traditional hot and cold data migration mechanisms ignore this spatial correlation, resulting in slow resource scheduling and migration response and affecting the overall performance of the data cleaning system. Existing data migration strategies are mostly periodic or manually triggered, lacking dynamic adaptability and unable to respond to data state changes in real time, resulting in low utilization efficiency of storage resources and computing resources and affecting the timeliness and accuracy of dynamic data cleaning.
[0005] In view of this, the present invention proposes a dynamic configurable data cleaning method and system based on big data to solve the above problems. Summary of the Invention
[0006] To overcome the above defects of the prior art and to achieve the above object, the present invention provides the following technical solution: A dynamic configurable data cleaning method based on big data, comprising: S1. Collect heterogeneous data sources, and connect to heterogeneous data sources by combining with data source adaptation interfaces; utilize the simulated antigen recognition and memory mechanism to identify abnormal patterns in heterogeneous data sources and encode them as antigen fingerprints; construct a data value density scoring function based on antigen fingerprints to obtain a data value density score; S2. Construct a three-dimensional storage topology space according to the data value density score, adopt a lattice migration algorithm to automatically identify the critical points of hot and cold data, trigger hierarchical migration when the heterogeneous data source cools down to a preset temperature threshold, introduce a thermodynamic diffusion mechanism, and simulate the migration process of the heterogeneous data source in the three-dimensional storage topology space as a heat diffusion path; S3. Construct a data cleaning flow chart by dragging, determine the execution order of cleaning operators in the data cleaning flow chart based on topological sorting, and automatically schedule cleaning tasks; S4. Construct a Darwinian evolution system for cleaning rules, evolve the cleaning rules periodically, define a cleaning rule fitness function, simulate the natural selection mechanism to dynamically optimize the cleaning strategy, and output an evolved rule instruction set; S5. Construct a heterogeneous processing architecture, customize a VHDL kernel to clean and process heterogeneous data sources; cooperate based on the heat diffusion path and the Spark batch reading mechanism, and call the evolved rule instruction set to complete a preset data cleaning task; S6. Construct a double-chain structure with an operation evidence chain and a quality certification chain to audit and verify the heterogeneous data source after cleaning processing; use a verifiable digest function to record the cleaning credibility of the cleaned heterogeneous data source; trigger a re-cleaning mechanism when the cleaning credibility is less than a preset cleaning credibility threshold, and update the cleaning rules.
[0007] Preferably, the method for identifying and encoding abnormal patterns in heterogeneous data sources as antigen fingerprints includes: Heterogeneous data sources include structured data, semi-structured data, unstructured data, streaming data, and system operation and maintenance data; perform standard deviation normalization processing on heterogeneous data sources to convert them into a standardized data format; Construct a simulated antibody library, and define each antibody in the simulated antibody library as an adaptable and mutable antibody template; the antibody template is any one of a template based on the K-nearest neighbor model, a template based on a clustering center vector, or a template based on a small-sample neural network; Use a similarity detection mechanism to calculate the matching degree between the heterogeneous data source and the antibody template; the similarity detection mechanism is any one of cosine similarity, Euclidean distance, or kernel function similarity; if the matching degree is less than a preset matching degree threshold, it is identified as an abnormal pattern and marked as a simulated antigen; The simulated antigen is fingerprinted, and the fingerprint encoding process includes feature dimension reduction, fingerprint hash mapping and feature vector normalization to generate an antigen fingerprint containing data identification, anomaly category, timestamp and feature hash value.
[0008] Preferably, the method for obtaining the data value density score includes: Extract multi-dimensional information from antigen fingerprints and construct data feature vectors. The multi-dimensional information includes antigen occurrence frequency, abnormality level, fingerprint similarity distribution, historical trigger records and context information entropy. A data value density scoring function is constructed based on the data feature vector. The data value density scoring function is a nonlinear weighted function. Each antigen fingerprint is input into the corresponding data value density scoring function to calculate the corresponding data value density score.
[0009] Preferably, the method for constructing the three-dimensional storage topological space includes: Design a three-dimensional coordinate system for the three-dimensional storage topological space, combine X, Y, and Z as the three-dimensional coordinates of the three-dimensional storage topological space, organize all three-dimensional coordinates into integer grids, and form a logical three-dimensional storage topological space; each dimension is defined as the X-axis, Y-axis, and Z-axis; the X-axis represents the value density dimension, the Y-axis represents the data structure dimension, and the Z-axis represents the time life cycle dimension; Establish grading and coding rules for each dimension. For the X-axis, normalize all data value density scores and linearly map them to the [0,1] scoring interval. Define the stratification levels according to the scoring interval, and preset the first threshold of data value density score, the second threshold of data value density score, and the third threshold of data value density score. If the data value density score is greater than or equal to the second data value density score threshold, and less than or equal to the third data value density score threshold, it is classified as a high value layer; if the data value density score is greater than or equal to the first data value density score threshold, and less than the second data value density score threshold, it is classified as a medium value layer; if it is less than the first data value density score threshold, it is classified as a low value layer; For the Y axis, set static mapping codes for the data structure types included in the data structure dimension, which include structured data, semi-structured data, and unstructured data; for structured data, let Y=0; for semi-structured data, let Y=1; for unstructured data, let Y=2; the data structure dimension is an extensible support subclass, and Y=3 indicates mixed structure data; For the Z-axis, define the time life cycle stage according to the most recent access time and set the mapping rules; preset the first threshold of the most recent access time, the second threshold of the most recent access time, and the third threshold of the most recent access time; when the most recent access time is less than the first threshold of the most recent access time, then let Z = 0; when the most recent access time is greater than or equal to the first threshold of the most recent access time and less than the second threshold of the most recent access time, then let Z = 1; when the most recent access time is greater than the second threshold of the most recent access time, then let Z = 2.
[0010] Preferably, the method for obtaining the heat diffusion path includes: In the three-dimensional storage topology space, regard each coordinate point as a lattice node, and the node contains m data units; for each data unit , define its current heat value as ; regard the three-dimensional coordinate system as a spatial lattice grid, and define the average heat of each lattice node as ; Considering the spatial correlation of heat diffusion, introduce the local heat diffusion coefficient , establish a lattice heat diffusion coupling model, and preset the temperature threshold. When the heterogeneous data source cools down to the preset temperature threshold, trigger the hierarchical migration determination; identify the heterogeneous data source with a temperature less than or equal to the preset temperature threshold as cold data, and identify the heterogeneous data source with a temperature greater than the preset temperature threshold as hot data; Define the data migration trigger mechanism, and perform hierarchical migration operations when the trigger conditions of the data migration trigger mechanism are met; migrate through the heat diffusion path of the heterogeneous data source from the high-temperature lattice node to the low-temperature lattice node, simulate the process of heat gradually releasing and the migration of data access hotspots, and the heat diffusion path is the spatial trajectory of the temperature change of the lattice node and the migration operation.
[0011] Preferably, the method for automatically scheduling cleaning tasks includes: Preset a data cleaning operator library, and select operator nodes from the preset data cleaning operator library through drag-and-drop operations to construct a data cleaning flow chart; generate directed edges of the data cleaning flow chart through the connection relationship between operator nodes, and automatically label the data dependency relationship for each directed edge; Formalize the data cleaning flow chart into a directed acyclic graph, and sort the operator nodes in the directed acyclic graph based on the topological sorting algorithm to generate an operator execution sequence that satisfies the data dependency relationship, where the operator nodes without dependency relationships are marked as executable in parallel; generate executable task units according to the operator execution sequence, and automatically schedule the cleaning tasks to the heterogeneous processing architecture for execution.
[0012] Preferably, the method for obtaining the evolved rule instruction set includes: Preset data cleaning rules, initialize the data cleaning rule set as a rule population, where the cleaning rules are represented in an encoded manner, and the initial rule population is generated from the historical cleaning rule library; define a rule fitness function that comprehensively considers cleaning accuracy, cleaning coverage, and execution efficiency, and comprehensively evaluates the rule performance through weighted summation of weights; In each evolution cycle, retain the rule with the highest fitness selected according to the rule fitness function; perform a crossover operation on the selected cleaning rules to generate new cleaning rules by combining the characteristics of the cleaning rules; apply a mutation operation to the newly generated cleaning rules after crossover; Preset a fitness threshold, replace the cleaning rules with fitness less than the preset fitness threshold with the newly evolved cleaning rules; after L evolution cycles, select the cleaning rules with fitness reaching the preset fitness threshold to form the final rule instruction set.
[0013] Preferably, the method of invoking the evolved rule instruction set to complete the preset data cleaning task includes: Construct a heterogeneous processing architecture consisting of an FPGA acceleration unit, a Spark distributed processing engine, and a data scheduling manager; divide the data cleaning process into three stages: preprocessing, core cleaning, and output verification. In the preprocessing stage, the Spark distributed processing engine batches, pre-caches, and binds cleaning rules to heterogeneous data; In the core cleaning stage, perform data cleaning operations on heterogeneous data through a customized VHDL kernel; in the output verification stage, the Spark distributed processing engine receives the cleaning results from the FPGA acceleration unit and integrates and outputs them, thereby completing the preset data cleaning task; Construct a VHDL cleaning pipeline containing N configurable logic modules based on a stream processing architecture, convert the evolved rule instruction set into an instruction format recognizable by the FPGA, and map it to the corresponding configurable logic module in the VHDL cleaning pipeline according to the instruction content; the configurable logic modules include a field recognition module, a missing value repair module, an outlier replacement module, and a format standardization module, and each module can be enabled as needed to achieve dynamic modular cleaning; Access heterogeneous data sources through the Spark distributed processing engine, construct a data heat perception scheduling mechanism based on the heat diffusion path, and push hot data to the FPGA acceleration unit for cleaning processing first through the data scheduling manager, and cache cold data waiting for scheduling; For each batch of pushed heterogeneous data sources, bind the corresponding rule instruction set identifier, drive the FPGA acceleration unit to load the corresponding configurable logic module and execute the cleaning operation; return the heterogeneous data sources that have completed cleaning from the FPGA acceleration unit to the Spark distributed processing engine for formatting and encapsulation, and output the cleaned heterogeneous data sources.
[0014] Preferably, the method for auditing and verifying the heterogeneous data source after cleaning processing includes: Construct a double-chain structure with an operation evidence chain and a quality certification chain. For each data cleaning operation process, record the data cleaning operation timestamp, the cleaning rule ID called, the data cleaning flow chart, the unique identifier of the heterogeneous data source being processed, the operator, and the algorithm main body ID in real time; use the verifiable digest function Merkle Tree to calculate the digest hash value of the data cleaning operation event to form an operation fingerprint. Write the operation fingerprint as a transaction into the operation evidence chain and form a blockchain structure by linking with the hash of the previous block. The operation evidence chain supports smart contracts. When the data cleaning operation is completed, the smart contract automatically triggers the auditing and verification of the heterogeneous data source after cleaning processing according to the operation fingerprint.
[0015] A dynamic configurable data cleaning system based on big data includes: An adaptive detection module for collecting heterogeneous data sources and accessing heterogeneous data sources in combination with a data source adaptation interface; using a simulated antigen recognition and memory mechanism to identify abnormal patterns in the heterogeneous data source and encode them as antigen fingerprints; constructing a data value density scoring function based on the antigen fingerprints to obtain a data value density score. An intelligent hierarchical storage module that constructs a three-dimensional storage topology space according to the data value density score, uses a lattice migration algorithm to automatically identify the critical point of hot and cold data, triggers hierarchical migration when the heterogeneous data source cools down to a preset temperature threshold, and introduces a thermodynamic diffusion mechanism to simulate the migration process of the heterogeneous data source in the three-dimensional storage topology space as a heat diffusion path. A data cleaning orchestration module that constructs a data cleaning flow chart by dragging and dropping, determines the execution order of cleaning operators in the data cleaning flow chart based on topological sorting, and automatically schedules cleaning tasks. An ecological rule evolution module that constructs a Darwinian evolution system for cleaning rules, evolves the cleaning rules periodically, defines a cleaning rule fitness function, simulates the natural selection mechanism to dynamically optimize the cleaning strategy, and outputs an evolved rule instruction set. A hardware-accelerated cleaning module that constructs a heterogeneous processing architecture, customizes a VHDL kernel to clean the heterogeneous data source; coordinates with the heat diffusion path and the Spark batch reading mechanism, and calls the evolved rule instruction set to complete a preset data cleaning task. A trusted cleaning evidence storage module that constructs a double-chain structure with an operation evidence chain and a quality certification chain to audit and verify the heterogeneous data source after cleaning processing; uses a verifiable digest function to record the cleaning credibility of the cleaned heterogeneous data source; triggers a re-cleaning mechanism when the cleaning credibility is less than a preset cleaning credibility threshold and updates the cleaning rules.
[0016] Compared with the prior art, the present invention has the following beneficial effects: By comprehensively considering the access frequency, update frequency, and data value density changes, the present invention constructs a comprehensive heat metric. Based on the local heat diffusion coupling model, it can capture the cold and hot critical points of multi-source heterogeneous data in the big data environment in real time, greatly improving the accuracy of cold and hot data partitioning, thereby optimizing the priority of data cleaning and the strategy selection. It dynamically reflects the temporal changes and phase transition processes of the data value density, enabling the data cleaning system to automatically adjust the cleaning rules and task scheduling according to the real-time changes in data value, enhancing the self-adaptability and intelligent level of the cleaning process.
[0017] By simulating the spatial diffusion process of heat in thermodynamics, it accurately reflects the migration trajectory of data access hotspots in the three-dimensional storage space, supports the dynamic hierarchical migration strategy based on the heat diffusion path, effectively avoids migration hysteresis, improves the utilization rate of storage resources and the execution efficiency of data cleaning tasks. Using the preset temperature threshold and migration trigger mechanism, it can automatically sense the data temperature change and trigger the migration operation, realizing the real-time dynamic management of cold and hot data, ensuring the efficient execution of data cleaning tasks at the appropriate storage level, and enhancing the robustness and response speed of the system. The dynamic migration mechanism based on the heat diffusion path replaces the traditional manual or periodic migration strategy, realizes the automatic sensing and execution of data cold and hot states and migration operations, greatly reducing the manual maintenance cost, and improving the stability and intelligent degree of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 It is a schematic flow chart of a dynamic configurable data cleaning method based on big data according to the present invention; Figure 2 It is a schematic structural diagram of a dynamic configurable data cleaning system based on big data according to the present invention; Figure 3 It is a schematic flow chart of a method for automatically scheduling cleaning tasks provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0019] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0020] Embodiment 1 Please refer to Figure 1 and Figure 3 As shown, Embodiment 1 further describes a dynamic configurable data cleaning method based on big data proposed by the present invention, including: With the rapid development of big data technology, data cleaning, as a key step in the data preprocessing process, is of great significance for improving data quality and enhancing data availability. Currently, cleaning tasks in the big data environment usually face challenges such as huge data volume, diverse types, complex structures, and strong dynamics. Especially in the scenario of multi-source heterogeneous data access, the accuracy and execution efficiency of cleaning strategies directly affect the effects of subsequent data analysis and intelligent decision-making.
[0021] Existing configurable data cleaning methods and their systems generally organize and schedule the cleaning process based on rule engines and static configurations, and have achieved certain application effects in some structured data cleaning. However, in the actual big data processing scenario, the following main technical problems need to be solved urgently: The number of heterogeneous data sources is huge, the structures are complex, and the access behaviors show highly dynamic and variable characteristics. Traditional hot and cold data management methods often rely on static indicators such as access frequency or last access time, divide data into "hot", "warm", or "cold" states by setting fixed thresholds, and perform data migration and hierarchical processing in an artificial or periodic triggering manner. When facing data with complex behavior patterns and non-linear value evolution, such methods often have difficulty accurately identifying the critical points of hot and cold state changes, resulting in misidentification of hot and cold data, and further affecting the priority judgment of cleaning tasks and the reasonable allocation of computing resources.
[0022] The intrinsic value density of data evolves over time and shows "phase transition" characteristics, that is, the data may change from high value to low value or in the opposite direction within a short time window. Traditional methods have not effectively modeled and responded to this dynamic process, resulting in lagging cleaning strategies and reducing the intelligence and adaptability of data processing.
[0023] The data access hotspots in the current data storage system usually show obvious spatial local aggregation, but traditional hot and cold data management strategies often take the global as the granularity, ignoring the local heat aggregation and diffusion behaviors, resulting in slow response in data migration and resource scheduling, and affecting the overall cleaning efficiency and performance of the system.
[0024] Existing systems mostly trigger data migration and cleaning operations in a fixed cycle or by manual intervention, lacking the ability to dynamically regulate based on real-time data status, and it is difficult to cope with sudden data changes or high-frequency access fluctuations, further restricting the resource utilization rate of the system and the timeliness of task processing.
[0025] In order to effectively solve the above problems, the present invention proposes a dynamic configurable data cleaning system based on big data, including: S1. Collect heterogeneous data sources, and access the heterogeneous data sources in combination with the data source adaptation interface; utilize the simulated antigen recognition and memory mechanism to identify abnormal patterns in the heterogeneous data sources and encode them as antigen fingerprints; construct a data value density scoring function based on the antigen fingerprints to obtain the data value density score. S2. Construct a three-dimensional storage topology space according to the data value density score, adopt the lattice migration algorithm to automatically identify the critical points of hot and cold data, trigger hierarchical migration when the heterogeneous data source cools down to the preset temperature threshold, introduce the thermodynamic diffusion mechanism, and simulate the migration process of the heterogeneous data source in the three-dimensional storage topology space as a heat diffusion path. S3. Construct a data cleaning flow chart by dragging, determine the execution order of the cleaning operators in the data cleaning flow chart based on topological sorting, and automatically schedule the cleaning tasks. S4. Construct a Darwinian evolution system for cleaning rules to evolve the cleaning rules periodically, define a cleaning rule fitness function, simulate the natural selection mechanism to dynamically optimize the cleaning strategy, and output the evolved rule instruction set. S5. Construct a heterogeneous processing architecture, customize the VHDL kernel to clean and process the heterogeneous data source; cooperate based on the heat diffusion path and the Spark batch reading mechanism, and call the evolved rule instruction set to complete the preset data cleaning task. S6. Construct a double-chain structure with an operation evidence chain and a quality certification chain to audit and verify the heterogeneous data source after cleaning; use a verifiable digest function to record the cleaning credibility of the cleaned heterogeneous data source; trigger the re-cleaning mechanism when the cleaning credibility is less than the preset cleaning credibility threshold, and update the cleaning rules.
[0026] The method for identifying abnormal patterns in heterogeneous data sources and encoding them as antigen fingerprints includes: Heterogeneous data sources include structured data, semi-structured data, unstructured data, streaming data, and system operation and maintenance data; perform standard deviation normalization processing on the heterogeneous data sources to convert them into a standardized data format. Construct a simulated antibody library, and define each antibody in the simulated antibody library as an adaptable and mutable antibody template; the antibody template is any one of a template based on the K-nearest neighbor model, a template based on the cluster center vector, or a template based on a small-sample neural network. Use a similarity detection mechanism to calculate the matching degree between the heterogeneous data source and the antibody template; the similarity detection mechanism is any one of cosine similarity, Euclidean distance, or kernel function similarity; if the matching degree is less than the preset matching degree threshold, it is identified as an abnormal pattern and marked as a simulated antigen. Fingerprint encoding is performed on the simulated antigen. The fingerprint encoding process includes feature dimensionality reduction, fingerprint hashing mapping, and feature vector normalization processing to generate an antigen fingerprint containing data identification, anomaly category, timestamp, and feature hash value.
[0027] The method for obtaining the data value density score includes: Extract multi-dimensional information from the antigen fingerprint to construct a data feature vector. The multi-dimensions include the antigen occurrence frequency, anomaly level (obtained based on the deviation between the matching degree and the preset matching degree threshold), fingerprint similarity distribution (the similarity distribution between the current antigen fingerprint and the historical known antigen fingerprints, used to determine whether the anomaly is a new anomaly or a variant anomaly), historical trigger records (whether this type of fingerprint triggered a cleaning task in the past and whether it caused abnormal results), and context information entropy; Construct a data value density scoring function based on the data feature vector. The data value density scoring function is a non-linear weighted function; input each antigen fingerprint into the corresponding data value density scoring function respectively to calculate the corresponding data value density score.
[0028] The method for constructing a three-dimensional storage topological space includes: Design a three-dimensional coordinate system for the three-dimensional storage topological space. Combine X, Y, and Z as the three-dimensional coordinates of the three-dimensional storage topological space. Organize all three-dimensional coordinates in an integer grid to form a logically three-dimensional storage topological space; each dimension is defined as the X-axis, Y-axis, and Z-axis respectively; the X-axis represents the value density dimension, the Y-axis represents the data structure dimension, and the Z-axis represents the time life cycle dimension; The value density dimension measures the potential value of data units in cleaning, mining, and analysis, which is derived from the previously calculated data value density score; the data structure dimension represents the structure type of the identified data, manifested as heterogeneous data sources; the time life cycle dimension represents the time attribute or active state of the data, used to depict its life cycle stage (such as real-time, near-time, medium-term, long-term); Establish a hierarchical and encoding rule for each dimension. For the X-axis, normalize all data value density scores and linearly map them to the [0,1] scoring interval; delimit hierarchical levels according to the scoring interval, and preset the first data value density score threshold, the second data value density score threshold, and the third data value density score threshold; If the data value density score is greater than or equal to the second data value density score threshold and less than or equal to the third data value density score threshold, it is designated as the high-value layer; if the data value density score is greater than or equal to the first data value density score threshold and less than the second data value density score threshold, it is designated as the medium-value layer; if it is less than the first data value density score threshold, it is designated as the low-value layer; For the Y-axis, set static mapping codes for the data structure types included in the data structure dimension. The data structure types include structured data, semi-structured data, and unstructured data. For structured data, let Y = 0; for semi-structured data, let Y = 1; for unstructured data, let Y = 2. The data structure dimension is an extensible support subclass, and when Y = 3, it represents hybrid structure data. For the Z-axis, define the time life cycle stage according to the most recent access time and set the mapping rules. Preset the first threshold of the most recent access time, the second threshold of the most recent access time, and the third threshold of the most recent access time. When the most recent access time is less than the first threshold of the most recent access time, then let Z = 0; when the most recent access time is greater than or equal to the first threshold of the most recent access time and less than the second threshold of the most recent access time, then let Z = 1; when the most recent access time is greater than the second threshold of the most recent access time, then let Z = 2.
[0029] The method for obtaining the heat diffusion path includes: In traditional hot and cold data management, threshold rules are often set through dimensions such as access frequency and last access time to classify data into hot data and cold data, and then the migration strategy is triggered manually or periodically. However, this method has the following limitations: it is unable to accurately identify the critical points of hot and cold changes of multi-source heterogeneous data under complex behavior patterns; it is difficult to perceive the phase transition state of the data value density changing over time; it does not have the ability to dynamically sense local spatial aggregation, resulting in a lag in the migration strategy and low resource utilization.
[0030] Therefore, introduce the lattice migration algorithm, combined with the three-dimensional storage topology, to achieve a more sensitive and accurate hot and cold migration mechanism. The lattice migration algorithm simulates the behavior of state migration caused by the thermal motion of particles in a solid lattice, and establishes a migration mechanism based on local spatial heat diffusion and critical phase transition perception. In the three-dimensional storage topology space, regard each coordinate point as a lattice node, and each node contains m data units. For each data unit , define its current heat value as ; ; where, represents the number of accesses to the data unit per unit time; represents the update (write or modify) frequency of the data unit per unit time; represents the change rate of the value density of the data unit per unit time; represents the current time point; represents the index of the data unit; represents the weight factor of the access frequency; represents the weight factor of the update frequency; The weight factor representing the rate of change of value density; according to the expert experience method, , and range from 0 to 1, and , , sum to 1; it should be noted that the current heat value is obtained through a weighted summation formula, taking into account the access frequency, update frequency, and information value change, which is a comprehensive measurement of data heat; Regarding the three-dimensional coordinate system as a spatial lattice grid, define the average heat of each lattice node as ; ; where represents the average heat of node in the three-dimensional storage topology space at time point ; represents the number of data units contained in lattice node ; considering the spatial correlation of heat diffusion, introduce the local heat diffusion coefficient , establish a lattice heat diffusion coupling model, and preset a temperature threshold. When the heterogeneous data source cools down to the preset temperature threshold, trigger the hierarchical migration determination; identify the heterogeneous data source with a temperature less than or equal to the preset temperature threshold as cold data, and identify the heterogeneous data source with a temperature greater than the preset temperature threshold as hot data; Define the data migration trigger mechanism. When the trigger condition of the data migration trigger mechanism is met, perform hierarchical migration operations; migrate through the heat diffusion path of the heterogeneous data source from the high-temperature lattice node to the low-temperature lattice node, simulating the process of heat gradually releasing and the migration of data access hotspots. The heat diffusion path is the spatial trajectory of the temperature change of the lattice node and the migration operation.
[0031] The lattice heat diffusion coupling model is: ; where represents the Laplace operator of the local heat; represents the cooling coupling factor, which adjusts the influence intensity of the difference between the heat and the temperature threshold. According to the expert experience method, ranges from 0 to 1; represents the temperature threshold; It should be noted that the lattice heat diffusion coupling model draws on the classical heat conduction equation, reflecting the natural trend of heat diffusion from the high-temperature region to the low-temperature region; adding a cooling mechanism is equivalent to a heat dissipation mechanism, making the system tend to a low temperature; the spatial heat diffusion + threshold dissipation jointly determine the heat evolution, which is the physical basis for cold and hot layer determination.
[0032] The data migration trigger mechanism is: ; where Represents a node in a three-dimensional storage topology space At a time point The average heat; Represents the time derivative of the heat change of heterogeneous data sources; Represents the cooling delay judgment time window, which is the shortest time that must maintain a low temperature; Represents a time point To the time point The total temperature integral within; Indicates that when the trigger condition of the data migration trigger mechanism is met, a hierarchical migration operation is triggered; Is a time integral variable, representing any moment in the time interval To Among; It should be noted that the data migration trigger mechanism is a triple judgment mechanism, enhancing robustness: Condition 1 ensures that it is currently cold; Condition 2 ensures that the heat is continuously decreasing to avoid short-term fluctuations; Condition 3 ensures that the cooling lasts long enough to rule out accidental temperature drops; The integral term introduces time stability: elevating the judgment criterion from an instantaneous value to the average behavior over a period of time; Solves the following problems existing in the prior art: In the dynamic configurable data cleaning process based on big data, the number of heterogeneous data sources is huge, the structure is complex, and the data access patterns are variable. Traditional hot and cold data management methods mainly rely on static thresholds to divide the hot and cold states of data, and perform data migration and hierarchical management through manual or periodic triggering. Multi-source heterogeneous data shows highly complex and dynamic behavior patterns in actual cleaning applications. Traditional threshold-based hot and cold partitioning is difficult to accurately capture the critical changes in the hot and cold states of data, resulting in confusion between hot and cold data, and affecting the priority scheduling and resource allocation of cleaning tasks.
[0033] The intrinsic value of data (such as accuracy, integrity, etc.) changes continuously during the big data cleaning process. Traditional methods fail to fully consider the phase transition characteristics of value density changing over time, making it difficult to intelligently adjust the cleaning strategy and affecting the cleaning effect and efficiency. In large-scale data storage, there are obvious spatial local aggregation characteristics in the access hotspots of data, but traditional hot and cold data migration mechanisms ignore this spatial correlation, resulting in slow resource scheduling and migration response, and affecting the overall performance of the data cleaning system. Existing data migration strategies are mostly periodic or manually triggered, lacking dynamic adaptability and unable to respond to data state changes in real time, resulting in low utilization efficiency of storage resources and computing resources, and affecting the timeliness and accuracy of dynamic data cleaning.
[0034] The beneficial effects compared with the prior art are as follows: By comprehensively considering the access frequency, update frequency, and data value density changes, a comprehensive heat metric is constructed. Based on the local heat diffusion coupling model, it can capture the cold and hot critical points of multi-source heterogeneous data in the big data environment in real time, greatly improving the accuracy of cold and hot data partitioning, thereby optimizing the priority and strategy selection of data cleaning. It dynamically reflects the temporal changes and phase transition processes of data value density, enabling the data cleaning system to automatically adjust the cleaning rules and task scheduling according to the real-time changes in data value, enhancing the self-adaptability and intelligent level of the cleaning process.
[0035] By simulating the spatial diffusion process of heat in thermodynamics, it accurately reflects the migration trajectory of data access hotspots in the three-dimensional storage space, supports the dynamic hierarchical migration strategy based on the heat diffusion path, effectively avoids migration hysteresis, and improves the utilization rate of storage resources and the execution efficiency of data cleaning tasks. Using the preset temperature threshold and migration trigger mechanism, it can automatically sense the data temperature change and trigger the migration operation, realizing the real-time dynamic management of cold and hot data, ensuring the efficient execution of data cleaning tasks at the appropriate storage level, and enhancing the robustness and response speed of the system. The dynamic migration mechanism based on the heat diffusion path replaces the traditional manual or periodic migration strategy, realizes the automatic sensing and execution of data cold and hot states and migration operations, greatly reduces the manual maintenance cost, and improves the stability and intelligent level of the system.
[0036] For example, in a large-scale intelligent manufacturing platform, there are more than 4,000 industrial CNC devices deployed. These devices upload multi-source heterogeneous data through IoT nodes every minute, including: Structured: operating parameters, temperature, current, vibration values, etc. (recorded at a second-level frequency); Semi-structured: sensor alarm logs, system alarm status; Unstructured: infrared / thermal imaging video clips, device camera images, etc.; Interactive data: engineer operation records, inspection tasks, etc. The platform is based on the data lake architecture, with a daily new data volume of more than 5TB, among which: about 70% is high-value data with frequent short-term access (such as vibration trends, critical alarms, etc.); about 20% of the data enters the long-term silent state within a few hours after initial collection; the remaining 10% is periodically active (such as monthly device reconciliation records or compliance inspection callbacks).
[0037] The traditional cold and hot hierarchical cleaning mechanism determines that data is cold data and cleans or migrates it based on "last access time > 30 days", resulting in some high-value alarm data being wrongly cleaned before participating in warning aggregation; some structured sensor data, although inactive, is still a key data for AI model training and is wrongly labeled as "cold"; the computing resource scheduling is seriously lagging, and the cleaning frequency during peak periods does not match the sudden increase in data volume.
[0038] 3D lattice construction: construct a 3D lattice topology according to "data source type × time dimension × storage level"; each lattice node (e.g., temperature sensor data, May 15, 2025, hot data area) contains approximately 1000 records on average.
[0039] The heat value of a certain node in the temperature sensor data is calculated as follows: ; where the number of accesses to the data unit per unit time is 12; the update frequency of the data unit per unit time is 0; the change rate of the value density of the data unit per unit time is 0.01; then the current heat value ; if the average heat value around this node is 6.5, but its own heat value continuously drops to 3.0 and remains so for more than the cooling delay judgment time window (30 min), then cold data migration is triggered.
[0040] The method for automatically scheduling cleaning tasks includes: Preset a data cleaning operator library, and select operator nodes from the preset data cleaning operator library through drag-and-drop operations to construct a data cleaning flow chart; generate directed edges of the data cleaning flow chart through the connection relationships between operator nodes, and each directed edge is automatically marked with data dependency relationships; The preset data cleaning operator library includes input operators, transformation operators, filtering operators, aggregation operators, and output operators, and constructs a data cleaning flow chart formed by connecting corresponding operator nodes on a blank canvas in response to drag-and-drop operations; Formalize the data cleaning flow chart into a directed acyclic graph, and sort the operator nodes in the directed acyclic graph based on the topological sorting algorithm to generate an operator execution sequence that satisfies data dependency relationships, where operator nodes with no dependency relationships are marked as executable in parallel; generate executable task units according to the operator execution sequence, and automatically schedule the cleaning tasks to be executed in a heterogeneous processing architecture.
[0041] The method for obtaining the evolved rule instruction set includes: Preset data cleaning rules, initialize the data cleaning rule set as a rule population, and the cleaning rules are represented by encoding. The initial rule population is generated from the historical cleaning rule library; define a rule fitness function, and the fitness function comprehensively considers cleaning accuracy, cleaning coverage, and execution efficiency, and comprehensively evaluates the rule performance through weighted summation of weights; The fitness function is ; where represents cleaning accuracy; represents cleaning coverage; represents execution efficiency; Indicates the th cleaning rule; Index indicating the number of cleaning rules; It should be noted that the cleaning accuracy rate is used to measure the ability of data cleaning rules to identify and repair abnormal or incorrect data. The acquisition method is as follows: count the number of data correctly identified and repaired in the cleaning result and the total amount of abnormal data; divide the number of data correctly identified and repaired in the cleaning result by the total amount of abnormal data to obtain the cleaning accuracy rate; The cleaning coverage rate is used to measure the proportion of cleaning rules applicable to the current data stream, mainly reflecting the generality and breadth of the rules. The acquisition method is as follows: in a batch of input data, count the number of data records adapted and processed by a certain rule and the total number of this batch of data; divide the number of data records adapted and processed by a certain rule by the total number of this batch of data to obtain the cleaning coverage rate; The execution efficiency is used to measure the performance consumption of cleaning rules during actual operation. The acquisition methods include the following indicators: count the processing time of the rule on a certain batch of data and the total number of this batch of data ; The unit data processing time can be introduced as an efficiency indicator. Then the execution efficiency ; That is, the shorter the unit data cleaning time, the higher the efficiency and the better the fitness.
[0042] In each evolution cycle, select and retain the rule with the highest fitness according to the rule fitness function; perform crossover operations on the selected cleaning rules to generate new cleaning rules by combining the characteristics of the cleaning rules; apply mutation operations to the newly generated cleaning rules after crossover; Preset a fitness threshold, and replace the cleaning rules with fitness less than the preset fitness threshold with the newly evolved cleaning rules; after L evolution cycles, select the cleaning rules with fitness reaching the preset fitness threshold to form the final rule instruction set.
[0043] The method for calling the evolved rule instruction set to complete the preset data cleaning task includes: Construct a heterogeneous processing architecture composed of an FPGA acceleration unit, a Spark distributed processing engine, and a data scheduling manager; the FPGA acceleration unit is used to execute high-concurrency data cleaning operations based on the rule instruction set, the Spark distributed processing engine is used for batch scheduling of large-scale heterogeneous data and encapsulation and output of data cleaning results, and the data scheduling manager is used for scheduling of data batches; The data cleaning process is divided into three stages: preprocessing, core cleaning, and output verification. In the preprocessing stage, the heterogeneous data is batch-divided, pre-cached, and bound with cleaning rules through the Spark distributed processing engine; in the core cleaning stage, the heterogeneous data is cleaned through a customized VHDL kernel; in the output verification stage, the Spark distributed processing engine receives the cleaning results from the FPGA acceleration unit and integrates and outputs them, thus completing the preset data cleaning task; Based on the stream processing architecture, a VHDL cleaning pipeline containing N configurable logic modules is constructed. The evolved rule instruction set is converted into an instruction format recognizable by the FPGA, and according to the instruction content, it is mapped to the corresponding configurable logic module in the VHDL cleaning pipeline; the configurable logic module includes a field recognition module, a missing value repair module, an outlier replacement module, and a format standardization module, and each module can be enabled as needed to achieve dynamic modular cleaning; The heterogeneous data source is accessed through the Spark distributed processing engine. Based on the heat diffusion path, a data heat perception scheduling mechanism is constructed. The hot data is first pushed to the FPGA acceleration unit for cleaning processing through the data scheduling manager, and the cold data is cached waiting for scheduling; For each batch of heterogeneous data sources pushed, the corresponding rule instruction set identifier is bound, driving the FPGA acceleration unit to load the corresponding configurable logic module and execute the cleaning operation; the heterogeneous data sources that have completed cleaning are returned from the FPGA acceleration unit to the Spark distributed processing engine, formatted and encapsulated, and the cleaned heterogeneous data sources are output.
[0044] The method for auditing and verifying the heterogeneous data sources after cleaning processing includes: Construct a double-chain structure with an operation evidence chain and a quality certification chain. For each data cleaning operation process, the data cleaning operation timestamp, the cleaning rule ID called, the data cleaning flow chart, the unique identifier of the heterogeneous data source being processed, the operator, and the algorithm main body ID are recorded in real time; the Merkle Tree of the verifiable digest function is used to calculate the digest hash value of this data cleaning operation event to form an operation fingerprint; The operation fingerprint is written as a transaction into the operation evidence chain and linked with the hash of the previous block to form a blockchain structure. The operation evidence chain supports smart contracts. When the data cleaning operation is completed, the smart contract automatically triggers the auditing and verification of the cleaned heterogeneous data sources according to the operation fingerprint.
[0045] The preset temperature threshold is set by the staff. By collecting different temperatures, the average value of multiple temperatures is taken as the preset temperature threshold. Similarly, the preset cleaning credibility threshold, preset matching degree threshold, first threshold of preset data value density score, second threshold of data value density score, third threshold of data value density score, first threshold of preset recent access time, second threshold of recent access time, third threshold of recent access time, and preset fitness threshold are set.
[0046] In this embodiment, by comprehensively considering the access frequency, update frequency, and data value density change, a comprehensive heat metric is constructed. Based on the local heat diffusion coupling model, it can capture the cold and hot critical points of multi-source heterogeneous data in the big data environment in real time, greatly improving the accuracy of cold and hot data division, thereby optimizing the priority of data cleaning and the strategy selection. It dynamically reflects the temporal change and phase transition process of data value density, enabling the data cleaning system to automatically adjust the cleaning rules and task scheduling according to the real-time change of data value, enhancing the self-adaptability and intelligent level of the cleaning process.
[0047] By simulating the spatial diffusion process of heat in thermodynamics, it accurately reflects the migration trajectory of data access hotspots in the three-dimensional storage space, supports the dynamic hierarchical migration strategy based on the heat diffusion path, effectively avoids migration hysteresis, and improves the utilization rate of storage resources and the execution efficiency of data cleaning tasks. Using the preset temperature threshold and migration trigger mechanism, it can automatically sense the data temperature change and trigger the migration operation, realizing the real-time dynamic management of cold and hot data, ensuring the efficient execution of data cleaning tasks at the appropriate storage level, and enhancing the robustness and response speed of the system. The dynamic migration mechanism based on the heat diffusion path replaces the traditional manual or periodic migration strategy, realizes the automatic perception and execution of data cold and hot states and migration operations, greatly reduces the manual maintenance cost, and improves the stability and intelligent level of the system.
[0048] Embodiment 2 Please refer to Figure 2 As shown in the figure, a dynamic configurable data cleaning system based on big data in this embodiment includes: An adaptive detection module, which is used to collect heterogeneous data sources, access heterogeneous data sources in combination with the data source adaptation interface; use the simulated antigen recognition and memory mechanism to identify abnormal patterns in the heterogeneous data sources and encode them as antigen fingerprints; construct a data value density scoring function based on the antigen fingerprints to obtain the data value density score; An intelligent hierarchical storage module, which constructs a three-dimensional storage topology space according to the data value density score, adopts a lattice migration algorithm, automatically identifies the cold and hot data critical points, triggers hierarchical migration when the heterogeneous data source cools down to the preset temperature threshold, introduces the thermodynamic diffusion mechanism, and simulates the migration process of the heterogeneous data source in the three-dimensional storage topology space as a heat diffusion path; The data cleaning orchestration module constructs a data cleaning flow chart in a drag-and-drop manner, determines the execution order of cleaning operators in the data cleaning flow chart based on topological sorting, and automatically schedules cleaning tasks; The ecological rule evolution module constructs a Darwinian evolution system for cleaning rules, evolves the cleaning rules periodically, defines a fitness function for cleaning rules, simulates the natural selection mechanism to dynamically optimize the cleaning strategy, and outputs an evolved rule instruction set; The hardware-accelerated cleaning module constructs a heterogeneous processing architecture, customizes a VHDL kernel to clean heterogeneous data sources; collaborates based on the heat diffusion path and the Spark batch reading mechanism, and calls the evolved rule instruction set to complete the preset data cleaning tasks; The trusted cleaning and archiving module constructs a double-chain structure with an operation archiving chain and a quality certification chain, audits and verifies the heterogeneous data sources after cleaning; uses a verifiable digest function to record the cleaning credibility of the cleaned heterogeneous data sources; triggers a re-cleaning mechanism when the cleaning credibility is less than the preset cleaning credibility threshold, and updates the cleaning rules.
[0049] Since the electronic device introduced in this embodiment is the electronic device used in implementing a big data-based dynamically configurable data cleaning method and its system in the embodiments of the present application, based on the big data-based dynamically configurable data cleaning method and its system introduced in the embodiments of the present application, those skilled in the art can understand the specific implementation manners and various variations of the electronic device in this embodiment. Therefore, the specific implementation of how this electronic device implements the method in the embodiments of the present application will not be described in detail here. As long as those skilled in the art implement the electronic device used in a big data-based dynamically configurable data cleaning method and its system in the embodiments of the present application, it falls within the protection scope of the present application.
[0050] The above formulas are all dimensionless and take their numerical values for calculation. The formulas are obtained by collecting a large amount of data for software simulation to get a formula closest to the real situation. The selection of preset parameters and thresholds in the formulas is set by those skilled in the art according to the actual situation.
[0051] The above are only the preferred implementation manners of the present invention. The protection scope of the present invention is not limited to the above embodiments. Any technical solutions falling within the idea of the present invention belong to the protection scope of the present invention. It should be noted that for ordinary technical users in the technical field, several improvements and refinements made without departing from the principle of the present invention should also be regarded as within the protection scope of the present invention.
Claims
1. A dynamic configurable data cleaning method based on big data, characterized in that Including: S1. Collect heterogeneous data sources and access the heterogeneous data sources in combination with the data source adaptation interface; Utilize the simulated antigen recognition and memory mechanism to identify the abnormal patterns in the heterogeneous data sources and encode them as antigen fingerprints; construct a data value density scoring function based on the antigen fingerprints to obtain the data value density score; S2. Construct a three-dimensional storage topology space according to the data value density score, adopt the lattice migration algorithm to automatically identify the critical points of hot and cold data, trigger hierarchical migration when the heterogeneous data source cools down to the preset temperature threshold, introduce the thermodynamic diffusion mechanism, and simulate the migration process of the heterogeneous data source in the three-dimensional storage topology space as a heat diffusion path; S3. Construct a data cleaning flow chart by dragging, determine the execution order of the cleaning operators in the data cleaning flow chart based on topological sorting, and automatically schedule the cleaning tasks; S4. Construct a Darwinian evolution system for cleaning rules, perform periodic evolution on the cleaning rules, define a cleaning rule fitness function, simulate the natural selection mechanism to dynamically optimize the cleaning strategy, and output the evolved rule instruction set; S5. Construct a heterogeneous processing architecture, customize a VHDL kernel to perform cleaning processing on the heterogeneous data source; cooperate with the heat diffusion path and the Spark batch reading mechanism, and call the evolved rule instruction set to complete the preset data cleaning task; S6. Construct a double-chain structure with an operation evidence chain and a quality certification chain to perform audit verification on the heterogeneous data source after cleaning processing; Utilize a verifiable digest function to record the cleaning credibility of the cleaned heterogeneous data source; trigger a re-cleaning mechanism when the cleaning credibility is less than the preset cleaning credibility threshold, and update the cleaning rules.
2. The dynamic configurable data cleaning method based on big data according to claim 1, characterized in that, The method for identifying and encoding the abnormal patterns in the heterogeneous data sources as antigen fingerprints includes: The heterogeneous data sources include structured data, semi-structured data, unstructured data, streaming data, and system operation and maintenance data; perform standard deviation normalization processing on the heterogeneous data sources to convert them into a standardized data format; Construct a simulated antibody library, and define each antibody in the simulated antibody library as an adaptable and mutable antibody template; the antibody template is any one of a template based on the K-nearest neighbor model, a template based on the clustering center vector, or a template based on a small sample neural network; Utilize a similarity detection mechanism to calculate the matching degree between the heterogeneous data source and the antibody template; the similarity detection mechanism is any one of cosine similarity, Euclidean distance, or kernel function similarity; if the matching degree is less than the preset matching degree threshold, it is identified as an abnormal pattern and marked as a simulated antigen; Perform fingerprint encoding processing on the simulated antigen, and the fingerprint encoding processing includes feature dimensionality reduction, fingerprint hash mapping, and feature vector normalization processing to generate an antigen fingerprint containing data identification, abnormal category, timestamp, and feature hash value.
3. A dynamic configurable data cleaning method based on big data according to claim 2, characterized in that The method for obtaining the data value density score includes: Perform multi-dimensional information extraction on the antigen fingerprints to construct a data feature vector, and the multi-dimensions include antigen occurrence frequency, abnormal level, fingerprint similarity distribution, historical trigger records, and context information entropy; A data value density scoring function is constructed based on the data feature vector. The data value density scoring function is a nonlinear weighted function. Each antigen fingerprint is input into the corresponding data value density scoring function to calculate the corresponding data value density score.
4. A dynamic configurable data cleaning method based on big data according to claim 3, characterized in that The method for constructing the three-dimensional storage topological space includes: Design a three-dimensional coordinate system for the three-dimensional storage topological space, combine X, Y, and Z as the three-dimensional coordinates of the three-dimensional storage topological space, organize all three-dimensional coordinates into integer grids, and form a logical three-dimensional storage topological space; each dimension is defined as the X-axis, Y-axis, and Z-axis; the X-axis represents the value density dimension, the Y-axis represents the data structure dimension, and the Z-axis represents the time life cycle dimension; Establish grading and coding rules for each dimension. For the X-axis, normalize all data value density scores and linearly map them to the [0,1] scoring interval. Define the stratification levels according to the scoring interval, and preset the first threshold of data value density score, the second threshold of data value density score, and the third threshold of data value density score. If the data value density score is greater than or equal to the second data value density score threshold, and less than or equal to the third data value density score threshold, it is classified as a high value layer; if the data value density score is greater than or equal to the first data value density score threshold, and less than the second data value density score threshold, it is classified as a medium value layer; if it is less than the first data value density score threshold, it is classified as a low value layer; For the Y axis, set static mapping codes for the data structure types included in the data structure dimension, which include structured data, semi-structured data, and unstructured data; for structured data, let Y=0; for semi-structured data, let Y=1; for unstructured data, let Y=2; the data structure dimension is an extensible support subclass, and Y=3 indicates mixed structure data; For the Z axis, the time life cycle stage is defined according to the most recent access time, and the mapping rules are set; the first threshold of the most recent access time, the second threshold of the most recent access time, and the third threshold of the most recent access time are preset; when the most recent access time is less than the first threshold of the most recent access time, Z=0 is set; when the most recent access time is greater than or equal to the first threshold of the most recent access time and less than the second threshold of the most recent access time, Z=1 is set; when the most recent access time is greater than the second threshold of the most recent access time, Z=2 is set.
5. A dynamic configurable data cleaning method based on big data according to claim 4, characterized in that The method for obtaining the heat diffusion path includes: In a three-dimensional storage topological space, each coordinate point is regarded as a lattice node, and the node contains m data units; for each data unit , its current heat value is defined as ; regarding the three-dimensional coordinate system as a spatial lattice grid, the average heat of each lattice node is defined as ; Considering the spatial correlation of heat diffusion, a local heat diffusion coefficient is introduced , a lattice heat diffusion coupling model is established, and a temperature threshold is preset. When the heterogeneous data source cools down to the preset temperature threshold, the hierarchical migration determination is triggered; the heterogeneous data source with a temperature less than or equal to the preset temperature threshold is marked as cold data, and the heterogeneous data source with a temperature greater than the preset temperature threshold is marked as hot data; Define a data migration trigger mechanism, and perform hierarchical migration operations when the trigger conditions of the data migration trigger mechanism are met; migrate heterogeneous data sources from high-temperature lattice nodes to low-temperature lattice nodes through the heat diffusion path, simulate the gradual release of heat and the migration process of data access hotspots, and the heat diffusion path is the spatial trajectory of lattice node temperature changes and migration operations.
6. A dynamic configurable data cleaning method based on big data according to claim 5, characterized in that, The method for automatically scheduling cleaning tasks comprises: Preset data cleaning operator library, select operator nodes from the preset data cleaning operator library by dragging and dropping to build a data cleaning flowchart; generate directed edges of the data cleaning flowchart through the connection relationship between operator nodes, and automatically annotate data dependency for each directed edge; Formalize the data cleaning flow chart into a directed acyclic graph, and sort the operator nodes in the directed acyclic graph based on the topological sorting algorithm to generate an operator execution sequence that satisfies the data dependency relationship, where the operator nodes without dependency relationships are marked as executable in parallel; generate executable task units according to the operator execution sequence, and automatically schedule the cleaning tasks to the heterogeneous processing architecture for execution.
7. A dynamic configurable data cleaning method based on big data according to claim 6, characterized in that The method for obtaining the evolved rule instruction set includes: Preset data cleaning rules, initialize the data cleaning rule set as a rule population, where the cleaning rules are represented by encoding, and the initial rule population is generated from the historical cleaning rule library; define a rule fitness function, which comprehensively considers the cleaning accuracy, cleaning coverage rate, and execution efficiency, and comprehensively evaluates the rule performance through weighted summation of weights; In each evolution cycle, retain the rule with the highest fitness according to the rule fitness function; perform a crossover operation on the selected cleaning rules to generate new cleaning rules by combining the characteristics of the cleaning rules; apply a mutation operation to the newly generated cleaning rules after crossover; Preset a fitness threshold, and replace the cleaning rules with fitness less than the preset fitness threshold with the newly evolved cleaning rules; after L evolution cycles, select the cleaning rules with fitness reaching the preset fitness threshold to form the final rule instruction set.
8. A dynamic configurable data cleaning method based on big data according to claim 7, characterized in that The method for calling the evolved rule instruction set to complete the preset data cleaning task includes: Construct a heterogeneous processing architecture composed of an FPGA acceleration unit, a Spark distributed processing engine, and a data scheduling manager; divide the data cleaning process into three stages: preprocessing, core cleaning, and output verification. In the preprocessing stage, the Spark distributed processing engine batches, pre-caches, and binds cleaning rules to heterogeneous data; In the core cleaning stage, perform data cleaning operations on heterogeneous data through a customized VHDL kernel; in the output verification stage, the Spark distributed processing engine receives the cleaning results of the FPGA acceleration unit and integrates and outputs them, thereby completing the preset data cleaning task; Build a VHDL cleaning pipeline containing N configurable logic modules based on the stream processing architecture, convert the evolved rule instruction set into an instruction format recognizable by the FPGA, and map it to the corresponding configurable logic module in the VHDL cleaning pipeline according to the instruction content; the configurable logic modules include a field recognition module, a missing value repair module, an outlier replacement module, and a format standardization module, and each module can be enabled as needed to achieve dynamic modular cleaning; Access heterogeneous data sources through the Spark distributed processing engine, build a data heat perception scheduling mechanism based on the heat diffusion path, and push hot data to the FPGA acceleration unit for cleaning processing through the data scheduling manager, and cache cold data waiting for scheduling; Bind the corresponding rule instruction set identifier to each batch of pushed heterogeneous data sources, drive the FPGA acceleration unit to load the corresponding configurable logic module and execute the cleaning operation; return the heterogeneous data sources that have completed cleaning from the FPGA acceleration unit to the Spark distributed processing engine for formatting and encapsulation, and output the cleaned heterogeneous data sources.
9. A dynamic configurable data cleaning method based on big data according to claim 8, characterized in that The method for auditing and verifying the heterogeneous data source after cleaning processing includes: Construct a double-chain structure with an operation evidence chain and a quality certification chain. For each data cleaning operation process, record the data cleaning operation timestamp, the cleaning rule ID called, the data cleaning flow chart, the unique identifier of the heterogeneous data source being processed, the operator, and the algorithm subject ID in real time; use the verifiable digest function Merkle Tree to calculate the digest hash value of the data cleaning operation event to form an operation fingerprint. Write the operation fingerprint as a transaction into the operation evidence chain and link it with the hash of the previous block to form a blockchain structure. The operation evidence chain supports smart contracts. When the data cleaning operation is completed, the smart contract automatically triggers the auditing and verification of the heterogeneous data source after cleaning processing according to the operation fingerprint.
10. A dynamic configurable data cleaning system based on big data, which is used to implement a dynamic configurable data cleaning method based on big data according to any one of claims 1 to 9, and is characterized in that, It includes: An adaptive detection module for collecting heterogeneous data sources and accessing heterogeneous data sources in combination with a data source adaptation interface. Utilize the simulated antigen recognition and memory mechanism to identify abnormal patterns in the heterogeneous data source and encode them as antigen fingerprints; construct a data value density scoring function based on the antigen fingerprints to obtain a data value density score. An intelligent hierarchical storage module constructs a three-dimensional storage topology space according to the data value density score, adopts a lattice migration algorithm, automatically identifies the critical point of hot and cold data, triggers hierarchical migration when the heterogeneous data source cools down to a preset temperature threshold, and introduces a thermodynamic diffusion mechanism to simulate the migration process of the heterogeneous data source in the three-dimensional storage topology space as a heat diffusion path. A data cleaning orchestration module constructs a data cleaning flow chart by dragging and dropping, determines the execution order of cleaning operators in the data cleaning flow chart based on topological sorting, and automatically schedules cleaning tasks. An ecological rule evolution module constructs a Darwinian evolution system for cleaning rules, evolves the cleaning rules periodically, defines a cleaning rule fitness function, simulates the natural selection mechanism to dynamically optimize the cleaning strategy, and outputs an evolved rule instruction set. A hardware-accelerated cleaning module constructs a heterogeneous processing architecture, customizes a VHDL kernel to clean the heterogeneous data source; coordinates with the heat diffusion path and the Spark batch reading mechanism, and calls the evolved rule instruction set to complete a preset data cleaning task. A trusted cleaning evidence storage module constructs a double-chain structure with an operation evidence chain and a quality certification chain to audit and verify the heterogeneous data source after cleaning processing. Use a verifiable digest function to record the cleaning credibility of the heterogeneous data source after cleaning; trigger a re-cleaning mechanism when the cleaning credibility is less than a preset cleaning credibility threshold and update the cleaning rules.
Citation Information
Patent Citations
Multi-source heterogeneous ecological environment big data processing method and system based on data lake
CN111459908A
Optical fiber data storage management system and method based on big data
CN120085812A
Digital twin system for power grid
WO2025086085A1
Cited By
Cold and hot data exchange method and system based on optical storage and storage medium
CN120428929A
Big data cleaning and process arrangement system, method and equipment and medium
CN120469992A
Multi-source data integration method based on data space
CN120578657A