Full link data cleaning and intelligent business insight integrated learning system
By integrating a full-chain data cleaning and intelligent business insight learning system, the problems of missing evidence chains and unstable insight results during the data cleaning process have been solved, thereby improving the authenticity of data and the stability of insights.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHAOQIANYUE TECHNOLOGY SERVICE HEBEI CO LTD
- Filing Date
- 2026-06-16
- Publication Date
- 2026-07-21
AI Technical Summary
The data cleaning process in existing technologies lacks end-to-end tracking, which leads to the inability of derived data to inherit the root evidence of the original data, resulting in feedback and self-verification phenomena, which damages the authenticity of the data. Furthermore, the side effects of cleaning are difficult to distinguish from real anomalies, and the insight results are unstable.
The system integrates end-to-end data cleaning and intelligent business insight learning. It encapsulates field values into atomic data units through multi-source access and atomic processing modules, ensures the integrity of the data evidence chain through evidence conservation and backfeeding blocking modules, quantifies and records cleaning actions through cleaning execution and post-effect marking modules, and generates mutually exclusive repair worlds and verifies the consistency of insights through repair world generation and cross-validation modules.
It enables full-chain tracing of data evidence, blocks the feedback and self-verification path, improves the reliability of data repair, and distinguishes between cleansing-induced anomalies and real business anomalies, thereby improving the stability and credibility of business insight results.
Smart Images

Figure CN122432152A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of enterprise data governance and intelligent analysis technology, and more specifically, to a full-chain data cleaning and intelligent business insight integrated learning system. Background Technology
[0002] This invention relates to the field of enterprise data governance and intelligent analysis technology, specifically to integrated processing technology for data cleaning and business insights. With the advancement of enterprise digital transformation, the cleaning of multi-source heterogeneous data and intelligent business insights have become core components of data-driven decision-making.
[0003] In existing technologies, data cleaning processes typically lack end-to-end tracking of the data evidence chain. Derived data cannot inherit the root evidence from the original data, making it prone to feedback loops and self-verification—where insights from the system's output are used to correct the original data, compromising data authenticity. Furthermore, existing systems do not quantify, record, or disseminate the aftereffects of cleaning actions, failing to distinguish between genuine anomalies in business insights and cleaning side effects, leading to misjudgments. Moreover, for scenarios with multiple uncertain repair hypotheses, existing systems often directly adopt a single repair result without verifying the consistency of insights under different repair hypotheses, resulting in insufficient stability of the insights.
[0004] To address this, the present invention proposes an integrated learning system for end-to-end data cleaning and intelligent business insights. Summary of the Invention
[0005] To address the shortcomings of existing technologies, the purpose of this invention is to provide an integrated learning system for end-to-end data cleaning and intelligent business insights.
[0006] To achieve the above objectives, the present invention provides the following technical solution: The end-to-end data cleaning and intelligent business insight integrated learning system includes a multi-source access and atomic processing module, an evidence conservation and backfeeding prevention module, a cleaning execution and post-effect marking module, a repair world generation and cross-validation module, and an insight output and feedback control module. The multi-source access and atomization processing module is used to receive multi-source data and encapsulate field values into atomic data units that record the original value, current value, source, time, root evidence set, cleaning action sequence, aftereffect marker set, and repair world identifier. The evidence conservation and backfeed blocking module is used to enable derived data to inherit the root evidence set of the parent data unit, and to verify the evidence of candidate repairs based on independent evidence support relationship and feedback loop overlap relationship. The cleaning execution and post-processing marking module is used to execute cleaning actions, record the post-processing markings corresponding to the cleaning actions, and propagate them along the data lineage. The repair world generation and cross-validation module is used to generate mutually exclusive repair worlds based on candidate repair hypotheses that have passed constraint checks, and run insight models in each repair world to determine model consistency and repair world consistency. The insight output and feedback control module is used to determine the insight state based on the root evidence state, the aftereffect matching result, the model consistency, and the repair world consistency, and to control the conditions for business feedback to enter repair verification or model update.
[0007] Furthermore, the multi-source access and atomic processing module is used to receive data from business databases, interface messages, logs, documents, message streams, and model recognition results; for structured data, it splits the data by field granularity; for semi-structured data, it parses field names, node paths, field values, and time information; for unstructured data, it obtains candidate field values through entity recognition, field extraction, layout parsing, or speech transcription, and adds parsing confidence information to the candidate field values.
[0008] Furthermore, the multi-source access and atomization processing module is also used to determine whether the atomic data unit originates from independently acquired facts; when the atomic data unit is generated by an independent observation, independent input, independent sensing, independent interface response, or independent manual verification, and its generation process does not depend on the cleaning results, insight results, or strategy suggestions previously output by the system, a root evidence identifier is generated based on the source channel, acquisition time, acquisition sequence number, original payload summary, and source signature; when independence cannot be confirmed, the corresponding data is marked as derived input.
[0009] Furthermore, the evidence conservation and backfeeding blocking module is used to determine the set of parent data units based on the input fields of the transformation operator, the aggregation window, the entity merging basis, the model input features, and the field lineage when the atomic data units are obtained by calculation, transformation, aggregation, completion, merging, or model inference, and to use the set of root evidence carried by the parent data units as the source of the derived root evidence; for candidate repair hypotheses, the independent evidence support is determined based on the deduplication result of the root evidence carried by the data units that support the candidate repair hypothesis, and the candidate repair hypothesis is not allowed to directly cover the master data when traceable evidence is lacking.
[0010] Furthermore, the evidence conservation and backfeeding blocking module is also used to determine the root evidence set of the target data to be repaired and the root evidence set on which the feedback data, insight results, or backfeed data supporting candidate repair depend, and to determine the feedback loop overlap degree accordingly; when the feedback loop overlap degree meets the backfeeding blocking condition, the corresponding feedback is not allowed to increase the repair confidence or be used as a training label to update the cleaning model; for candidate repairs that pass the evidence verification, a repair imprint is generated including the candidate repair hypothesis identifier, root evidence summary, cleaning action type, action parameter summary, rule version, model version, action field, action time segment, and affected lineage subgraph.
[0011] Furthermore, the cleaning execution and post-processing marking module is used to receive atomic data units to be processed and cleaning actions. The cleaning actions include field normalization, unit conversion, outlier handling, missing value completion, duplication resolution, entity merging, time alignment, encoding conversion, and feature enhancement. When each cleaning action is executed, the input field, output field, triggering condition, action type, action parameters, affected fields, and affected data range are recorded. For cleaning actions involving overwriting original values, shifting time, merging entities, or replacing anomalies, candidate repair hypotheses are first generated and then verified in the repair world.
[0012] Furthermore, the aftereffect markers include value change intensity, distribution compression intensity, time rearrangement intensity, substitution component intensity, and scope intensity; the cleaning execution and aftereffect marker module is used to pass aftereffect markers to output features and insight inputs based on the contribution weight of atomic data units to output data, the sensitivity of output fields to input fields, and the degree of propagation decay of cleaning actions when atomic data units participate in aggregation, connection, model input, or feature derivation; for fields that affect primary key matching, entity merging, event sorting, or window partitioning, aftereffect markers are propagated according to structural influence.
[0013] Furthermore, the cleaning execution and aftereffect marking module is also used to generate an abnormal pattern vector of the insight target before the business insight is output, and extract the synthetic aftereffect vector related to the cleaning action from the lineage path of the insight target; when the matching relationship between the abnormal pattern vector and the synthetic aftereffect vector satisfies the aftereffect matching condition, and the corresponding cleaning action is located in the lineage path of the insight target, the insight target is marked as a cleaning-induced candidate insight.
[0014] Furthermore, the repair world generation and cross-validation module is used to generate candidate repair hypotheses for uncertain anomalies, and to perform constraint checks on the candidate repair hypotheses on field type, value boundaries, event order, root evidence status, source consistency, primary key relationship, and affected lineage subgraph; when multiple candidate repair hypotheses act on the same data unit, the same entity relationship, the same event time, or the same aggregation caliber and cannot be true at the same time, different repair worlds are generated; the repair world adopts a difference layer storage method, and the difference layer stores the data units that have changed relative to the main world, cleaning actions, aftereffect markers, and evidence status.
[0015] Furthermore, the repair world generation and cross-validation module is also used to replay the affected lineage subgraphs in each repair world and run multiple insight models to obtain insight values with uniform dimensions. It determines model consistency based on the differences in the outputs of multiple insight models within the same repair world and determines repair world consistency based on the differences in insight outputs between different repair worlds. The insight output and feedback control module is used to determine the comprehensive credibility based on repair world consistency, model consistency, independent evidence support, feedback loop overlap, and aftereffect matching, and classifies the insight state into stable insight, repair sensitive insight, cleansing and inducing candidate insight, or evidence blocking state. After receiving business feedback, the business feedback is first converted into feedback data units and its root evidence set is determined. When the business feedback comes from independently collected facts and does not trigger the backfeeding blocking condition, it is used to participate in the candidate repair hypothesis verification; otherwise, it is only processed as derived input or audit record.
[0016] Compared with the prior art, the present invention has the following beneficial effects: 1. To address the issues of missing derived data evidence and compromised data authenticity due to feedback loop self-verification in existing technologies, this invention employs a multi-source access and atomization processing module to encapsulate field values into atomic data units containing a root evidence set. Through an evidence conservation and feedback loop blocking module, derived data inherits the root evidence set of its parent data unit. Evidence verification of candidate repairs is performed based on independent evidence support and feedback loop overlap. This technical feature enables end-to-end tracking of the data evidence chain, blocks feedback loop self-verification paths, and helps improve the reliability of data repair. 2. To address the issues in existing technologies where it's difficult to distinguish between the side effects of data cleaning and genuine business anomalies, and where uncertain repairs lead to unstable insight results, this invention quantifies and records the aftereffects of cleaning actions through a cleaning execution and aftereffect labeling module, propagating them along data lineage to detect the matching relationship between aftereffects and insight anomalies. Furthermore, a repair world generation and cross-validation module generates mutually exclusive repair worlds and runs multiple insight models to determine model consistency and repair world consistency. This technical feature can distinguish between cleaning-induced anomalies and genuine business anomalies, helping to improve the stability and reliability of business insight results. Attached Figure Description
[0017] Figure 1 A block diagram of a learning system integrating end-to-end data cleaning and intelligent business insights; Figure 2 This is a flowchart illustrating the implementation of the cleaning execution and post-marking module of the present invention. Figure 3 This is a flowchart illustrating the implementation of the world generation and cross-validation module of this invention. Detailed Implementation
[0018] Example, refer to Figure 1 The full-link data cleaning and intelligent business insight integrated learning system in this embodiment includes a multi-source access and atomization processing module, an evidence conservation and backfeeding blocking module, a cleaning execution and post-effect marking module, a repair world generation and cross-validation module, and an insight output and feedback control module. Mod1, Multi-source Access and Atomization Processing Module.
[0019] Data from different sources is transformed into unified atomic data units, and the original value, current value, source, time, root evidence (an indivisible original evidence unit generated from independently collected facts, which is the ultimate basis for the authenticity and independence of the data), cleaning action, aftereffect markers, and repair world identifiers are recorded at the system entry point.
[0020] S11. Receive data from business databases, interface messages, logs, documents, message streams, or model recognition results; for structured data, split it by field granularity; for semi-structured data, parse field names, node paths, field values, and time information; for unstructured data, first obtain candidate field values through entity recognition, field extraction, layout parsing, or speech-to-text, and then attach parsing confidence information; the above parsing process can be implemented using existing text recognition, entity recognition, or field mapping technologies, as long as the field values, source, event time, and original payload summary can be obtained; S12. Encapsulate each field value into an atomic data unit: ; in, Represents atomic data units; Indicates the atomic data unit identifier; Represents the original value that has not been covered; Indicates the value currently being calculated; Indicates the field identifier; Indicates the source system, device, or interface channel; Indicates the time when the business event occurred; Indicates the system receiving time; This represents the set of root evidence carried by the atomic data unit; This indicates the sequence of cleaning actions that have been applied to this atomic data unit; This represents the set of aftereffect tags carried by the atomic data unit; This indicates the repair world identifier to which the atomic data unit belongs; for master data that has not yet entered the repair world, Take the overworld identifier; for data that has not yet been cleaned, and same, and It can be an empty set; S13. Determine whether the atomic data unit originates from an independently acquired fact (the data is generated by a single independent observation, independent entry, independent sensing, independent interface response, or independent manual verification, and its generation process does not depend on the cleaning results, insight results, or strategy recommendations previously output by this system; the judgment criterion is that the triggering conditions, acquisition equipment, processing flow, and storage location of the data generation have no direct causal relationship with the output results of this system); when this condition is met, the system generates a root evidence identifier based on the source channel, acquisition time, acquisition sequence number, original payload summary, and source signature. and order ; Root evidence identifier Generated using hash combination method: ; in, Indicates the source channel identifier; Indicates the time of data collection; Indicates the collection sequence number; This represents a summary of the original payload; Indicates the signature from the source; This represents a string concatenation operation; This represents a one-way hash function; The original payload digest is obtained through one-way digest functions such as SHA series and MD series. In this embodiment, the SHA256 algorithm is used. The digest object includes at least the original field value, the source payload fragment, the source channel identifier, and the received time fragment. For return data whose independence cannot be confirmed, the system can mark it as a derived input and not generate new root evidence for it.
[0021] Mod2, Evidence Conservation and Recharge Blocking Module.
[0022] Constrained derived data (data obtained by calculation, transformation, aggregation, completion, merging, or model inference from one or more original data or other derived data) can only inherit existing evidence and cannot increase the number of independent evidence due to copying, aggregation, completion, report return, or model inference.
[0023] S21, If atomic data unit If a data unit is calculated, transformed, aggregated, completed, merged, or inferred from one or more parent data units, then the Evidence Conservation and Reinjection Blocking Module determines its set of parent data units. The parent data unit is determined by the input fields of the transformation operator, the aggregation window, the entity merging criteria, the model input features, and the field lineage. The root evidence set of the output data unit is inherited according to the following formula: ; in, Indicates the output data unit; This represents the set of parent data units that the output data unit depends on. Indicates the parent data unit; This represents the set of root evidence carried by the parent data unit; This represents the union of sets; this expression indicates that derivation processing can only pass on existing root evidence and cannot generate new root evidence due to data copying or format changes. S22, Regarding candidate repair hypotheses Determine the set of data units that support this hypothesis. Support relationships can come from cross-source consistency, field constraints, invariant residuals, sensor measurements, external receipts, manual verification, or model interpretation results; the system will The root evidence carried by each data unit is placed into a multiset. The same piece of evidence can appear in different supporting data. The independent evidence support for the candidate repair hypothesis is as follows: (The last part, "repeated occurrences," appears to be a typo and can be omitted.) ; in, Indicates candidate repair hypothesis The degree of independent evidence supporting the claim; This indicates a multiple set of root evidence supporting the data. Indicates to The set of root evidence obtained after deduplication; Indicates the number of times the root evidence appears in the multiset; Indicates the number of root evidences after deduplication; For dimensionless values, the normalized range is ;like An empty value indicates that the candidate repair lacks traceable evidence, and the system does not allow the repair to directly overwrite the master data. Independent evidence support threshold The range of values is In this embodiment, repair actions are divided into three levels: low risk, medium risk, and high risk. Low-risk repair actions include field format conversion and unit conversion. It can be set to a lower value; medium-risk repair actions include outlier replacement and missing value completion. It can be set to a medium value; high-risk repair actions such as entity merging and time reordering, It can be set to a higher value; Adjustments can be made based on historical manual review results. When the manual review pass rate for a certain type of repair action is lower than the preset standard, the corresponding pass rate for that type of repair action can be increased. ; S23. Detect whether the candidate repair has self-verification through re-feedback; for candidate repair hypotheses... The system determines the root evidence set of the target data to be repaired. And identify the set of root evidence upon which the feedback data, insights, or reflow data used to support the repair depend. Calculate the overlap of the feedback loop: ; in, Indicates candidate repair hypothesis The degree of overlap of feedback loops; This represents the set of root evidence for the target data or segment to be repaired. This represents the set of root evidence upon which the feedback, insights, or reflow data used to support the fix depends; Represents the intersection of sets; For dimensionless values, the normalized range is ;like If the result is empty, the system will mark the candidate repair as evidence that cannot be verified. Feedback loop threshold The range of values is In this embodiment, Set to a medium value; when the feedback loop overlap is... Exceed This indicates that the feedback supporting the repair is highly dependent on the data to be repaired itself or its derived results. The system does not allow the feedback to directly increase the repair confidence, nor does it allow it to be used as a training label to directly update the cleaning model. Adjustments can be made based on the results of business replays. When the system experiences numerous errors due to repeated self-verification during re-implementation, the impact can be reduced. ; S24. Generate repair imprints for candidate repairs that pass evidence verification. Repair imprints include at least the candidate repair hypothesis identifier, root evidence summary, cleansing action type, action parameter summary, rule version, model version, action field, action time segment, and affected lineage subgraph. When any subsequent data flows back into the system, the system compares the repair imprint it carries with the insight imprint on which the repair to be performed depends. If the two have the same or highly overlapping root evidence summaries and there is a backflow path of the same repair action or the same insight result, the backflow data is only treated as derived data and not as new independent evidence.
[0024] Mod3, the cleaning execution and post-processing marking module.
[0025] like Figure 2 As shown, while cleaning the data, the manual traces that may be generated by the cleaning process are recorded, enabling downstream insights to distinguish between real business anomalies and cleaning side effects (aftereffect markers, which are marker vectors that record the manual traces generated by the cleaning action on the data).
[0026] S31, Receive atomic data units to be processed and perform cleaning actions. Cleaning actions may include field normalization, unit conversion, outlier handling, missing value completion, duplicate elimination, entity merging, time alignment, encoding conversion, and feature enhancement. Each cleaning action records the input fields, output fields, triggering conditions, action type, action parameters, affected fields, and affected data range during execution. For routine field format conversions, the current value can be directly generated. For high-risk actions involving overwriting original values, shifting time, merging entities, or replacing anomalies, the system first generates candidate repair hypotheses and then enters the repair world for verification. S32. After each cleaning action is executed, the cleaning execution and post-effect marking module generates a post-effect marking vector: ; in, Indicates the cleaning action Acting on atomic data units The resulting post-generic label vector; each component is a dimensionless numerical value, with a normalized range of [value missing]. Normalization can be performed based on the historical fluctuation range of the field, the distribution of data in the same batch, the conversion scale of the field unit, or the upper limit of the action configuration. The specific calculation methods for each component are as follows: Value change intensity : For numeric fields: ; in, Representation field The historical maximum value range is determined by the difference between the maximum and minimum values of the field in the historical data. For text fields: ; in, Indicates edit distance, Representation field The longest historical text length; For categorical fields: ; in, Indicates the level of the category value. Representation field The maximum number of category ranks; Distributed compressive strength : ; in, Indicates the cleaning action Before execution, atomic data units The variance of the data within the local window; Indicates the cleaning action After execution, the variance of the data within the local window and the size of the local window are determined based on field characteristics and business requirements. Time rearrangement intensity : ; in, Indicates the time of the business event prior to the cleaning; Indicates the time of occurrence of the business event after the cleanup; This represents the largest time window size in the business system, which is determined based on the business scenario. strength of substitute ingredients : ; in, This represents the sum of weights corresponding to the components in the current value formed by completion, inference, merging, or external correction; This represents the sum of the weights of all components of the current value; for values derived entirely from the original data, =0, =0; for values obtained entirely through completion or deduction, equal , =1; Range of action intensity : ; in, Indicates the action of being cleaned The total number of data units, fields, and feature nodes affected; This indicates the total number of data units, fields, and feature nodes in the current lineage subgraph; S33. The after-effect label propagates along the data lineage. If an atomic data unit participates in aggregation, connection, model input, or feature derivation, the cleaning execution and after-effect labeling module will pass the corresponding after-effect label to the output feature and insight input according to its contribution weight to the output data, field sensitivity, and the degree of propagation decay of the cleaning action. Aftereffect label propagation weight The calculation formula is: ; in, The contribution weight of the atomic data unit to the output data is determined by methods such as linear regression and gradient analysis. This indicates the sensitivity of the output field to the input field, determined through perturbation analysis. This represents the propagation attenuation factor of the cleaning action, with a value range of [value range missing]. The determination is based on the type of cleaning action and the distance of blood relation; For fields that do not directly participate in numerical calculations but affect primary key matching, entity merging, event sorting, or window partitioning, their aftereffect markers are still propagated according to structural impact. If the same insight input is affected by multiple cleansing actions, the system synthesizes aftereffect markers according to lineage contribution and does not include cleansing actions with no lineage relationship in the aftereffect judgment of the insight. S34. Before outputting business insights, test the matching relationship between the post-cleaning effect and insight anomalies; for insight objectives... The system generates anomaly pattern vectors. The vector contains five components, each of which is normalized to 0. : Indicator of the intensity of the mutation: ; in, This indicates the current indicator value; Indicates the baseline value of the indicator; Indicates the baseline fluctuation range of the indicator; Indicates the intensity of local distribution anomalies: ; in, This represents the entropy of the current local data; This represents the entropy of local data at the baseline state; Indicating the intensity of time concentration: ; in, This indicates the number of events within the abnormal time window; This indicates the total number of events within the statistical period; Indicates the threshold proximity intensity: ; in, This indicates the preset threshold values for business metrics; Indicates the concentration intensity of the source: ; in, This indicates the number of abnormal events originating from the same source. Indicates the total number of abnormal events; For cleaning actions The system gains insights Extract the synthetic aftereffect vector related to the action from the bloodline path. The matching degree between the two is: ; in, Indicates insight into the target With cleaning action The degree of subsequent matching; Represents anomaly pattern vector The One component; Represents the synthesized aftereffect vector The Each component; if the denominator is zero, it indicates that there are no abnormal patterns or aftereffect markers available for matching, let ; Aftereffect matching threshold The range of values is In this embodiment, Set to a medium value; when the subsequent matching degree Exceed and the cleaning action Located in Insight When the insight is in the lineage path, mark it as a cleansing-induced candidate insight, instead of directly outputting it as a real business anomaly. The false alarm rate can be adjusted based on historical data. When the system frequently misjudges cleaning side effects as business anomalies, the false alarm rate can be reduced. .
[0027] Mod 4: Fixed the world generation and cross-validation modules.
[0028] like Figure 3 As shown, when uncertain cleaning results occur, multiple mutually exclusive repair worlds (independent data views formed based on different candidate repair hypotheses for the same original dataset) are retained, and cross-world verification is used to determine whether the insight is stable.
[0029] S41. Generate candidate repair hypotheses for uncertain anomalies; candidate repair hypotheses can represent retaining the original value, replacing field values, moving event time, splitting entity relationships, merging duplicate records, adjusting unit caliber, canceling a certain cleaning action, or temporarily suspending overwriting. The system performs hard constraint checks on candidate repair hypotheses. Hard constraint checks are performed using the following steps: 1. Field type check: Verify whether the data type of the candidate repaired field is consistent with the field definition type; 2. Value boundary check: Verify whether the candidate repaired data value is within the preset value range of the field; 3. Event sequence check: Verify whether the occurrence time of the candidate repaired events conforms to the business logic sequence; 4. Root evidence status check: Verify whether the root evidence on which the candidate repair depends is valid and has not been revoked; 5. Source consistency check: Verify whether the data after candidate repair is consistent with relevant data from other independent sources; 6. Primary key relationship check: Verify whether the candidate repaired data has damaged the primary key uniqueness and foreign key reference relationship; 7. Affected lineage subgraph check: Verify whether candidate repairs will lead to circular dependencies or logical contradictions in the lineage subgraph; If a candidate repair violates any of the above hard constraints, the candidate repair will be marked as invalid and will not proceed to the subsequent verification process. S42. Perform mutual exclusion judgment on candidate repair hypotheses; if two candidate repair hypotheses apply to the same data unit, the same entity relationship, the same event time, or the same aggregation caliber, and the two cannot be true at the same time, then generate different repair worlds; The criterion for determining whether two candidate repair hypotheses can both be true is: 1. The two candidate repair hypotheses assign different values to the same field in the same data unit; 2. The two candidate repair hypotheses assign different occurrence times to the same event, and these two times cannot coexist. 3. The two candidate repair hypotheses assign different associated objects to the same entity relationship; 4. The two candidate repair hypotheses assign different calculation rules to the same aggregation caliber; 5. Executing one candidate repair hypothesis will cause the premise of another candidate repair hypothesis to be invalid. If multiple candidate repairs are merely parameter changes under the same repair mechanism, the system can merge them into a parameterized repair world. The repair world uses a difference layer storage method. The main world stores the original data and confirmed cleaning results, while the difference layer only stores the data units, cleaning actions, aftereffect markers, and evidence status that the repair world has changed relative to the main world. The world difference repair layer uses a key-value pair storage structure, where the key is an identifier for an atomic data unit. The value represents the changes made to this atomic data unit relative to the Overworld in the current repair world, including the current value. Cleaning action sequence Aftereffect tag set and root evidence set This saving method avoids copying the entire data and facilitates playback of only the affected lineage subgraphs. Follow these steps to replay the bloodline subgraph: 1. Determine the extent of the lineage subgraph affected by candidate repairs; 2. Load the original data of the bloodline subgraph in the main world; 3. Apply the changes in the current repair world's difference layer to generate lineage subgraph data for that repair world; 4. Run the insight model on the generated kinship subgraph data; 5. Save the insights and compare them with those from other world-restoring efforts; S43, in each repaired world Multiple insight models are run within the system. These models can be rule-based, statistical detection, tree-based, time-series, graphical, linear, or neural network models. To avoid inconsistencies in the output dimensions of different models, the system first compares the insights from each model with the target insights. The output is transformed into a unified dimensionless insight value. After conversion The range of values is ; The specific conversion method for dimensionless insight values is as follows: For classification models, The original output probabilities are converted into calibrated probabilities using the Platt calibration method, which calibrates the original probabilities by training a logistic regression model. For numerical prediction models: ; in, Indicates the model's predicted value; Indicates the actual observed value; This indicates the historical normal fluctuation range of the indicator; For rule-based models: ; in, Indicates the rule trigger strength, with a value range of 100%. ; Indicates the confidence level of the rule, with a value range of 100%. ; S44, the Repair World Generation and Cross-Validation module calculates model consistency within the same repair world and repair consistency between different repair worlds, respectively: ; in, Meaning to repair the world Multiple models within the system provide insights into the target. Consistency; This indicates the interaction between multiple repair worlds and the insight target. Consistency; Indicates participation in insight objectives A collection of models; Indicates insight into the target The collection of worlds involved in the restoration; Representation Model In repairing the world The dimensionless insight value output by the middle; Meaning to repair the world The aggregated calibration values output by multiple models within the model; Indicates variance; Indicates the model's difference scale parameter; This indicates the parameter for repairing global disparity. Calibrate aggregate value The median, truncated mean, or calibrated weighted value can be used; in this embodiment, when a calibrated weighted value is used, the weight calculation formula is as follows: ; in, Representation Model Accuracy on the historical validation set; Indicates participation in insight objectives A collection of models; The formula for calculating the calibration aggregate value is: ; Model difference scale parameters The range of values is ; The range of normal fluctuations in historical stable insights is determined by collecting insights that have been manually confirmed as stable in the past, calculating the variance of multiple model outputs within the same repaired world, and taking the average of these variances as the mean. The initial value; Adjustments can be made based on the validation set playback results. When there is a significant difference between the model consistency calculation results and human judgment, adjustments can be made. The value; Repairing world difference scale parameters The range of values is ; The range of normal fluctuations in historical stable insights is determined by collecting insights that have been manually confirmed as stable in the past, calculating the variance of insight outputs across multiple repair worlds, and taking the average of these variances as the mean. The initial value; Adjustments can be made based on the validation set replay results. Adjustments can be made when the calculated world consistency differs significantly from manual judgment. The value; like lower and A higher value indicates that the main uncertainty stems from discrepancies in the model output; if higher and The lower values indicate that the main uncertainty comes from different repair assumptions; the system then triggers model calibration, supplements independent evidence, preserves the repair world, or prevents overwrite cleaning, respectively.
[0030] Mod5, Insight Output and Feedback Control Module.
[0031] The insight status is divided based on cross-world validation results, root evidence status, and post-matching results, and the direct impact of feedback flow on the cleaning model and remediation strategy is limited.
[0032] S51, Understanding the Target Calculate the overall credibility: ; in, Indicates insight into the target Overall credibility; This signifies restoring global consistency; This represents the median value indicating the consistency of the model across different repaired worlds; Indicates insight into the target The minimum level of independent evidence supporting the candidate repair hypothesis can be determined by relevant [relevant factors]. The smaller value is determined if Without relying on any candidate repair, then Take 1; Indicates insight into the target The highest feedback loop overlap among relevant candidate repairs can be determined by the relevant The larger value is determined if there is no feedback loop. Set to 0; Indicates insight into the target The highest aftereffect matching degree in the relevant cleaning actions can be determined by the relevant The larger value is determined; if no subsequent match exists, then... Set to 0; S52, according to and Classify the state of insight; Overall credibility output threshold The range of values is In this embodiment, Set to a medium value; when the overall confidence level is... Greater than or equal to At that time, the insight can be output as a stable insight; Adjustments can be made based on business needs. When the business has high requirements for the accuracy of insights, improvements can be made. When the business has high requirements for insight coverage, the [value] can be reduced. ; If an insight maintains a consistent direction across multiple repair worlds that satisfy hard constraints, and its overall credibility meets the output conditions, then the output is a stable insight. If the model consistency is high within the same repair world, but there are significant differences between different repair worlds, then the output is a repair-sensitive insight, along with candidate repair hypotheses, key data units, and supplementary root evidence that affect the insight. If an insight highly matches the post-cleaning effect, and the relevant cleaning action is located in the insight's lineage path, then the output is a cleaning-induced candidate insight, along with cleaning actions that may produce artificial traces. If the candidate repairs on which the insight depends have a high degree of overlap in feedback loops, then the output is an evidence-blocking state, and the insight cannot be used as a basis for automatic repair or a label for model training. S53. After receiving business feedback, the system first generates feedback data units using the multi-source access and atomization processing module, and then determines their root evidence set using the evidence conservation and backfeedback blocking module. If the feedback comes from independently collected facts and does not overlap with the root evidence exceeding the feedback loop threshold of the target data to be repaired, the feedback can participate in the verification of candidate repair hypotheses. If the feedback comes from insights, reports, strategy suggestions previously output by this system, or backflow data directly driven by the results, the feedback is only processed as a derived input and does not improve the repair credibility independently. For feedback that passes the evidence verification, the system can update the cleaning action parameters, the aftereffect label propagation weight, the repair world generation priority, or the model calibration parameters. For feedback that fails the evidence verification, the system retains its display and audit purposes, but does not use it as a direct basis for comprehensive cleaning or model training.
[0033] Through the detailed description of the above embodiments, the end-to-end data cleaning and intelligent business insight integrated learning system of the present invention achieves standardized processing of the entire process from data access to business insight output by constructing an integrated technical architecture that includes multi-source data atomic processing, end-to-end evidence chain control, quantitative propagation of post-cleaning effects, cross-validation of repair worlds, and closed-loop control of insight feedback. The system encapsulates each field value into an atomic data unit carrying complete metadata, ensuring the traceability of data sources; avoids data distortion caused by feedback self-verification through root evidence conservation and backfeedback blocking mechanisms; achieves identifiable cleaning side effects through post-cleaning effect labeling and propagation; quantifies the impact of different repair hypotheses on insight results through the generation and cross-validation of mutually exclusive repair worlds; and finally, classifies insight states and controls the feedback process based on multi-dimensional indicators, providing a systematic technical solution for enterprise data governance and business insights.
[0034] The preset parameters in the above formulas shall be set by those skilled in the art according to the actual situation.
[0035] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more sets of available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium. A semiconductor medium can be a solid-state drive.
[0036] It should be understood that in the various embodiments of this application, the order of the above-mentioned processes does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0037] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0038] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0039] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0040] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0041] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A full-link data cleaning and intelligent business insight integrated learning system, characterized in that: It includes a multi-source access and atomization processing module, an evidence conservation and backfeeding prevention module, a cleaning execution and after-effect marking module, a repair world generation and cross-validation module, and an insight output and feedback control module; The multi-source access and atomization processing module is used to receive multi-source data and encapsulate field values into atomic data units that record the original value, current value, source, time, root evidence set, cleaning action sequence, aftereffect marker set, and repair world identifier. The evidence conservation and backfeed blocking module is used to enable derived data to inherit the root evidence set of the parent data unit, and to verify the evidence of candidate repairs based on independent evidence support relationship and feedback loop overlap relationship. The cleaning execution and post-processing marking module is used to execute cleaning actions, record the post-processing markings corresponding to the cleaning actions, and propagate them along the data lineage. The repair world generation and cross-validation module is used to generate mutually exclusive repair worlds based on candidate repair hypotheses that have passed constraint checks, and run insight models in each repair world to determine model consistency and repair world consistency. The insight output and feedback control module is used to determine the insight state based on the root evidence state, the aftereffect matching result, the model consistency, and the repair world consistency, and to control the conditions for business feedback to enter repair verification or model update.
2. The end-to-end data cleaning and intelligent business insight integrated learning system according to claim 1, characterized in that, The multi-source access and atomic processing module is used to receive data from business databases, interface messages, logs, documents, message streams, and model recognition results; for structured data, it splits the data by field granularity; for semi-structured data, it parses field names, node paths, field values, and time information. For unstructured data, candidate field values are obtained through entity recognition, field extraction, layout parsing, or speech transcription, and parsing confidence information is added to the candidate field values.
3. The end-to-end data cleaning and intelligent business insight integrated learning system according to claim 2, characterized in that, The multi-source access and atomization processing module is also used to determine whether the atomic data unit originates from independently acquired facts. When the atomic data unit is generated by an independent observation, independent input, independent sensing, independent interface response, or independent manual verification, and its generation process does not depend on the cleaning results, insight results, or strategy suggestions previously output by the system, a root evidence identifier is generated based on the source channel, acquisition time, acquisition sequence number, original payload summary, and source signature. When independence cannot be confirmed, the corresponding data is marked as derived input.
4. The end-to-end data cleaning and intelligent business insight integrated learning system according to claim 1, characterized in that, The evidence conservation and backfeeding blocking module is used to determine the parent data unit set based on the input fields of the transformation operator, the aggregation window, the entity merging basis, the model input features, and the field lineage when the atomic data unit is obtained by calculation, transformation, aggregation, completion, merging, or model inference. The root evidence set carried by the parent data unit is used as the source of the derived root evidence. For candidate repair hypotheses, the independent evidence support is determined based on the deduplication result of the root evidence carried by the data unit that supports the candidate repair hypothesis. In the absence of traceable evidence, the candidate repair hypothesis is not allowed to directly cover the master data.
5. The end-to-end data cleaning and intelligent business insight integrated learning system according to claim 4, characterized in that, The evidence conservation and backfeeding blocking module is also used to determine the root evidence set of the target data to be repaired and the root evidence set on which the feedback data, insight results or backfeed data supporting candidate repair depend, and to determine the feedback loop overlap degree accordingly; when the feedback loop overlap degree meets the backfeeding blocking condition, the corresponding feedback is not allowed to increase the repair confidence or be used as a training label to update the cleaning model; for candidate repairs that pass the evidence verification, a repair imprint is generated including the candidate repair hypothesis identifier, root evidence summary, cleaning action type, action parameter summary, rule version, model version, action field, action time segment and affected lineage subgraph.
6. The end-to-end data cleaning and intelligent business insight integrated learning system according to claim 1, characterized in that, The cleaning execution and post-processing marking module is used to receive atomic data units to be processed and cleaning actions. The cleaning actions include field normalization, unit conversion, outlier handling, missing value completion, duplicate resolution, entity merging, time alignment, encoding conversion, and feature enhancement. When each cleaning action is executed, the input field, output field, triggering condition, action type, action parameters, affected fields, and affected data range are recorded. For cleaning actions involving overwriting original values, moving time, merging entities, or replacing anomalies, candidate repair hypotheses are first generated and then validated in the repair world.
7. The end-to-end data cleaning and intelligent business insight integrated learning system according to claim 6, characterized in that, The aftereffect markers include value change intensity, distribution compression intensity, time rearrangement intensity, substitution component intensity, and scope intensity; the cleaning execution and aftereffect marker module is used to pass aftereffect markers to output features and insight inputs when atomic data units participate in aggregation, connection, model input, or feature derivation, based on the contribution weight of atomic data units to output data, the sensitivity of output fields to input fields, and the degree of propagation attenuation of cleaning actions. For fields that affect primary key matching, entity merging, event sorting, or window partitioning, mark them according to the structural impact propagation effect.
8. The end-to-end data cleaning and intelligent business insight integrated learning system according to claim 7, characterized in that, The cleaning execution and aftereffect marking module is also used to generate an abnormal pattern vector of the insight target before the business insight is output, and extract the synthetic aftereffect vector related to the cleaning action from the lineage path of the insight target; when the matching relationship between the abnormal pattern vector and the synthetic aftereffect vector satisfies the aftereffect matching condition, and the corresponding cleaning action is located in the lineage path of the insight target, the insight target is marked as a cleaning-induced candidate insight.
9. The end-to-end data cleaning and intelligent business insight integrated learning system according to claim 1, characterized in that, The repair world generation and cross-validation module is used to generate candidate repair hypotheses for uncertain anomalies and to perform constraint checks on the candidate repair hypotheses on field type, value boundaries, event order, root evidence status, source consistency, primary key relationship, and affected lineage subgraph. When multiple candidate repair hypotheses act on the same data unit, the same entity relationship, the same event time, or the same aggregation caliber and cannot be true at the same time, different repair worlds are generated. The repair world adopts a difference layer storage method, in which the difference layer stores the data units that have changed relative to the main world, cleaning actions, aftereffect markers, and evidence status.
10. The end-to-end data cleaning and intelligent business insight integrated learning system according to claim 9, characterized in that, The repair world generation and cross-validation module is also used to replay the affected lineage subgraphs in each repair world and run multiple insight models to obtain insight values with uniform dimensions. It determines model consistency based on the differences in the outputs of multiple insight models within the same repair world and determines repair world consistency based on the differences in insight outputs between different repair worlds. The insight output and feedback control module is used to determine the comprehensive credibility based on repair world consistency, model consistency, independent evidence support, feedback loop overlap, and aftereffect matching, and classifies the insight status into stable insight, repair sensitive insight, cleansing and inducing candidate insight, or evidence blocking state. After receiving business feedback, the business feedback is first converted into feedback data units and its root evidence set is determined. When the business feedback comes from independently collected facts and does not trigger the backfeeding blocking condition, it is used to participate in the candidate repair hypothesis verification; otherwise, it is only processed as derived input or audit record.