An industrial data quality dynamic self-looping processing method and device

CN121901582BActive Publication Date: 2026-09-11CHINA ACADEMY OF INFORMATION & COMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511792996.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-01
Publication Date
2026-09-11
Estimated Expiration
2045-12-01

AI Technical Summary

Technical Problem

[0010]本申请提出一种工业数据质量动态自闭环加工方法和装置,解决现有工业数据质量治理方法依赖人工规则、缺乏语义理解能力、无法动态自适应及持续自我优化的问题

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121901582B_ABST
    Figure CN121901582B_ABST
Patent Text Reader

Abstract

The application provides an industrial data quality dynamic self-closed loop processing method and device, solves the problems that the existing method relies on manual rules, lacks semantic understanding ability, cannot dynamically adapt and continuously optimize itself. The method obtains multi-source heterogeneous data and determines abnormal semantic information including abnormal type, business context and causal prompt; determines the target processing strategy based on the information and the strategy selection basis, and executes the corresponding operator sequence; dynamically adjusts the strategy basis or the abnormal detection parameter according to the difference between the execution effect and the expected target. The application realizes the fundamental change from static rules to dynamic autonomy, significantly improves the accuracy, adaptability and intelligent level of data governance through semantic understanding and closed loop feedback mechanism, and effectively reduces the cost of manual intervention.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of industrial data governance technology, and in particular to a dynamic self-closed-loop processing method and apparatus for industrial data quality. Background Technology

[0002] With the deep integration of industrial internet, intelligent manufacturing, and artificial intelligence technologies, industrial enterprises' production processes are rapidly evolving towards data-driven, model-based decision-making, and intelligent autonomy. Industrial data has become a key production factor for enterprises to achieve lean management, process optimization, and intelligent decision-making, and its quality directly determines its availability and reliability at each stage. High-quality data is the cornerstone supporting key application scenarios such as predictive maintenance, production scheduling optimization, and energy consumption control, ensuring the accuracy of model algorithms and the stable operation of business systems. Conversely, low-quality data will directly lead to algorithm inaccuracies, control command failures, and even production safety hazards, becoming a significant bottleneck restricting the improvement of industrial intelligence.

[0003] However, in the current industrial data governance system, data quality processing is generally in a rudimentary stage of "static patching and manual intervention." Traditional methods mainly rely on predefined cleaning rules, periodic processing scripts, or manual intervention to correct anomalies, missing data, and drift issues that occur during data collection, transmission, and aggregation. This passive governance model was applicable in the early stages when the data scale was small and the application scenarios were simple, but as industrial scenarios become more complex, data types become more diverse, and real-time requirements increase, its inherent limitations become increasingly apparent, specifically in the following aspects: First, it lacks self-understanding capabilities and cannot identify complex semantic anomalies. Traditional data cleaning methods are mostly based on numerical rules or static threshold settings, which cannot identify and understand the semantic logic behind the anomalies. For example, in equipment operation data, sudden changes in temperature parameters may originate from noise at the measuring point, or they may reflect actual changes in equipment operating conditions or fault precursors. Static rules cannot distinguish the semantic attributes of these anomalies, which can easily lead to misjudgments or incorrect repairs, causing biases in subsequent model training and affecting the accuracy of decision-making.

[0004] Second, the lack of dynamic adaptability makes it difficult to cope with changes in operating conditions and data drift. The data characteristics of industrial systems are significantly dynamic and nonlinear, constantly changing with production cycles, environmental conditions, or equipment aging. Once static processing rules are set, they are difficult to adjust themselves to changes in data distribution. For example, factors such as increased acquisition frequency, sensor replacement, and equipment aging can all cause data distribution shifts. If the processing strategy cannot be dynamically updated accordingly, it will lead to distorted repair and data failure, and the governance effect will rapidly diminish.

[0005] Third, the lack of self-feedback and continuous learning mechanisms prevents optimization of the processing. Existing data governance systems typically operate on a unidirectional open-loop chain of "anomaly detection—cleaning and repair—result output," lacking feedback and learning mechanisms based on the quality of the processed results. The system cannot evolve through experience accumulation and model iteration, and the processing process cannot be self-evaluated or its effectiveness quantified, resulting in short-term and fragmented governance behaviors that hinder continuous improvement in governance levels.

[0006] Fourth, there is a lack of system collaboration and closed-loop linkage mechanisms. The data governance process is often scattered across multiple stages, such as the acquisition end, the transmission layer, and the application system, lacking a unified scheduling control and quality feedback mechanism. When data flows between multiple systems, abnormal information cannot be effectively transmitted back to the source, and repair experience cannot be shared between different stages, resulting in a fragmented state of "local governance, overall insensitivity," causing repeated errors and wasting governance resources.

[0007] In the new generation of industrial internet systems, the explosive growth of data volume and the increasing complexity of data flow paths have amplified the drawbacks of traditional static repair methods. In typical complex scenarios such as multi-station sensing in intelligent production lines, energy monitoring, predictive equipment maintenance, and multi-factory collaborative manufacturing, problems such as data temporal misalignment, semantic drift, signal deviation, and inconsistent definitions are particularly prominent. Relying on manual rules or periodic data cleaning is no longer sufficient to meet the real-time and accurate data quality assurance requirements of dynamic, high-frequency industrial scenarios.

[0008] Although some intelligent data cleaning methods have emerged in the industry, such as machine learning-based anomaly detection, rule engine-based autofill, and model-driven predictive repair, these methods still suffer from the core problem of "one-way processing and lack of closed loop": the anomaly detection algorithm and the repair module are isolated from each other and lack unified logical coordination; the repair results fail to feed back to the anomaly detection model to achieve self-calibration; and the various processing strategies exist in isolation, making it difficult to form a dynamic processing path that can be continuously optimized.

[0009] Therefore, the field of industrial data quality governance urgently needs a breakthrough technological solution that can achieve dynamic self-awareness, self-adaptation, and self-evolution of data quality. This requires the system to no longer rely on fixed rules or continuous human intervention, but to automatically identify the semantic and structural features of data through a self-understanding mechanism, automatically match and execute the optimal processing strategy chain based on a self-processing mechanism, and conduct quality assessment and model retraining of the processing results through a self-feedback mechanism. Ultimately, this forms a dynamic self-closed-loop governance system integrating "perception-processing-feedback-optimization," thereby fundamentally improving the intelligence level and long-term effectiveness of industrial data governance. Summary of the Invention

[0010] This application proposes a dynamic self-closed-loop processing method and apparatus for industrial data quality, which solves the problems of existing industrial data quality governance methods relying on manual rules, lacking semantic understanding capabilities, and being unable to dynamically adapt and continuously self-optimize.

[0011] In a first aspect, embodiments of this application provide a dynamic self-closed-loop processing method for industrial data quality, comprising the following steps: Acquire multi-source heterogeneous data to determine abnormal semantic information; the abnormal semantic information includes the abnormal type, business context, and causal indication. Based on the abnormal semantic information and the predefined strategy selection criteria, a target processing strategy is determined from multiple candidate processing strategies, and the data processing operator sequence corresponding to the target processing strategy is executed on the multi-source heterogeneous data; Evaluate the execution effect of the data processing operator sequence, and dynamically adjust the strategy selection criteria and / or anomaly detection parameters based on the difference between the evaluation results and the expected goals.

[0012] Furthermore, it also includes the following steps: The process involves a cyclical process of acquiring, determining, executing, evaluating, and adjusting. The loop stops when the difference between the evaluation result and the expected target is less than a set threshold.

[0013] In one embodiment, the step of determining the abnormal semantic information includes: Based on the multi-layer semantic mapping relationship of measurement points, equipment, processes and working conditions, the multi-source heterogeneous data is processed. The abnormal semantic information also includes abnormal location information.

[0014] In one embodiment, the strategy selection is based on a multi-objective optimization function; The multi-objective optimization function is configured to calculate the utility score of candidate processing strategies based at least on the amount of data quality improvement and the amount of business indicator improvement.

[0015] In one embodiment, the step of dynamically adjusting the criteria for strategy selection includes: Based on the feedback score of the execution effect, update the parameter weights in the strategy selection criteria.

[0016] In one embodiment, the step of dynamically adjusting the anomaly detection parameters includes: The probability threshold for anomaly detection is adaptively adjusted based on the combined changes in anomaly detection frequency and feedback scores.

[0017] In one embodiment, the dynamic adjustment step also triggers a feedforward governance instruction; The feedforward governance instructions are used to optimize the configuration parameters of the data acquisition terminal or the network transmission layer.

[0018] Secondly, embodiments of this application also provide an industrial data quality dynamic self-closed-loop processing device for implementing the industrial data quality dynamic self-closed-loop processing method described in any embodiment of the first aspect, comprising: an acquisition module for acquiring multi-source heterogeneous data; a determination module for determining abnormal semantic information; wherein the abnormal semantic information includes anomaly type, business context, and causal indication; and is further configured to determine a target processing strategy from multiple candidate processing strategies based on the abnormal semantic information and predefined strategy selection criteria; and an execution module for executing a data processing operator sequence corresponding to the target processing strategy on the multi-source heterogeneous data; and is further configured to evaluate the execution effect of the data processing operator sequence, and dynamically adjust the strategy selection criteria and / or anomaly detection parameters based on the difference between the evaluation result and the expected target.

[0019] Thirdly, embodiments of this application also provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method described in any one of the embodiments of the first aspect.

[0020] Fourthly, embodiments of this application also provide an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the method as described in any embodiment of the first aspect.

[0021] The above-described technical solutions adopted in the embodiments of this application can achieve the following beneficial effects: This application provides a dynamic self-closed-loop processing method and apparatus for industrial data quality. By constructing a closed-loop mechanism of "self-understanding-self-processing-self-feedback", it effectively solves the technical problems of traditional industrial data governance methods that rely on manual rules, lack semantic understanding capabilities, and cannot dynamically adapt and continuously self-optimize. Traditional methods usually rely on preset rule bases and fixed thresholds, and cannot automatically update governance strategies according to changes in operating conditions, which makes the processing results easily fail due to equipment fluctuations, noise interference, or process adjustments. Specifically, this application first establishes a multi-layered semantic mapping relationship of "measuring point-equipment-process-operating condition" to achieve deep association between data and business. It innovatively adopts a semantic anomaly quadruple that includes anomaly type, business context, and causal indication, enabling the system to recognize anomalies at the semantic level. This semantic enhancement mechanism can identify potential business anomalies before data deviations occur, giving the system a "proactive recognition" capability for industrial business logic, rather than merely remaining at the detection stage of data pattern changes. Second, based on predefined strategy selection criteria and a multi-objective optimization mechanism, it dynamically determines the optimal processing scheme from multiple candidate processing strategies, overcoming the shortcomings of fixed rules that are difficult to adapt to complex operating condition changes. Finally, by evaluating the execution effect and comparing it with the expected target, the system can automatically adjust the strategy selection criteria and anomaly detection parameters. It also innovatively introduces a feedforward governance mechanism, which optimizes data acquisition and transmission while completing data repair. The feedforward mechanism can proactively adjust the acquisition frequency, filtering parameters, or network transmission priority based on the current cycle's quality shortcomings, reducing the probability of anomalies and noise generation from the data source and achieving collaborative governance of "source prevention + end-point reinforcement." This complete technical solution transforms industrial data quality governance from the traditional "static repair, manual driving" model to a new paradigm of "dynamic autonomy, intelligent evolution." It not only significantly improves the accuracy and efficiency of data quality governance and greatly reduces the cost of manual intervention, but also provides continuous and reliable high-quality data support for intelligent manufacturing systems, effectively ensuring the stable operation and continuous optimization of intelligent industrial applications. Attached Figure Description

[0022] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings: Figure 1 This is a flowchart of a dynamic self-closed-loop processing method for industrial data quality according to an embodiment of this application; Figure 2 This is a structural diagram of an industrial data quality dynamic self-closed-loop processing device according to an embodiment of this application; Figure 3 This is a schematic diagram of the structure of an industrial data quality dynamic self-closed-loop processing device according to an embodiment of this application; Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0023] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0024] The technical solutions provided by the various embodiments of this application are described in detail below with reference to the accompanying drawings.

[0025] Figure 1 A flowchart of a dynamic self-closed-loop processing method for industrial data quality provided in this application embodiment includes the following steps: Step 110 to Step 130.

[0026] Step 110: Obtain multi-source heterogeneous data and determine abnormal semantic information; the abnormal semantic information includes the abnormal type, business context, and causal indication. This step is the core of achieving the "self-understanding" function, which aims to transform raw data into anomaly descriptions with business semantics.

[0027] Acquiring multi-source heterogeneous data: This step collects raw data from multiple sources, including the Enterprise Manufacturing Execution System (MES), Supervisory Control and Data Acquisition System (SCADA / DCS), equipment logs, and energy and quality systems. This multi-source heterogeneous data is diverse in type and origin, specifically referring to: "Multi-source" refers to the diversity of data sources, such as data coming from multiple independent industrial systems, including MES, SCADA / DCS, equipment controllers, sensor networks, and quality inspection systems.

[0028] Heterogeneity refers to the heterogeneity of data types and formats. For example, data includes continuous sensor readings (such as real numbers like temperature, pressure, and vibration), discrete equipment status (such as enumerated variables like running / stopped / faulted), event-based alarm logs, and batch-based production work order information. These data have different formats, sampling frequencies, and communication protocols.

[0029] This application uses a unified time base, caliber mapping, and master data alignment to perform preliminary processing of the aforementioned raw data, constructing a structured data representation that can be used for subsequent processing.

[0030] Determine the semantic information of the anomaly: The aforementioned anomaly semantic information is a structured data object, crucial for transitioning from numerical detection to semantic understanding. The core of this information lies in its inclusion of business semantics and causal relationships not found in traditional methods.

[0031] Core components (corresponding to steps 110-130): This abnormal semantic information contains at least the following three elements: Anomaly Type: A qualitative description of the pattern of anomalous data behavior. Examples include "missing data," "numerical mutation," "trend drift," or "data delay." This element is generated by the pattern classification submodule.

[0032] Business context: A feature vector describing the relevant business state and environmental conditions when an anomaly occurs. For example, {Process: "Fine Processing", Equipment Status: "High-Speed ​​Operation", Ambient Temperature: 35°C}. This element is provided by the semantic modeling module and is fundamental to understanding the meaning of the anomaly.

[0033] Causal hint: Indicative information about the possible root causes of an anomaly, provided by the causal inference model. For example, "may be related to a decrease in cooling water flow." This element is generated by the causal inference submodule and provides direction for subsequent remediation.

[0034] It should be noted that the acquisition steps include accessing structured / semi-structured / time-series data from different sources such as sensor streams, device logs, business systems, and edge acquisition nodes, and performing unified parsing and semantic modeling on the data. The system constructs a multi-level semantic graph of measurement points, equipment, processes, and operating conditions to identify entity relationships between data, and determines abnormal semantic information by combining contextual states, numerical fluctuation trends, and historical behavior patterns. The abnormal semantic information includes the abnormality type, business context, and causal indications, and further includes the location of the abnormality and suspected impact links, providing accurate decision-making basis for subsequent processing strategies.

[0035] In one embodiment, step 110, determining the abnormal semantic information, includes: Step 110-1: Based on the multi-layer semantic mapping relationship of measurement points, equipment, processes and working conditions, process the multi-source heterogeneous data; This is a key technical means to achieve "self-understanding". The system establishes a four-layer semantic mapping relationship π of "measuring point - equipment - process - operating condition", deeply binding the raw data with its corresponding physical equipment, technological process, and operating condition characteristics. Through this mapping, the raw data sequence is... Inject business context This forms an input sequence with contextual semantics. The aforementioned The business context feature vector corresponding to time t can include equipment operating mode, production plan cycle time, material batch attributes, environmental parameters, etc.

[0036] The system acquires raw datasets from multiple sources, including the Manufacturing Execution System (MES), Supervisory Control and Data Acquisition System (SCADA / DCS), Energy Monitoring System, Quality Inspection System, and Log Acquisition System. ,in For a moment The data (multi-source sampling point observation vector) includes equipment state variables, process variables, sensor readings, etc., which can be continuous (real numbers) or discrete (enumerated variables).

[0037] By unifying the time base, caliber mapping, and aligning with the master data, a structured data representation is constructed.

[0038] Establish semantic mapping relationships. This involves binding data to its physical equipment, process flow, and operating conditions to form an input sequence with contextual semantics. ,in: The context-enhanced input sequence is a structured dataset formed after semantic modeling and feature binding, used for subsequent anomaly identification and policy generation.

[0039] For the moment The corresponding business context feature vector includes contextual information such as equipment operating status, production plan, material batch, and environmental parameters.

[0040] Step 110-2: The abnormal semantic information also includes abnormal location information.

[0041] Preferably, in this embodiment, the abnormal semantic information is specifically implemented as a more complete semantic abnormality quadruple. Its generation process is as follows: The system uses a composite detector D_θ to perform anomaly identification on the context-enhanced input sequence X' obtained in step 110-1. Anomalies are determined by calculating the anomaly probability p(y_t) = σ(f_θ(x_t, c_t)) and comparing it with a threshold τ. Once an anomaly is confirmed, the system combines semantic mapping and a causal inference model to generate a structured result containing four elements: anomaly type, anomaly location, business context, and causal indication—that is, a semantic abnormality quadruple.

[0042] To clearly define the complete structure and technical implications of this semantically anomalous quadruple, the definitions, examples, and sources of its constituent elements are summarized in Table 1 below: Table 1

[0043] This semantically abnormal quadruple Type, location, context, causal hints This constitutes the core output of the system's "self-understanding" capability. It not only describes "what happened" (type) of the anomaly, but also clarifies "where it happened" (location), "in what environment it happened" (context), and "why it might have happened" (causal hint), thus providing accurate and interpretable input for subsequent "self-processing" and "self-feedback" steps.

[0044] The generation of the semantic anomaly quadruple relies on a backend integrated composite anomaly detection and causal inference model. The specific implementation process is as follows: The system uses a composite detector. The input data undergoes multidimensional feature analysis and anomaly detection. The detector combines rule constraints and a time series learning model to jointly judge the changing trends, contextual states, and historical patterns of multi-source data, and outputs the anomaly probability value at each time step. .

[0045] Formula 1 in: ( ) indicates a parameter. The detection function is used to extract data features and calculate anomaly scores; This is a normalization function that maps the output to the [0,1] interval; For a moment Input data; For a moment The corresponding business context feature vector.

[0046] when The time was judged as abnormal, among which The value is 0.8, and the abnormal quadruple results are output based on semantic modeling information: Formula 2 Used to describe the type of exception, where it occurred, the context, and possible causal clues.

[0047] in Generated by a causal inference model on a semantic graph, it is used to indicate possible causes of anomalies (such as sensor drift, changes in operating conditions, or interface delays), thereby achieving intelligent and semantic anomaly identification.

[0048] Step 120: Based on the abnormal semantic information and the predefined strategy selection criteria, determine the target processing strategy from multiple candidate processing strategies, and execute the data processing operator sequence corresponding to the target processing strategy on the multi-source heterogeneous data; This step is the core of the "self-processing" function, and its purpose is to dynamically select and execute the optimal data repair and optimization scheme based on the intelligent understanding of the anomaly.

[0049] The strategy selection criteria refer to the set of rules, functions, or models used by the system to evaluate and select the best data processing strategy.

[0050] For example, the strategy selection criteria are predefined, but its internal parameters (such as weights) can be dynamically adjusted during operation through subsequent feedback steps.

[0051] It should be noted that the strategy selection is based on a multi-objective optimization model, integrating data quality improvement, business metric enhancement, operator latency cost, and operator resource overhead. The comprehensive benefit of each candidate strategy is calculated using a utility scoring function. After determining the target strategy, the system automatically generates and schedules a sequence of data processing operators corresponding to the strategy, based on the required processing flow. This enables the coordinated execution of multiple types of operators, including missing value repair, anomaly filtering, confidence fusion, feature enhancement, and consistency alignment.

[0052] The candidate processing strategy refers to a series of selectable data processing flows formed by combining basic processing units (operators) in a pre-built strategy operator library in different ways.

[0053] For example, each of the candidate processing strategies is a directed acyclic graph (DAG) that represents a specific data processing path.

[0054] The target processing strategy refers to the strategy selected by the system from all candidate processing strategies based on the strategy selection criteria, which is considered optimal under the current abnormal semantic information and business context.

[0055] The data processing operator sequence, which is the specific execution step defined by the target processing strategy, is a sequence composed of one or more data processing operators arranged in a specific order.

[0056] The sequence of data processing operators will be called sequentially by the execution engine to complete the actual repair, correction or enhancement of abnormal data.

[0057] The detailed implementation process of step 120 is as follows: This step automatically generates processing strategy chains for different types of anomalies, enabling dynamic data repair and optimization.

[0058] This invention establishes a strategy operator library consisting of various operators for different anomaly characteristics, including operators such as data interpolation, time alignment, drift correction, confidence fusion, caliber standardization, and reversible desensitization.

[0059] Each operator is encapsulated as an independent callable module, possessing attributes such as delay cost, reversibility score, confidence output, and security label.

[0060] The system automatically combines the optimal operator chain (i.e., policy chain) to process data based on the anomaly type, contextual characteristics, and historical feedback.

[0061] During execution, the system records the link identifier, parameter configuration, and execution results, enabling full-process lineage tracing and rollback control.

[0062] Unlike traditional cleaning methods that rely on manual configuration and fixed rules, this invention enables dynamic generation of strategy chains, autonomous decision-making, and verifiable processes, significantly improving the autonomy and reliability of industrial data processing.

[0063] Strategy operator library design The system incorporates various data processing operators, including interpolation, denoising, temporal realignment, caliber correction, fragment replacement, and confidence-weighted fusion. A strategy operator library supports the dynamic combination and optimal scheduling of data processing actions in the self-processing stage. The operator library adopts a modular design, with each operator encapsulated as an independently callable processing unit, featuring standardized input / output interfaces and supporting parameter self-adjustment and reversible execution.

[0064] Operator libraries mainly include the following categories: 1) Data correction operators: used to correct numerical deviations and errors in physical measurement points.

[0065] Robust interpolation operator ( ): Based on quantile regression, Kalman filtering, or local linear models, imputation is performed on missing or outlier values.

[0066] Abnormal replacement operator ( ): Use the average of adjacent time windows or similar working condition samples to replace paragraphs.

[0067] Drift correction operator ( ): Long-term sensor offset is eliminated by sliding baseline estimation and trend fitting.

[0068] 2) Timing alignment operators: used to correct data misalignment caused by asynchronous sampling or clock drift.

[0069] Time realignment operator ( Phase matching of multi-source signals is achieved using the Dynamic Time Warping (DTW) algorithm. Sampling synchronization operator ( ): Automatic resampling and calibration of sampling frequency differences based on the principle of maximum mutual information.

[0070] 3) Operators for caliber and semantic correction: used to unify the measurement, units and semantic calibers between different systems.

[0071] Unit normalization operator ( Based on equipment metadata, it unifies the conversion of different units of measurement (such as MPa / bar, ℃ / K); Semantic mapping operator ( ): Automatic alignment and standardized expression of indicator semantics are achieved through knowledge graph mapping.

[0072] 4) Fusion and confidence enhancement operators: used to improve the credibility and robustness of processed data.

[0073] Confidence-weighted fusion operator ( ): Analyze multi-source signals according to the confidence matrix Weighted summation: Formula 3 This is the weighted and integrated comprehensive estimate.

[0074] The confidence weights are dynamically adjusted based on real-time noise characteristics.

[0075] When effective noise statistics are lacking, confidence-weighted fusion is performed using fixed quantization weights (example): The two sources are [0.60, 0.40][0.60, 0.40][0.60, 0.40]. The three sources are [0.50, 0.35, 0.15][0.50, 0.35, 0.15][0.50, 0.35, 0.15]. The four sources are [0.40, 0.30, 0.20, 0.10][0.40, 0.30, 0.20, 0.10][0.40, 0.30, 0.20, 0.10] (with a weighted sum of 1).

[0076] When a data source is determined to be suspicious or abnormal, its weight is reduced by 0.5 times and normalized as a whole (lower limit 0.05) to suppress the influence of low confidence signals on the results.

[0077] Multimodal consistency operator ( ): Joint verification of multimodal data, including visual, acoustic, and temporal data, is performed to ensure cross-modal consistency.

[0078] Data augmentation and desensitization operators: used to generate high-quality training samples and ensure security and compliance.

[0079] Data augmentation operators ( ): Improve data coverage through perturbation generation and distribution expansion; Reversible desensitization operator ( ): Performs encryption masking on sensitive fields and retains the reverse recovery key for post-audit.

[0080] Each operator node Carrying execution attribute vectors: Formula 4 in: : Execution delay cost; Reversibility score; Confidence output score; Safety compliance level.

[0081] In one embodiment, the strategy selection in step 120 is based on a multi-objective optimization function; The multi-objective optimization function is configured to calculate the utility score of candidate processing strategies based at least on the amount of data quality improvement and the amount of business indicator improvement.

[0082] In this embodiment, policy selection is formalized as a multi-objective optimization problem. The multi-objective optimization function is used to quantify the expected combined utility of each candidate policy.

[0083] The multi-objective optimization function is a mathematical function that comprehensively considers multiple, sometimes conflicting, objectives.

[0084] The utility score is a function used to calculate the numerical results of ranking and selecting candidate strategies.

[0085] Policy chain construction and adaptive scheduling The operators in the operator library are organized in a directed acyclic graph (DAG) manner, forming a chain of candidate policies. .

[0086] The system automatically selects operator combinations based on anomaly type, semantic context, and data characteristics, and optimizes the function through a multi-objective approach. Formula 5 The candidate chains are scored. Among them: Data quality improvement; : Increase in business KPIs; Computing cost; Link delay; The weighting coefficients are shown in Appendix Table 2. Table 2

[0087] To achieve efficient adaptive optimization of the policy chain, this invention introduces a hybrid strategy of Bayesian search and reinforcement learning in the process of policy chain generation and execution.

[0088] This strategy balances global optimization with rapid local exploration capabilities, and can dynamically adjust the algorithm path and parameter weights based on real-time feedback, forming an intelligent optimization closed loop of "exploration-evaluation-convergence".

[0089] Bayesian search mechanism: used to quickly determine high-yield strategy combinations in a multidimensional parameter space, achieving global approximate optimum.

[0090] Reinforcement learning mechanism: The probability of policy selection is dynamically updated through historical execution feedback to continuously optimize the convergence direction of the algorithm.

[0091] The system will automatically generate a traceable record chain after the policy is executed. This includes information such as strategy chain number, key parameters, timestamps, and rollback identifiers, enabling the processing process to be traceable, rollbackable, and verifiable.

[0092] Formula 6 This design ensures the intelligence and interpretability of the strategy chain optimization, enabling the system to have adaptive and self-evolving optimization capabilities when facing diverse industrial data tasks. It achieves traceability, rollback capability, and verifiability throughout the entire processing workflow.

[0093] Step 130: Evaluate the execution effect of the data processing operator sequence, and dynamically adjust the strategy selection criteria and / or anomaly detection parameters based on the difference between the evaluation results and the expected goals.

[0094] This step drives the continuous evolution and source-end governance of the system through feedback results, achieving long-term stable optimization of data quality. When a decrease in the data quality improvement rate or an increase in the abnormal repetition rate is detected, the system automatically triggers lightweight model retraining and policy reordering, forming an adaptive optimization mechanism. Simultaneously, this invention incorporates a feedforward governance component: when the system identifies high-frequency root causes (such as specific equipment drift, sampling period mismatch, network latency, etc.), it automatically generates feedforward governance instructions and sends them to the acquisition layer or network end for structural optimization. This mechanism realizes a shift from "post-event repair" to "pre-event prevention," constructing an industrial data governance system with self-healing characteristics.

[0095] This step enables the system to self-learn and optimize by calculating the quality changes and business improvements before and after processing. This step is used to evaluate the effectiveness of data processing and establish a result-based feedback learning mechanism to continuously optimize processing strategies and models.

[0096] This invention calculates the quality improvement by comparing data quality indicators (including completeness, consistency, accuracy, timeliness, and availability) before and after data processing, and combines this with changes in key performance indicators (KPIs) of downstream business systems to form a feedback score. The system automatically adjusts the strategy chain weights and detection thresholds based on the feedback results. When the feedback results are consistently positive, the priority of that strategy chain is strengthened; otherwise, model calibration or strategy replacement is triggered. This invention's self-feedback module achieves dynamic optimization by "feeding back the process with results," enabling the system to have self-learning and self-adjustment capabilities during continuous operation, thereby ensuring the long-term stability of data governance effects.

[0097] The anomaly detection parameters include at least a probability threshold for determining an anomaly.

[0098] The anomaly detection parameters constitute the core configuration set for the system's intelligent perception and adaptive optimization, primarily including the probability threshold, detection window size, feature weight coefficients, and sensitivity parameters. Among these, the probability threshold (τ) is the key criterion for anomaly identification, initially set to 0.8. The system determines the data to be abnormal when the anomaly probability p(yt) > τ. This threshold is dynamically adjusted through a mechanism τ. new = τ old + μ × (F - F ref To achieve adaptive optimization, where μ is the regulation rate and F... ref The target feedback value is defined by the detection window size, which determines the time range for anomaly detection. Feature weight coefficients balance the contribution of multi-dimensional features in anomaly identification, while the sensitivity parameter controls the system's response to data fluctuations. By continuously monitoring the coordinated changes in anomaly identification frequency and feedback scores, the system establishes a multi-parameter linkage adjustment mechanism, enabling the anomaly detection capability to continuously evolve during operation and ensuring optimal detection performance under different working conditions.

[0099] In one embodiment, step 130, the step of dynamically adjusting the criteria for strategy selection, includes: Step 130-1: Update the parameter weights in the strategy selection criteria based on the feedback score of the execution effect.

[0100] Quality and effect evaluation The system compares data before and after processing to calculate the overall quality improvement rate. Formula 7 in , These represent the data quality scores before and after processing, which are automatically generated based on indicators such as data integrity, accuracy, consistency, timeliness, and availability.

[0101] Business performance feedback The system inputs the processed data into downstream business systems (such as equipment predictive maintenance, energy consumption analysis, quality traceability, etc.) and automatically monitors key business indicators (KPIs).

[0102] If business indicators improve If the value is greater than the preset threshold, it indicates that the data processing is effective; otherwise, a strategy adjustment will be triggered.

[0103] Self-feedback and strategy update mechanism Based on the comprehensive feedback results, the system automatically generates a feedback score: Formula 8 in, and For the system's self-learning weights, To improve learning weights for quality improvement, Weighting for business performance.

[0104] By default, =0.6, the system focuses more on improving data quality in the initial stage; =0.4, and will be dynamically adjusted by the system through self-learning, with a focus on comprehensive optimization.

[0105] when When >0, the system records successful strategy chains and parameter configurations, and updates the experience base; when If <0, automatically roll back to the previous version and retrain the parameters.

[0106] The feedback module also supports dynamic learning of the link: the system will count the average feedback effect of the same type of anomaly under different processing strategies, and use it for the next strategy optimization, realizing a self-evolution mechanism of "optimizing one type with one use".

[0107] This module enables the system to possess intelligent features such as automatic self-learning and rapid convergence in real-world industrial scenarios through a result-oriented, lightweight feedback mechanism.

[0108] In one embodiment, step 130, the step of dynamically adjusting the anomaly detection parameters, includes: Step 130-2: Adaptively adjust the probability judgment threshold used for anomaly identification based on the combined changes in anomaly identification frequency and feedback scores.

[0109] Threshold adaptive adjustment The system automatically adjusts the detection threshold based on the frequency of anomaly detection and changes in feedback scores over a recent period. Formula 9 in This is the threshold adjustment rate, used to control the convergence speed of threshold updates. The target feedback value (default is 0.8 initially). The threshold value is the one used for the previous detection (the initial default value is 0.6). This linear adjustment method is more flexible than the traditional fixed threshold method and can automatically stabilize to the optimal range depending on the scenario.

[0110] Lightweight self-calibration of the model The system makes minor adjustments to the parameters of the detection model and repair operator based on high-quality samples accumulated in the experience base.

[0111] This process does not require full retraining; it only updates the influence weights and confidence intervals, achieving a lightweight update that "learns as it runs".

[0112] Policy priority reordering The system periodically evaluates the average feedback score of various policy chains and dynamically adjusts their priorities based on their long-term benefits, giving the optimal policy a higher scheduling weight. This mechanism ensures that the system tends towards a stable and predictable optimal state over time.

[0113] In one embodiment, in step 130, the dynamic adjustment step further triggers a feedforward governance command; The feedforward governance instructions are used to optimize the configuration parameters of the data acquisition terminal or the network transmission layer.

[0114] Feedforward governance mechanism When the feedback module continuously identifies the same root cause (such as sensor malfunction or excessive interface latency in a device), the system automatically generates feedforward governance instructions and sends them to the source end to execute optimization tasks, including: Adjust the sampling frequency or time synchronization strategy; Increase redundancy at key measurement points; Optimize data buffers or network bandwidth.

[0115] The feedforward mechanism enables a shift from "post-event repair" to "pre-event prevention," giving data quality governance self-healing and forward-looking characteristics.

[0116] Furthermore, it also includes the following steps: Step 140: Repeat the steps of acquiring, determining, executing, evaluating, and adjusting. This step defines the basic working mode of the aforementioned dynamic self-closed-loop processing method for industrial data quality. Specifically, the "steps of acquisition, determination, execution, evaluation, and adjustment" refer to the cyclical execution of steps 110 to 130 in the embodiments of this application. Through this cyclical operating mechanism, the system constructs a continuous "perception-cognition-decision-execution-learning" closed loop, making data quality governance no longer a one-off task, but an autonomous process capable of adapting to changes in operating conditions and continuously self-optimizing.

[0117] Step 150: In response to the difference between the evaluation result and the expected target being less than a set threshold, stop the loop.

[0118] This step defines the intelligent termination condition of the loop, which is key to the system's "goal-driven" and "economical" characteristics. The "evaluation result" comes from the quantitative output of the data processing effect in step 130 (e.g., comprehensive feedback score F or data quality improvement ΔQ). The "expected goal" is a preset quality or business performance benchmark of the system. When the difference between the two is less than a set threshold, it indicates that the data quality has reached a stable and satisfactory state, and the system stops the loop in response to this condition, thereby avoiding unnecessary consumption of computing resources. This mechanism ensures that the system can automatically enter a standby state after achieving the governance goal, until a new anomaly or change triggers a new round of processing loop.

[0119] Figure 2 This is a structural diagram of an industrial data quality dynamic self-closed-loop processing device provided in an embodiment of this application.

[0120] An industrial data quality dynamic self-closed-loop processing apparatus, used to implement the industrial data quality dynamic self-closed-loop processing method described in any embodiment of the first aspect, comprising: The acquisition module 201 is used to acquire multi-source heterogeneous data.

[0121] The determination module 202 is used to determine abnormal semantic information; the abnormal semantic information includes abnormal type, business context and causal indication; it is also used to determine the target processing strategy from multiple candidate processing strategies based on the abnormal semantic information and predefined strategy selection criteria.

[0122] The execution module 203 is used to execute the data processing operator sequence corresponding to the target processing strategy on the multi-source heterogeneous data; it is also used to evaluate the execution effect of the data processing operator sequence, and dynamically adjust the strategy selection criteria and / or anomaly detection parameters based on the difference between the evaluation result and the expected target.

[0123] In one embodiment, the acquisition module includes a first acquisition unit for acquiring multi-source heterogeneous data.

[0124] The determining module includes a first determining unit, used to determine abnormal semantic information; the abnormal semantic information includes the abnormal type, business context, and causal indication.

[0125] It also includes a second determining unit, used to determine a target processing strategy from multiple candidate processing strategies based on the abnormal semantic information and predefined strategy selection criteria.

[0126] The execution module includes a first execution unit, which is used to execute the data processing operator sequence corresponding to the target processing strategy on the multi-source heterogeneous data.

[0127] It also includes a second execution unit for evaluating the execution effect of the data processing operator sequence, and dynamically adjusting the strategy selection criteria and / or anomaly detection parameters based on the difference between the evaluation results and the expected target.

[0128] In one embodiment, the execution module further includes a third execution unit for cyclically running the functions of the acquisition module, the determination module, and the execution module; and stopping the loop in response to the difference between the evaluation result and the expected target being less than a set threshold.

[0129] The above embodiments are used to implement the contents of steps 140 and 150 in the specification.

[0130] In one embodiment, the first determining unit is specifically used to process the multi-source heterogeneous data based on the multi-layer semantic mapping relationship of measuring points, equipment, processes and working conditions; and the abnormal semantic information includes abnormal location information.

[0131] The above embodiments are used to implement the contents of steps 110-1 and 110-2 in the specification.

[0132] In one embodiment, the second determining unit is specifically used to calculate the utility score of the candidate processing strategy through a multi-objective optimization function; the multi-objective optimization function is configured to calculate based at least on the amount of data quality improvement and the amount of business indicator improvement.

[0133] The above embodiments are used to implement the content regarding the multi-objective optimization function in step 120 of the specification.

[0134] In one embodiment, the second execution unit is specifically used to update the parameter weights in the strategy selection criteria based on the feedback score of the execution effect.

[0135] The above embodiments are used to implement the content of step 130-1 in the specification.

[0136] In one embodiment, the second execution unit is further configured to adaptively adjust the probability determination threshold for anomaly identification based on the combined change of anomaly identification frequency and feedback score.

[0137] The above embodiments are used to implement the content of step 130-2 in the specification.

[0138] In one embodiment, the second execution unit is further configured to trigger a feedforward governance instruction; the feedforward governance instruction is used to optimize the configuration parameters of the data acquisition end or the network transmission layer.

[0139] The above embodiments are used to implement the content regarding the feedforward governance mechanism in step 130 of the specification.

[0140] Figure 3 This is a schematic diagram of the structure of the industrial data quality dynamic self-closed-loop processing system provided in an embodiment of this application. (See attached diagram.) Figure 3 As shown, the system specifically includes the following modules: The data access and semantic modeling module is used to implement the functions of the first acquisition unit in the acquisition module. The data access and semantic modeling module supports unified access to sensor stream data, equipment operation logs, production execution system data, quality inspection data, and network-side monitoring data; and transforms heterogeneous data into a unified intermediate structure through protocol parsing, time sequence alignment, and semantic association for subsequent processing; the data access and semantic modeling module can also communicate with edge nodes or field acquisition devices to realize feedforward adjustment of data acquisition frequency, caching strategy, and sampling quality.

[0141] Furthermore, the data access and semantic modeling module can automatically select the optimal access protocol stack based on the protocol type, sampling period, and business importance of the data source, and supports automatic identification and conversion of various industrial communication protocols such as OPC UA, MQTT, Modbus, PROFINET, and HTTP / REST.

[0142] Meanwhile, the data access and semantic modeling module has a data quality screening function, which can quickly detect the integrity, legality, noise level and timestamp accuracy of the data during the access stage, and provide quality identification labels for the subsequent data processing module.

[0143] The data access and semantic modeling module can also integrate a lightweight edge AI model to perform real-time compression, noise reduction, anomaly filtering, and event triggering on-site data, thereby reducing the amount of data sent and improving network-side transmission efficiency.

[0144] The data access and semantic modeling module can automatically adjust the field acquisition strategy according to the feedforward governance strategy, such as increasing the sampling frequency of key process measurement points, reducing the acquisition redundancy in a stable state, and enabling field data buffering to cope with network jitter, thereby realizing the dynamic self-optimization capability of the acquisition end.

[0145] In addition, the data access and semantic modeling module can synchronize data with the digital twin model or industrial control platform in the industrial field, so that the status of the physical equipment on site is consistent with the data access status, and supports automatic updates of the collection point table and equipment topology information when equipment is maintained, process is switched or production line is changed.

[0146] By working in conjunction with the system's semantic modeling function, the data access and semantic modeling module can automatically inject semantic tags of "measurement point → equipment → process → operating condition" during the data access phase, providing a structured semantic foundation for the upper-level modules of anomaly identification and strategy selection. Specifically, it is responsible for collecting raw data from multiple sources, including the production execution system, monitoring and data acquisition system, equipment logs, and energy and quality system, and constructing a structured data representation through a unified time benchmark, caliber mapping, and master data alignment, while establishing a multi-layered semantic mapping relationship of "measurement point-equipment-process-operating condition".

[0147] The data access and semantic modeling module not only enables unified access and high-quality acquisition of multi-source heterogeneous data, but also provides intelligent self-optimization capabilities that can be adjusted forward, semantically enhanced, and edge-coordinated.

[0148] The anomaly perception and cognition module is used to implement the functions of the first and second determination units in the determination module. This anomaly perception and cognition module can combine rule engine, time series prediction model, semantic embedding model and causal inference algorithm to jointly judge abnormal fluctuations in data, cross-device correlation, process deviation, sudden changes in equipment status, etc., thereby generating a semantic anomaly vector containing anomaly type, trigger line, impact link and possible root cause, providing structured perception input for subsequent strategy optimization.

[0149] The anomaly perception and cognition module is further used to determine the target processing strategy from multiple candidate processing strategies based on the anomaly semantic information and predefined strategy selection criteria. This module analyzes the anomaly semantic information to identify key business processes, equipment topology locations, and data link characteristics involved in the anomaly. It then combines this with a predefined multi-objective optimization function to evaluate the utility of different strategy combinations and automatically selects the processing scheme that best matches the current anomaly characteristics. Strategy selection criteria may include multi-dimensional factors such as data quality improvement rate, business performance improvement potential, operator execution cost, latency constraints, and upstream and downstream dependencies, enabling the system to dynamically select the optimal data processing path under complex operating conditions.

[0150] Furthermore, the anomaly perception and cognition module can maintain an anomaly knowledge base and a strategy knowledge base. By continuously accumulating anomaly features, processing strategies, and feedback effects, it can achieve rapid response and intelligent recommendation for similar anomalies in the future. Simultaneously, it supports a strategy weighting and filtering mechanism based on historical representations, enabling the strategy decision-making process to possess both interpretability and self-evolutionary capabilities. Specifically, it is responsible for processing multi-source heterogeneous data based on semantic mapping relationships, using a composite detector for anomaly identification, and outputting a semantic anomaly quadruple containing anomaly type, location, context, and causal hints. Simultaneously, it determines the target processing strategy based on the anomaly semantic information and predefined strategy selection criteria.

[0151] The strategy chain generation and execution module is used to implement the function of the first execution unit in the execution module, specifically responsible for executing the corresponding data processing operator sequence on multi-source heterogeneous data according to the target processing strategy.

[0152] The quality assessment and feedback module implements the functions of the second execution unit in the execution module. This module is responsible for evaluating the execution effect of the data processing operator sequence. The evaluation process includes calculating the data quality improvement (ΔQ) and business indicator improvement (ΔK) before and after processing, and combining operator overhead and execution latency to obtain a comprehensive feedback score F, forming a quantifiable representation of the processing effect. Specifically, it is responsible for evaluating the execution effect of the data processing operator sequence, calculating the data quality improvement and business indicator improvement, and generating a comprehensive feedback score.

[0153] The evolution control module implements the functions of the third execution unit in the execution module. Based on the deviation between the feedback score and the preset target score, this module triggers adaptive adjustments to the weight parameters in the strategy selection criteria. Simultaneously, it automatically corrects anomaly detection parameters (including probability threshold, window size, confidence interval, etc.) based on data fluctuation levels and anomaly identification accuracy. This module can also construct an execution performance evaluation model based on historical feedback results, enabling prediction and advance optimization of future operator chain performance, allowing the system to gradually develop self-learning, self-calibration, and self-evolutionary processing capabilities.

[0154] Furthermore, the evolution control module can also be linked with feedforward governance to automatically generate feedforward governance instructions when structural quality problems at the data source level (such as time drift, noise accumulation, and missing sampling) are identified, enabling dynamic adjustment of upstream parameters such as acquisition frequency, edge caching strategy, and network transmission priority. Specifically, it is responsible for adaptive threshold adjustment, lightweight model self-calibration, and policy priority reordering based on feedback results, and triggering feedforward governance instructions when high-frequency root causes are identified.

[0155] The genealogy tracing and auditing module is used to record the entire process operation log and version information, supporting traceability and rollback of the processing process.

[0156] The interface and display module provides a visual monitoring interface and open API interfaces.

[0157] Each module interacts collaboratively through event streams and message buses, and can be deployed independently in the cloud, at the edge, or locally to jointly complete the dynamic self-closed-loop processing of industrial data quality.

[0158] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0159] Therefore, this application also proposes a computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the methods described in any embodiment of this application.

[0160] Furthermore, this application also proposes an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the method described in any embodiment of this application.

[0161] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0162] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0163] These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable apparatus for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0164] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory. Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0165] Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. The electronic device 600 shown is merely an example and should not impose any limitations on the function and scope of use of the embodiments of this application. It includes: one or more processors 420; and a storage device 410 for storing one or more programs. When the one or more programs are executed by the one or more processors 420, the one or more processors 420 implement the dynamic self-closed-loop processing method for industrial data quality provided in the embodiments of this application. The method includes: Acquire multi-source heterogeneous data to determine abnormal semantic information; the abnormal semantic information includes the abnormal type, business context, and causal indication. Based on the abnormal semantic information and the predefined strategy selection criteria, a target processing strategy is determined from multiple candidate processing strategies, and the data processing operator sequence corresponding to the target processing strategy is executed on the multi-source heterogeneous data; Evaluate the execution effect of the data processing operator sequence, and dynamically adjust the strategy selection criteria and / or anomaly detection parameters based on the difference between the evaluation results and the expected goals.

[0166] The electronic device 400 also includes an input device 430 and an output device 440; the processor 420, storage device 410, input device 430 and output device 440 in the electronic device can be connected by a bus or other means, as shown in the figure, which is connected by a bus 450.

[0167] Storage device 410, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and module units, such as the program instructions corresponding to the dynamic self-closed-loop processing method for industrial data quality in this embodiment. Storage device 410 may mainly include a program storage area and a data storage area. The program storage area may store the operating system and at least one application program required for a function; the data storage area may store data created based on terminal usage. Furthermore, storage device 410 may include high-speed random access memory and non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some instances, storage device 410 may further include memory remotely located relative to processor 420, and these remote memories can be connected via a network. Examples of such networks include, but are not limited to, the Internet, enterprise intranets, local area networks, mobile communication networks, and combinations thereof.

[0168] Input device 430 can be used to receive input digital, character, or voice information, and to generate key signal inputs related to user settings and function control of the electronic device. Output device 640 may include electronic devices such as a display screen and a speaker.

[0169] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0170] Those skilled in the art will understand that, unless otherwise stated, the singular forms “a,” “an,” “the,” and “the” used herein may also include the plural forms. It should be understood that when a device or component is “connected” to another device or component, it may be directly connected to the other device or component, or there may be an intermediary device or component. Furthermore, the term “connection” as used herein may include partially wireless connections as well as partially wired connections.

[0171] In the description of this application, it should be understood that the terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance. Those skilled in the art can understand the specific meaning of the above terms in this application based on the specific circumstances. Furthermore, in the description of this application, unless otherwise stated, "multiple" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship.

[0172] The above description is merely an embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the embodiments of this application.

Claims

1. A dynamic self-closed-loop processing method for industrial data quality, characterized in that, Includes the following steps: Acquire multi-source heterogeneous data to determine abnormal semantic information; the abnormal semantic information includes the abnormal type, business context, and causal indication. The causal indication is an indicative information about the root cause of the anomaly, provided by a causal inference model. Based on the abnormal semantic information and the predefined strategy selection criteria, a target processing strategy is determined from multiple candidate processing strategies, and the data processing operator sequence corresponding to the target processing strategy is executed on the multi-source heterogeneous data; the strategy selection criteria are constructed using a multi-objective optimization model; the utility of different strategy combinations is evaluated, and the processing scheme that best matches the current abnormal characteristics is automatically selected; Evaluate the execution effect of the data processing operator sequence, and dynamically adjust the strategy selection criteria and / or anomaly detection parameters based on the difference between the evaluation results and the expected goals; the anomaly detection parameters include at least a probability judgment threshold for determining anomalies.

2. The industrial data quality dynamic self-closed-loop processing method according to claim 1, characterized in that, It also includes the following steps: The process involves a cyclical process of acquiring, determining, executing, evaluating, and adjusting. The loop stops when the difference between the evaluation result and the expected target is less than a set threshold.

3. The industrial data quality dynamic self-closed-loop processing method according to claim 1, characterized in that, The steps for determining the abnormal semantic information include: Based on the multi-layer semantic mapping relationship of measurement points, equipment, processes and working conditions, the multi-source heterogeneous data is processed. The abnormal semantic information also includes abnormal location information.

4. The industrial data quality dynamic self-closed-loop processing method according to claim 1, characterized in that, The strategy selection criteria include a multi-objective optimization function; The multi-objective optimization function is configured to calculate the utility score of candidate processing strategies based at least on the amount of data quality improvement and the amount of business indicator improvement.

5. The industrial data quality dynamic self-closed-loop processing method according to claim 1, characterized in that, The steps for dynamically adjusting the criteria for strategy selection include: Based on the feedback score of the execution effect, update the parameter weights in the strategy selection criteria.

6. The industrial data quality dynamic self-closed-loop processing method according to claim 1, characterized in that, The steps for dynamically adjusting the anomaly detection parameters include: The probability threshold for anomaly detection is adaptively adjusted based on the combined changes in anomaly detection frequency and feedback scores.

7. The industrial data quality dynamic self-closed-loop processing method according to claim 1, characterized in that, The dynamic adjustment step also triggers a feedforward governance command; The feedforward governance instructions are used to optimize the configuration parameters of the data acquisition terminal or the network transmission layer.

8. An industrial data quality dynamic self-closed-loop processing device, used to implement the industrial data quality dynamic self-closed-loop processing method according to any one of claims 1 to 7, characterized in that, include: The acquisition module is used to acquire heterogeneous data from multiple sources. The determination module is used to determine the abnormal semantic information; the abnormal semantic information includes the abnormal type, business context, and causal indication. It is also used to determine the target processing strategy from multiple candidate processing strategies based on the abnormal semantic information and predefined strategy selection criteria; The execution module is used to execute the data processing operator sequence corresponding to the target processing strategy on the multi-source heterogeneous data; it is also used to evaluate the execution effect of the data processing operator sequence, and dynamically adjust the strategy selection criteria and / or anomaly detection parameters based on the difference between the evaluation result and the expected target.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method as described in any one of claims 1-7.

10. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Systems and methods for evaluating and repairing data using data quality indicators

    CN116097628A

  • Database anomaly detection method and device based on multi-modal time sequence fusion

    CN120892317A