Intelligent operation and maintenance tracing monitoring system and method for industrial computer failure
Patent Information
- Application Number
- CN202610916230.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-24
- Publication Date
- 2026-09-01
AI Technical Summary
数据采集仅简单采集单维度硬件参数,缺少环境干扰剔除逻辑;故障识别依赖海量人工标注故障样本,采用固定静态图谱匹配故障;运维环节仅依据预设固定规则执行处置动作,故障处置完成后无数据回流优化流程,各技术环节完全割裂,不存在协同联动的一体化架构设计
Smart Images

Figure CN122673013A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of industrial intelligent operation and maintenance technology, specifically to an intelligent operation and maintenance tracing and monitoring system and method for industrial computer faults. Background Technology
[0002] Existing publicly available documents on industrial computer operation and maintenance technologies all adopt a modular and independent design approach. The data acquisition module, fault identification module, and operation and maintenance execution module operate independently, and there is no bidirectional data linkage or reverse iteration mechanism between the modules. Data acquisition simply collects single-dimensional hardware parameters and lacks logic for eliminating environmental interference; fault identification relies on massive amounts of manually labeled fault samples and uses fixed static graphs to match faults; the operation and maintenance process only performs handling actions according to preset fixed rules, and there is no data feedback optimization process after fault handling is completed. The various technical links are completely isolated, and there is no integrated architecture design with collaborative linkage.
[0003] The existing distributed operation and maintenance architecture has given rise to a unified core technical problem that cannot be addressed simultaneously. That is, the various independent technologies currently available cannot build a complete closed-loop system that coordinates multi-source data purification, small-sample fault tracing, hierarchical adaptive operation and maintenance, and autonomous model iteration. In industrial production sites with strong interference and rare, sporadic faults, this can simultaneously lead to false alarms and missed alarms in monitoring, the inability to locate the underlying root cause of cascading hidden faults, mismatch between operation and maintenance solutions and fault scenarios, and multiple derivative defects caused by the repeated occurrence of similar faults. This lack of a closed loop is the fundamental reason why the existing technical system cannot solve all the pain points of industrial control operation and maintenance at the same time, and it is also the only core technical challenge that the overall solution of this invention focuses on overcoming.
[0004] In view of the above, this application is hereby submitted. Summary of the Invention
[0005] The purpose of this invention is to provide an intelligent operation and maintenance traceability monitoring system and method for industrial computer faults, so as to solve the problems mentioned in the background art.
[0006] To solve the above-mentioned technical problems, the present invention provides an intelligent operation and maintenance tracing and monitoring system for industrial computer faults, which includes a multi-source hierarchical data acquisition and purification module, a dynamic fault tracing and analysis module, an adaptive closed-loop operation and maintenance decision-making module, and a visual early warning display module. The multi-source hierarchical data acquisition and purification module builds a four-layer synchronous acquisition architecture consisting of hardware, system, business, and environmental layers. It constructs a fixed environmental interference feature library and uses a residual comparison algorithm to remove abnormal data corresponding to environmental interference in the industrial field, thereby purifying the real operating data of the industrial computer itself. The dynamic fault tracing and analysis module builds a small-sample comparative learning model based on normal operation samples, mines hidden fault characteristics, dynamically constructs a fault propagation map, and matches a three-level fault hierarchy attribution system to complete the quantitative location of fault root causes. The adaptive closed-loop operation and maintenance decision module binds the fault quantification and tracing results to match the hierarchical operation and maintenance strategy, executes the corresponding fault handling operations, and simultaneously accumulates operation and maintenance data to iteratively optimize the front-end data purification rules and tracing model parameters. The visualization and early warning display module shows the fault tracing results, fault propagation paths, and operation and maintenance execution status, and outputs fault risk warning information. Through a core linkage architecture of four-layer hierarchical data purification, small-sample dynamic graph quantification tracing, and tracing-driven operation and maintenance closed-loop iteration, it overcomes the shortcomings of existing technologies such as large data noise interference, high fault sample dependence, low tracing accuracy, rigid operation and maintenance strategies, and lack of autonomous optimization capabilities. It realizes accurate monitoring, root cause tracing, and adaptive intelligent operation and maintenance of all types of industrial computer faults under complex industrial conditions, forming a complete and autonomously evolving fault management closed loop. The overall architecture is different from the fragmented technical solutions of traditional single monitoring or single tracing.
[0007] Furthermore, the multi-source hierarchical data acquisition and purification module includes a hardware data acquisition unit, a system data acquisition unit, a business data acquisition unit, an environmental data acquisition unit, and a data interference stripping and purification unit. The hardware data acquisition unit collects real-time operating status data of the industrial computer's motherboard, network card, hard drive, and power supply. The system data acquisition unit collects process, port, driver, and system operation log data of the industrial computer. The business data acquisition unit collects industrial business interaction, data transmission, and instruction execution data of the industrial computer. The environmental data acquisition unit collects environmental data such as electromagnetic intensity, temperature, humidity, vibration, and dust concentration in the industrial site. The data interference stripping and purification unit calls the environmental interference feature library and uses a residual comparison algorithm to remove false abnormal data caused by environmental factors, outputting a structured and clean fault dataset. By accurately collecting full-dimensional operating and environmental data through multi-unit hierarchical layering, the module achieves full coverage of industrial computer operating data collection. By accurately stripping environmental interference data through a dedicated purification unit, the module eliminates false alarms and missed alarms from the data source, providing high-quality and high-accuracy data support for subsequent fault tracing and analysis.
[0008] Furthermore, the dynamic fault tracing and analysis module includes a small-sample comparative learning and recognition unit. This unit trains the model based on samples of normal, routine operation of industrial computers and uses the distribution characteristics of normal data to identify abnormal fault features, thus accurately identifying both overt and covert faults. It abandons the traditional fault identification model's reliance on massive amounts of labeled fault samples, adapts to the industry characteristics of low-probability fault occurrence and scarce fault samples in industrial computers, effectively solves the technical problems of fault identification failure and low accuracy in small-sample scenarios, and overcomes inherent technical biases in the industry.
[0009] Furthermore, the dynamic fault tracing and analysis module includes a dynamic fault propagation graph construction unit. This unit correlates the temporal change characteristics of hierarchical operational data in real time, dynamically updates the fault node weights and node associations, and reconstructs the complete propagation path of the fault from the hardware layer to the business layer in real time. It replaces the traditional static fault knowledge graph, realizes the dynamic real-time reconstruction of the fault propagation process, accurately captures the diffusion patterns of multi-level chain faults, and solves the problems of traditional technologies being unable to track the dynamic propagation links of faults and the chaotic tracing of chain faults.
[0010] Furthermore, the dynamic fault tracing and analysis module includes a three-level fault quantification and attribution unit. This unit constructs a three-level fault hierarchy system comprising surface phenomena, mid-level logic, and underlying root causes. Through weighted scoring, correlation scoring, and risk value scoring, it distinguishes between primary and secondary faults and locates the core root causes. This enables quantitative and standardized fault tracing, eliminating subjective errors inherent in manual tracing, accurately distinguishing between surface fault phenomena and underlying core fault causes, providing a quantitative basis for developing differentiated operation and maintenance strategies, and avoiding superficial but ineffective maintenance solutions.
[0011] Furthermore, the adaptive closed-loop operation and maintenance decision module includes a hierarchical adaptive operation and maintenance unit; the hierarchical adaptive operation and maintenance unit matches corresponding predictive maintenance, immediate repair, and emergency isolation and handling operation and maintenance strategies according to the fault quantification risk level and fault level; it realizes the precise adaptation of operation and maintenance strategies to fault scenarios, replaces the traditional fixed and rigid operation and maintenance mode, takes into account the continuity of industrial production and the safety of equipment operation, and improves the pertinence and effectiveness of fault operation and maintenance.
[0012] Furthermore, the adaptive closed-loop operation and maintenance decision-making module includes a reverse iterative optimization unit; the reverse iterative optimization unit collects operation and maintenance process data and fault recurrence data, iteratively updates environmental interference feature library parameters, fault propagation map weights and small sample comparative learning model feature parameters; constructs a system autonomous evolution closed loop, allowing the system adaptability and accuracy to be continuously optimized over time, solving the problems of solidified adaptability and repeated recurrence of similar faults in traditional operation and maintenance systems, and significantly reducing the long-term equipment failure rate and operation and maintenance costs.
[0013] An intelligent operation and maintenance source tracing and monitoring method for industrial computer faults, applied to an intelligent operation and maintenance source tracing and monitoring system for industrial computers, includes the following steps: S1. Layered collection of multi-source operational data and industrial site environmental data from industrial computers, stripping away environmental interference data, and purifying the actual fault operation data; S12. Synchronous collection of industrial computer operational data at the hardware, system, and business layers, as well as site operating condition data at the environmental layer; S13. Based on an environmental interference feature library and residual comparison algorithm, elimination of false abnormal data caused by environmental interference, generating a structured and clean dataset; S2. Completion of intelligent fault identification and dynamic source tracing based on the clean dataset, quantifying and locating the core root cause of the fault; S24. Backward mining of data anomaly features through a small-sample comparative learning model to identify explicit and implicit faults; S25. Dynamic construction of a fault propagation map to reconstruct the fault. The time-series propagation path and node association relationship; S23 completes the fault hierarchy classification and core root cause location through a three-level fault quantification attribution system, and outputs quantitative source tracing results; S3 adaptively matches hierarchical operation and maintenance strategies based on quantitative source tracing results and executes corresponding fault operation and maintenance handling operations; S4 accumulates operation and maintenance process data, reverse iteratively optimizes data purification rules and fault source tracing models, and realizes closed-loop intelligent operation and maintenance; through hierarchical steps and processes of layered purification, small sample source tracing, hierarchical operation and maintenance, and closed-loop iteration, the entire process of industrial computer faults from data collection, fault identification, root cause tracing to intelligent operation and maintenance and system optimization is automated and controlled. Each step is deeply coordinated and linked, forming a full-dimensional fault self-healing capability that cannot be achieved by a single step, and significantly improving the intelligence, accuracy and autonomy of industrial control computer fault operation and maintenance in complex industrial scenarios.
[0014] Furthermore, S1 specifically includes: S13 performing hierarchical labeling on the collected multi-source heterogeneous data, and classifying and archiving the data according to hardware anomalies, system anomalies, business anomalies, and environmental anomalies; S14 filtering effective fault feature data based on hierarchical classification data to provide structured data support for fault tracing analysis; through hierarchical labeling and classification archiving of data, orderly management of multi-source heterogeneous data is achieved, accurately distinguishing fault-induced data at different levels, further improving the accuracy of subsequent fault identification and tracing, and avoiding analytical bias caused by data chaos.
[0015] Furthermore, S2 specifically includes: S24 updating the node association weights of the fault propagation graph in real time and dynamically correcting the fault propagation path; S25 determining the fault risk level and the scope of fault impact based on the fault quantification score results and generating a standardized source tracing report; by dynamically correcting the fault propagation path and quantifying the fault risk, the dynamic, accurate, and standardized output of the fault source tracing results is achieved, providing accurate and standardized decision-making basis for the subsequent matching of adaptive operation and maintenance strategies, and ensuring the scientific and targeted nature of operation and maintenance.
[0016] Compared with the prior art, the beneficial effects of the present invention are: 1. This solution adopts a four-layer multi-source data purification combined with a small-sample comparative learning collaborative architecture, which is different from the conventional single-dimensional collection and massive sample training mode. It removes false anomalies caused by the industrial environment from the source, and relies on the normal data of equipment to reverse identify various hidden faults. It overcomes the industry's inherent dependence on fault labeling samples, greatly expands the coverage of fault identification, and fundamentally reduces the situation of false alarms and missed alarms in monitoring.
[0017] 2. This solution combines a dynamic fault propagation map with a three-level quantitative attribution system, breaking through the limitations of qualitative judgment based on static maps. It restores the fault propagation path layer by layer in real time, distinguishes between the fault manifestations and the underlying root causes through multi-dimensional scoring, solves the problem of difficulty in identifying the primary and secondary faults in cascading concurrent faults, provides standardized quantitative basis for operation and maintenance, and avoids handling only at the fault surface level.
[0018] 3. This solution sets up a source-driven hierarchical operation and maintenance system and a reverse iteration closed loop to enable the fault analysis results to be directly adapted to differentiated handling methods. At the same time, it accumulates full-process data to continuously optimize the entire system rules, forming an autonomous evolutionary operation logic. This solves the industry shortcomings of fragmented monitoring, source tracing and operation and maintenance links and continuous degradation of system performance with operation, and completes a comprehensive upgrade of industrial control faults from passive handling to proactive prediction and autonomous optimization. Attached Figure Description
[0019] Figure 1 A schematic diagram of an intelligent operation and maintenance traceability monitoring system for industrial computer faults. Figure 2 A flowchart for an intelligent operation and maintenance traceability and monitoring method for industrial computer faults. Detailed Implementation
[0020] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0021] Please see Figures 1 to 2This invention provides a technical solution: an intelligent operation and maintenance tracing and monitoring system and method for industrial computer faults, adapted to the 24 / 7 uninterrupted operation scenario of industrial computer clusters in intelligent manufacturing continuous production lines. Such industrial scenarios involve complex and harsh operating environments, with persistent electromagnetic interference, diurnal temperature and humidity fluctuations, equipment mechanical vibration, and dust accumulation. Industrial computers, as core control equipment, operate continuously year-round, and equipment faults often exhibit sporadic, latent, and cascading characteristics. The industry generally suffers from a scarcity of fault labeling samples. Traditional monitoring and maintenance technologies are simplistic and prone to false alarms and missed alarms, unclear fault location, poor adaptability of maintenance strategies, and repeated recurrence of similar faults, failing to meet the high-precision operation and maintenance requirements of modern intelligent manufacturing.
[0022] Currently, publicly available industrial computer operation and maintenance technologies generally employ a traditional technical system based on fixed-threshold single-dimensional data monitoring, static fault rule matching, supervised training with massive manually labeled samples, post-event manual tracing and analysis, and passive operation and maintenance with fixed strategies. Existing publicly available technologies generally suffer from inherent technical defects such as weak data resistance to environmental interference, poor adaptability to complex industrial conditions, low fault identification accuracy in small-sample scenarios, lack of quantitative judgment standards for cascading fault tracing, disconnect between operation and maintenance execution and tracing analysis, and lack of autonomous iterative optimization capabilities. This invention, relying on an integrated innovative technical architecture of four-layer hierarchical data purification, small-sample reverse feature mining, dynamic fault map modeling, three-level quantitative attribution tracing, adaptive hierarchical operation and maintenance, and closed-loop iterative optimization, systematically overcomes the long-standing industry pain points of existing technologies.
[0023] Intelligent Operation and Maintenance Traceability Monitoring System for Industrial Computer Faults: Existing publicly available industrial computer operation and maintenance technologies mostly adopt fragmented and independent architectures with single monitoring modules or single traceability modules. Each functional module operates independently, data cannot be shared, and functions cannot be linked or coupled, failing to form a complete closed-loop management system for fault monitoring, traceability, operation and maintenance, and optimization. Traditional fixed technical architectures have limited functionality and poor linkage, unable to adapt to the full-process operation and maintenance needs of multi-interference, multi-level, and multi-type cascading faults under complex industrial operating conditions. Overall, the system has low intelligence, high reliance on manual intervention, and poor overall fault operation and maintenance effectiveness. This system relies on the layered coupling operation mechanism of industrial computer hardware and software, multi-source heterogeneous data fusion theory, dynamic network correlation modeling technology, adaptive decision control theory, and big data closed-loop iterative optimization mechanism. Specific technical means are as follows: The overall architecture of this system comprises a multi-source hierarchical data acquisition and purification module, a dynamic fault tracing and analysis module, an adaptive closed-loop operation and maintenance decision-making module, and a visualization and early warning display module. These four modules form a complete closed-loop operation chain from top to bottom: data input, intelligent analysis, decision execution, visualization output, and autonomous iteration. The multi-source hierarchical data acquisition and purification module establishes a four-layer synchronous acquisition architecture encompassing hardware, system, business, and environmental layers. It constructs a dedicated standardized environmental interference feature library and uses residual comparison algorithms to accurately extract abnormal data corresponding to environmental interference in industrial settings, purifying the actual operational fault data of the industrial computer itself, providing a high-quality, high-accuracy data foundation for backend intelligent analysis. The dynamic fault tracing and analysis module builds a small-sample comparative learning model based on samples of normal equipment operation, achieving accurate mining of latent fault features without the support of massive fault-labeled samples. It synchronously and dynamically constructs a real-time updated fault propagation map, and combines it with a three-level fault hierarchy attribution system to complete the quantitative location of fault root causes, outputting standardized tracing results. The adaptive closed-loop operation and maintenance decision-making module binds the results of fault quantification and tracing to adaptively match hierarchical operation and maintenance strategies, automatically executing fault handling operations of the corresponding level. It simultaneously accumulates closed-loop operation and maintenance data throughout the entire process, and iteratively optimizes the front-end data purification rules and core parameters of the tracing model, enabling the system to autonomously upgrade and iterate. The visual early warning display module dynamically displays fault tracing results, fault propagation paths, and real-time operation and maintenance execution status in real time, continuously outputting fault risk classification early warning information, achieving visualized management and control of the entire operation and maintenance status.
[0024] Example: An intelligent manufacturing production line for automotive parts is equipped with multiple industrial computers, fully responsible for the core tasks of the entire production line process, including real-time control of equipment, bidirectional transmission of production data, and workshop production scheduling. All industrial computers maintain uninterrupted operation year-round. The welding equipment in the production workshop continuously generates high-frequency electromagnetic interference; the start-up and shutdown processes of the production line cause periodic equipment vibrations; alternating day and night temperature differences cause significant fluctuations in workshop temperature and humidity; long-term continuous production leads to dust accumulation on equipment surfaces and internally. The superposition of multiple complex operating conditions causes frequent false abnormal data fluctuations in the industrial computers, accompanied by various hidden cascading failures such as hardware aging, system process anomalies, and business data delays. Traditional operation and maintenance systems can only identify explicit faults, frequently outputting invalid alarms, unable to capture hidden fault risks, unable to accurately locate the core root cause of cascading failures, and their fixed and singular operation and maintenance solutions cannot adapt to different fault scenarios, making it easy for faults to recur after temporary handling. This system, through the deep collaborative operation of four core modules, purifies environmental interference noise from the data source, accurately identifies faults across all dimensions, quantifies and locates the core root cause of faults, matches adaptive operation and maintenance strategies, and continuously iterates and optimizes, achieving unmanned, precise, and intelligent operation and maintenance of industrial computers throughout the entire process.
[0025] The unique technical approach of this solution lies in its distinct approach compared to the fragmented and independent functional architectures disclosed in existing literature. This solution adopts an integrated and interconnected architecture with four layers of data purification, small-sample dynamic tracing, adaptive closed-loop operation and maintenance, and autonomous iterative optimization. The modules are deeply coupled, mutually empowering, and synergistically enhancing each other, forming an autonomous evolutionary fault management closed loop that cannot be achieved by existing traditional technologies. It breaks through the performance bottlenecks of traditional operation and maintenance technologies from all dimensions of data processing, fault analysis, decision execution, and system optimization, significantly improving the accuracy, stability, and intelligence level of industrial control computer fault operation and maintenance under complex industrial conditions.
[0026] Multi-source hierarchical data acquisition and purification module: Existing publicly available technologies generally adopt a single-dimensional data acquisition mode, only selectively collecting basic hardware operating parameters of industrial control computers (ICS). This fails to cover multi-dimensional core data such as system operation, business interactions, and the on-site environment, easily leading to severe data silos and failing to fully reflect the true operating status of the ICS. Furthermore, existing technologies lack a systematic and standardized environmental interference removal mechanism, making it impossible to accurately distinguish between false anomalies caused by environmental interference and genuine equipment faults. The collected data has high heterogeneity, with a very low proportion of effective fault data, directly causing serious deviations in subsequent fault identification and tracing analysis, resulting in frequent misjudgments and missed diagnoses. This module addresses the issue based on the hierarchical operating characteristics of ICS, multi-source heterogeneous data synchronous acquisition technology, residual data comparative analysis theory, and industrial environmental interference feature matching mechanisms. Specific technical methods are as follows: The multi-source, layered data acquisition and purification module includes a hardware data acquisition unit, a system data acquisition unit, a business data acquisition unit, an environmental data acquisition unit, and a data interference removal and purification unit. The hardware data acquisition unit continuously collects real-time operating status data of the industrial computer's motherboard, network card, hard drive, and power supply, comprehensively covering the dynamic operating status of the industrial control computer's core hardware. The system data acquisition unit collects the industrial computer's process running status, port interaction data, driver matching status, and system operation log data throughout the entire process, completely capturing various operational anomalies at the system level. The business data acquisition unit collects real-time industrial business interaction data, data transmission status, and instruction execution feedback data from the industrial computer, fully covering the upper-level production business operating conditions. The environmental data acquisition unit collects continuous dynamic operating status data of electromagnetic intensity, temperature and humidity, vibration amplitude, and dust concentration in the industrial field in real time. The data interference removal and purification unit uniformly aggregates the raw heterogeneous data collected by the four units, calls a pre-built standardized environmental interference feature library, and compares the raw data with the standard environmental interference feature data frame by frame using a residual comparison algorithm, accurately eliminating all false abnormal data caused by environmental factors, and finally outputting a layered, labeled, and well-categorized structured clean fault dataset.
[0027] Example: During the continuous operation of a smart manufacturing production line, high-frequency welding equipment in the workshop generates high-intensity electromagnetic radiation. Workshop ventilation control and alternating day and night temperatures cause large-scale temperature and humidity fluctuations. Long-term operation of the production line equipment results in continuous mechanical vibration and dust accumulation. Traditional monitoring systems only collect basic data such as hardware temperature and voltage, directly identifying minor voltage fluctuations caused by electromagnetic interference and hardware temperature deviations caused by temperature and humidity changes as equipment malfunctions, frequently triggering invalid alarms and interfering with normal operation and maintenance. This module, through the collaborative operation of five units, synchronously collects data from all dimensions of hardware, system, business, and environment. It accurately identifies voltage and temperature fluctuations as false anomalies caused by environmental interference, automatically eliminates invalid noise data, and retains only the actual equipment malfunction data such as motherboard aging, hard drive read / write anomalies, process lag, and business instruction delays, providing clean and effective data support for subsequent high-precision fault analysis.
[0028] The unique technical approach of this solution lies in its departure from the extensive data acquisition and processing mode of existing publicly available technologies that do not distinguish between dimensions and levels. This solution adopts an integrated technical architecture of five-unit hierarchical specialized acquisition and precise purification, achieving full coverage and synchronous acquisition of industrial control computer operation data and on-site working condition data. Through a dedicated residual comparison and purification mechanism, environmental interference and noise data are completely removed from the data source, solving the industry pain point of high noise and low effective data ratio in fault data under complex industrial working conditions, and eliminating the problems of false alarms and missed alarms in fault monitoring from the root.
[0029] Small-sample contrastive learning recognition unit: Existing publicly available intelligent fault identification technologies all rely on massive amounts of manually labeled fault samples for supervised model training, resulting in a serious flaw of fault sample dependency. Industrial computer faults are low-probability, sporadic events, and the industry has long faced the problem of scarce labeled fault samples and uneven distribution of fault types. Traditional intelligent recognition models are prone to training failure and a significant drop in recognition accuracy under small-sample operating conditions. They can only identify explicit, high-identity faults and cannot accurately capture gradual, implicit faults such as gradual shifts in hardware parameters and hidden system stutters, resulting in a large number of fault identification blind spots. Based on self-supervised contrastive learning theory, normal sample feature modeling mechanism, and abnormal feature reverse mining technology, this unit is adapted to industrial small-sample fault identification application scenarios. The specific technical means are as follows: The small-sample contrastive learning recognition unit completely abandons the traditional fault sample supervised training mode. It relies on the massive amount of normal, routine samples generated by the long-term stable operation of industrial computers to complete the autonomous training of the model, and builds a fixed and accurate model of the distribution characteristics of normal equipment data. During the model operation, it continuously receives the structured and clean dataset after front-end purification, and performs a comprehensive and accurate comparison between the real-time running data characteristics and the preset normal data distribution characteristics. It reverse-engineers all abnormal data characteristics that deviate from the normal operation rules, and simultaneously distinguishes between instantaneous and sudden explicit faults and long-term gradual implicit faults, achieving accurate identification of all types of faults without discrimination. No new fault-labeled samples are needed to participate in model training and iteration throughout the entire process.
[0030] Example: Industrial control computers (ICCs) on intelligent manufacturing production lines operate stably and continuously year-round, with very few failures. The number of labeled fault samples available for model training is far less than the training threshold for traditional models. Traditional intelligent identification models cannot complete effective training iterations and can only identify obvious faults such as system errors and equipment crashes, failing to recognize potential faults such as gradual shifts in hardware parameters, hidden process lags, and minor business delays. This unit utilizes the massive amount of normal operating data accumulated from the daily operation of ICCs to train the model. Through reverse comparison logic of normal features, it accurately captures various subtle abnormal changes, identifying hidden faults in advance. This achieves high-precision, multi-dimensional intelligent fault identification even under harsh operating conditions with insufficient labeled fault samples.
[0031] The unique technical approach of this solution lies in overcoming the inherent technical bias of relying on massive amounts of fault-labeled samples for intelligent fault identification, which has long existed in the industry. It adopts a brand-new technical logic of self-supervised training with normal samples and reverse anomaly identification, which solves the problem of intelligent identification failure caused by the scarcity of fault samples in industrial control computers. It fills the technical gap in intelligent identification of hidden faults in small sample scenarios and greatly expands the coverage and accuracy of fault identification.
[0032] Dynamic Fault Propagation Graph Construction Unit: Existing publicly available fault tracing technologies generally use static fault knowledge graphs to complete fault matching analysis. The overall structure, node relationships, and weight parameters of the graph are all fixed at the factory, which cannot adapt to the complex operating characteristics of dynamic fault propagation and multi-level cascading in industrial computers. Traditional static graphs cannot capture the complete process of fault temporal diffusion and gradual evolution. When facing multi-level cascading faults, they are prone to problems such as chaotic tracing paths and distorted node relationships, failing to fully reconstruct the entire process of fault formation and propagation, resulting in extremely poor authenticity and accuracy of tracing results. This unit is based on time-series data association analysis, complex dynamic network modeling technology, fault chain propagation mechanisms, and real-time network dynamic update mechanisms. Specific technical means are as follows: The dynamic fault propagation graph construction unit is based on a three-layer architecture of industrial control computer hardware, system, and business nodes, building a multi-level interconnected fault association network. The unit connects in real-time to the cleaned time-series dataset from the front end, continuously associating the dynamic change characteristics of data at each layer. Based on the time series of abnormal data fluctuations and the strength of correlations, it dynamically updates the weight parameters of each node in the fault network and the relationships between nodes, correcting the fault propagation path in real time. The unit synchronously tracks the entire diffusion process of a fault from the underlying hardware to the middle-layer system and upper-layer business, iteratively updating the graph network structure in real time to ensure that the fault evolution status displayed in the graph is completely synchronized with the actual dynamic evolution of the fault on the industrial control computer.
[0033] Example: A surface-level fault occurs in the industrial control computer of a smart manufacturing production line, resulting in the failure of data transmission for end-user business processes. Traditional static knowledge graphs can only match single fault types and cannot trace the triggering factors at the upstream hardware and system levels. Maintenance personnel can only address specific business modules, failing to completely eradicate the fault. This unit, through real-time time-series analysis of dynamic knowledge graphs, accurately captures the complete time-series link of the fault's cascading propagation. It fully presents the multi-level chain propagation process where small fluctuations in power supply voltage cause instability in the motherboard's operating conditions, leading to driver adaptation anomalies, background process lag, and ultimately, business transmission failure. It completely reconstructs the entire dynamic evolution of the fault, providing real and effective dynamic data support for accurately locating the core root cause.
[0034] The unique technical approach of this solution lies in its use of a time-series dynamic, real-time updated fault map construction logic, which differs from the static fault map schemes with fixed structures and parameters in existing technologies. This solution can accurately adapt to the dynamic propagation and gradual evolution of cascading faults in industrial control computers, fully restore the time-series path of multi-level faults spreading layer by layer, and solve the core technical defects of traditional static maps that cannot track dynamic faults and whose cascading fault source tracing is distorted.
[0035] Three-level fault quantitative attribution unit: Existing publicly available fault tracing technologies all adopt qualitative and subjective judgment modes, which can only simply describe the fault symptoms and fault types, and lack a standardized fault hierarchy classification system and quantitative judgment standards. When faced with multi-node concurrent faults and multi-level cascading fault scenarios, it is impossible to effectively distinguish between surface fault phenomena and underlying core fault roots. The determination of fault priority relies entirely on manual operation and maintenance experience, which is highly subjective and lacks consistent judgment standards. This easily leads to the industry problem of operation and maintenance only addressing surface faults while failing to thoroughly eliminate core root causes, ultimately resulting in recurring faults. Based on the fault hierarchy transmission mechanism, multi-dimensional risk scoring model, quantitative attribution judgment theory, and a mechanism for accurately distinguishing fault priority, the specific technical means are as follows: The three-level fault quantification and attribution unit constructs a standardized three-level fault hierarchy system: surface phenomena, mid-level logic, and underlying root causes. This system uniformly categorizes all fault types in industrial computers and assigns them to each level. The unit sets three independent quantitative dimensions: hierarchical weight score, node correlation score, and fault risk value score. For each fault event, it automatically performs multi-dimensional parameter calculations and comprehensive analysis. Through the comprehensive scoring results, it accurately distinguishes the primary and secondary relationships of faults, precisely locates the underlying core root causes in multi-level fault scenarios, and quantifies the severity, scope of impact, and risk of spread of the fault, forming standardized, quantifiable, and traceable source determination results.
[0036] Example: In a smart manufacturing production line, the industrial control computer simultaneously exhibits three concurrent anomalies: network card data latency, system process lag, and business instruction execution failure. Traditional tracing methods cannot distinguish the order and priority of these faults, and blindly addressing superficial business faults fails to solve the problem at its root. This unit utilizes a three-level quantitative attribution system to accurately determine that business instruction failure is a superficial phenomenon, process lag is a mid-level logical anomaly, and unstable motherboard power supply is the underlying core root cause. Quantitative scoring identifies underlying hardware faults as the core triggering factors, guiding maintenance work to prioritize addressing the core fault root cause, completely eliminating the source of the fault, and preventing recurrence.
[0037] The unique technical approach of this solution lies in breaking through the existing qualitative and subjective fault tracing and judgment model. It adopts a standardized attribution system with three-level hierarchical division and multi-dimensional quantitative scoring to achieve digital and standardized judgment of fault primary and secondary distinction and root cause location. This solves the industry pain points of chaotic fault tracing and concurrent fault tracing and large subjective bias, and provides a reliable quantitative basis for accurate hierarchical operation and maintenance.
[0038] Hierarchical Adaptive Operation and Maintenance Unit: Existing publicly available intelligent operation and maintenance technologies employ a rigid handling mode that matches fixed fault levels with fixed operation and maintenance strategies, lacking scenario-adaptive capabilities. Excessive operation and maintenance for low-risk, latent faults can disrupt normal production processes, causing production losses; delayed operation and maintenance of high-risk cascading faults can lead to fault propagation and production line downtime, failing to simultaneously ensure industrial production continuity and equipment operational safety, resulting in extremely low targeting and effectiveness of operation and maintenance work. This unit addresses the need for a fault risk classification and control theory, a scenario-adaptive decision-making mechanism, and a precise matching logic between fault levels and operation and maintenance strategies. Specific technical methods are as follows: The hierarchical adaptive operation and maintenance unit is directly linked to the front-end quantitative traceability results. Based on fault level, quantitative risk level, and fault propagation scope as core judgment criteria, three differentiated adaptive operation and maintenance strategies are constructed. For low-risk latent faults with clear underlying causes and no risk of propagation, a predictive maintenance strategy is implemented, completing preventative operation and maintenance operations such as equipment calibration, parameter optimization, and hardware testing during production line idle windows. For medium-risk explicit faults with mid-level logical anomalies, an immediate repair strategy is implemented, completing precise repair operations such as process restart, driver adaptation, and parameter reset in real time to ensure continuous and stable operation of the production line. For high-risk faults with hardware damage and multi-level cascading anomalies, an emergency isolation and disposal strategy is implemented, quickly isolating faulty industrial control computer nodes, switching to backup equipment, blocking the fault propagation chain, and avoiding the risk of large-scale production downtime.
[0039] Example: In a smart manufacturing production line, the industrial control computer detects a low-risk, latent fault—a slight, gradual shift in hard drive read / write parameters. The system automatically matches a predictive maintenance strategy, completing hard drive testing, defragmentation, and parameter calibration during shift handover downtime to prevent the fault from worsening. If the system detects a medium-risk fault—process lag caused by driver incompatibility—it immediately performs driver restart and parameter adaptation repair, ensuring continuous production without interruption. If the system detects a high-risk, multi-level cascading fault caused by abnormal motherboard power supply, it immediately isolates the faulty industrial control computer and switches to backup equipment, ensuring stable operation of the entire production line.
[0040] The unique technical approach of this solution lies in its ability to differentiate itself from the fixed, rigid, and undifferentiated unified operation and maintenance model of existing technologies. This solution implements an adaptive hierarchical operation and maintenance mechanism driven by traceability results, which allows operation and maintenance strategies to be precisely adapted to fault levels, risk levels, and scenario characteristics. It takes into account both the continuity of industrial production and the safety of equipment operation, and completely solves the technical defects of traditional operation and maintenance, such as blindness, poor adaptability, and inaccurate handling.
[0041] Reverse Iterative Optimization Unit: Existing publicly available intelligent operation and maintenance monitoring systems all employ static, fixed architectures. Model parameters, feature library rules, and judgment criteria are permanently fixed after equipment deployment, lacking data accumulation and autonomous optimization iteration capabilities. As industrial equipment ages over time and on-site conditions dynamically change, the original adaptation rules gradually become incompatible with the conditions, leading to a continuous decline in model recognition and tracing accuracy, recurring similar faults, and a sustained deterioration in the system's long-term adaptability and stability. This unit addresses this issue by employing big data closed-loop iterative theory, online model optimization mechanisms, adaptive operating condition update technology, and reverse empowerment logic based on operation and maintenance data. Specific technical methods are as follows: The reverse iterative optimization unit collects fault identification results, source analysis data, operation and maintenance records, fault repair status, and fault recurrence statistics throughout the entire process, constructing a complete operation and maintenance closed-loop database. Based on the accumulated massive closed-loop operation data, the unit periodically iteratively updates the matching parameters and judgment thresholds of the environmental interference feature library, dynamically corrects the node weights and correlations of the fault propagation graph, simultaneously optimizes the abnormal feature identification parameters and comparison standards of the small-sample comparative learning model, and continuously optimizes data purification rules and fault judgment logic, achieving dynamic upgrades to the overall system performance and autonomous adaptation to operating conditions.
[0042] Example: After long-term operation of a smart manufacturing production line, subtle dynamic changes occur in the patterns of dust accumulation, temperature and humidity fluctuations, and electromagnetic interference intensity in the workshop. This leads to a decrease in the accuracy of the original environmental interference feature database, resulting in a small number of misjudgments. This unit autonomously updates the feature parameters corresponding to dust, temperature, humidity, and electromagnetic interference through long-term accumulated operation and maintenance data and fault data, optimizing the residual comparison algorithm's judgment threshold to accurately adapt to new on-site conditions. Simultaneously, for recurring similar latent gradual faults, the unit optimizes the small-sample model feature recognition system, significantly improving the accuracy of identifying similar faults and continuously reducing the probability of fault recurrence.
[0043] The unique technical approach of this solution lies in its construction of a complete operational data reverse iterative optimization closed loop, which differs from existing static systems with fixed architectures, parameters, and no autonomous optimization capabilities. This enables continuous autonomous upgrades of system rules, model parameters, and feature libraries, allowing the system's adaptability and accuracy to continuously improve over time. This completely solves the long-standing industry pain point of traditional systems where accuracy decreases with use and similar faults recur repeatedly.
[0044] Intelligent Operation and Maintenance Traceability Monitoring Method for Industrial Computer Faults: Existing publicly available monitoring and maintenance methods consist of fragmented and independent steps. Data collection, fault analysis, source tracing, and maintenance handling are isolated and lack coordination, failing to form a complete closed-loop workflow. Traditional methods involve numerous manual interventions, have low automation levels, and poor step-by-step connections, making integrated intelligent operation and maintenance impossible. Overall operational efficiency and fault handling accuracy are extremely low, failing to meet the demands of modern industrial intelligent operation and maintenance. This new method is based on a hierarchical data purification process, small-sample intelligent identification logic, dynamic source tracing mechanism, adaptive operation and maintenance decision rules, and a closed-loop iterative optimization system.
[0045] S1 collects multi-source operational data from industrial computers and industrial site environmental data in a layered manner, removing environmental interference data and refining the actual fault operation data. Through a four-layer synchronous parallel acquisition architecture, it comprehensively collects all dimensions of industrial control computer operational data and field operating condition data. Relying on a standardized environmental interference feature library and residual comparison algorithm, it eliminates all false anomaly data caused by environmental interference, refines the actual fault data of the equipment itself, and generates a structured, clean dataset, providing a standardized data foundation for subsequent intelligent fault analysis.
[0046] The S11 synchronously collects industrial computer operation data from the hardware, system, and business layers, as well as on-site operating data from the environmental layer, ensuring that multi-source heterogeneous data is synchronized in time, complete in dimensions, and without missing or delayed data, achieving accurate data collection with full coverage across all scenarios.
[0047] Based on an environmental interference feature library and a residual comparison algorithm, S12 compares the original collected data with standard environmental interference feature parameters frame by frame, accurately identifies and removes false abnormal data caused by environmental interference, retains the true fault data of the equipment itself, and completes the data standardization and purification process.
[0048] S2 performs intelligent fault identification and dynamic source tracing based on clean datasets, quantitatively locating the core root cause of faults. Relying on purified, high-quality structured data, it completes intelligent identification of all types of faults through a small-sample comparative learning model. Combined with a dynamic fault propagation map, it reconstructs the dynamic diffusion path of faults and accurately locates the underlying core fault root cause through a three-level quantitative attribution system, outputting standardized quantitative source tracing results.
[0049] S21 uses a small-sample comparative learning model to reverse-engineer abnormal data features. Based on the normal distribution patterns of equipment data, it accurately identifies explicit sudden faults and implicit gradual faults that deviate from normal operating conditions, achieving high-precision fault identification without the support of a massive number of fault samples.
[0050] S22 dynamically constructs a fault propagation map, dynamically updates the node association weights and relationships based on real-time data time-series changes, and fully restores the time-series propagation path and node linkage relationships of faults from the underlying hardware to the upper-layer business.
[0051] S23 uses a three-level fault quantitative attribution system to classify faults into surface phenomena, mid-level logic, and underlying root causes. Through multi-dimensional quantitative scoring, it determines the primary and secondary faults and locates the core root causes, outputting standardized quantitative tracing results that include fault level, risk level, and scope of impact.
[0052] S3 adaptively matches hierarchical operation and maintenance strategies based on quantitative traceability results and executes corresponding fault operation and maintenance operations. The system reads standardized quantitative traceability reports and automatically matches predictive maintenance, immediate repair, and emergency isolation handling strategies according to fault level, risk level, and scope of spread, completing fault operation and maintenance work in a fully automated manner.
[0053] S4 accumulates data from the entire operation and maintenance process, iteratively optimizing data purification rules and fault tracing models to achieve closed-loop intelligent operation and maintenance. The system retains all data from the entire process of data collection, fault identification, tracing analysis, operation and maintenance handling, and fault recurrence. Based on the accumulated data, it continuously iterates and optimizes data purification algorithms, environmental interference feature libraries, fault spectrum weights, and small sample model parameters, achieving autonomous upgrades and optimizations of the methodology.
[0054] Example: During the daily operation of the industrial control computer on a smart manufacturing production line, the system automatically executes a closed-loop operation and maintenance method. The first step involves synchronously collecting four layers of multi-dimensional operational and environmental data, accurately eliminating false anomalies caused by electromagnetic interference, temperature, humidity, and vibration in the workshop, and generating a standardized, clean fault dataset. The second step uses a small-sample model to identify latent faults caused by the long-term gradual aging of hard drives, reconstructs the temporal path of the fault's gradual spread through dynamic graphs, and determines hard drive aging as the underlying core cause of the fault through three-level quantitative attribution. The third step involves the system matching a low-risk predictive maintenance strategy, completing hard drive inspection and replacement operations during production downtime. The fourth step involves accumulating the fault data and operation and maintenance data, autonomously optimizing the latent fault identification features and data purification thresholds to improve the accuracy and efficiency of identifying and handling similar faults in the future.
[0055] The unique technical approach of this solution lies in its distinct approach compared to existing publicly available methods that are fragmented, with isolated steps, independent functions, and no iterative optimization. This solution constructs a comprehensive, multi-level, interconnected, and integrated methodology system encompassing data purification, intelligent tracing, adaptive operation and maintenance, and closed-loop iteration. Each step deeply collaborates and empowers the others, resulting in a full-dimensional fault self-healing effect that cannot be achieved by a single step operating independently. This enables a comprehensive upgrade of industrial control computer fault operation and maintenance from manual, passive handling to fully automated, intelligent prediction and control, and autonomous optimization.
[0056] Data stratification, labeling, and archiving refinement: Existing publicly available data processing methods only perform simple format conversion and storage of raw data, lacking a stratified, categorized, labeled, and archived mechanism. This results in a jumbled accumulation of multi-source, heterogeneous data, making it impossible to accurately distinguish the corresponding equipment fault levels. Using this disorganized raw data directly for fault analysis leads to chaotic fault level matching, significant errors in source tracing and location, and extremely low effective data utilization, severely impacting the accuracy of fault analysis and source determination. This is addressed by employing multi-source, heterogeneous data stratified and categorized management technology, data level-fault level correspondence matching rules, and a structured data preprocessing mechanism.
[0057] S13 performs hierarchical labeling on the collected multi-source heterogeneous data, and completes accurate classification and labeling according to four fixed labels: hardware anomaly, system anomaly, business anomaly, and environmental anomaly. This ensures that each set of data has clear hierarchical attributes and anomaly type identifiers, achieving a one-to-one correspondence between data hierarchy and fault hierarchy.
[0058] S14 filters and removes invalid, redundant, and duplicate data from the hierarchically categorized and labeled dataset, retains the valid core data with fault characteristic associations, and archives and stores them according to hierarchical classification to form a structured and orderly fault dataset, providing standardized data support for subsequent fault identification and source tracing analysis.
[0059] Example: In a smart manufacturing production line, the industrial control computer synchronously collects four types of raw data: hardware voltage fluctuations, system process delays, business command lags, and on-site electromagnetic fluctuations. The system labels these data at the hardware, system, business, and environmental levels, and then archives and stores them accordingly. During subsequent fault analysis, the analysis module can directly retrieve the data at the corresponding level, quickly distinguishing between invalid environmental interference data and valid data indicating actual equipment faults, accurately locating the fault's level, and completely avoiding biases in the tracing analysis caused by mixed data.
[0060] The unique technical approach of this solution lies in its distinction from the indiscriminate and extensive data processing and storage model of existing technologies. This solution adopts a refined data preprocessing mechanism of hierarchical labeling and classification archiving to establish a precise correspondence between data levels and fault levels. This ensures the accuracy of subsequent fault identification and source tracing analysis from the data structure level, and significantly improves the effective utilization rate of multi-source heterogeneous data and the accuracy of fault analysis.
[0061] Dynamic Fault Map Correction and Standardized Report Generation: Existing publicly available fault tracing methods rely on fixed fault propagation paths, failing to adapt to the complex characteristics of dynamic and gradual fault spread. Fault propagation analysis results suffer from severe lag and distortion. Furthermore, current technologies lack a standardized fault tracing result output mechanism, resulting in disorganized and unfocused information. Fault risk level and impact scope determination rely entirely on human experience, leading to strong subjectivity and poor standardization, failing to provide a unified, standardized, and accurate reference for operational decisions. This solution addresses the need for dynamic network real-time update technology, multi-dimensional quantitative risk assessment rules, and an automatic standardized report generation mechanism.
[0062] S24 monitors the real-time dynamic fluctuation trend of industrial control computer fault data. Based on the real-time changes in the strength of data anomalies and the scope of correlation, it continuously updates the node correlation weights of the fault propagation map, dynamically corrects the fault propagation path and the scope of fault spread, and ensures that the fault evolution process displayed in the map is completely synchronized with the actual fault status of the equipment.
[0063] Based on the three-level fault quantification scoring results, S25 automatically determines the fault risk level and the scope of fault impact, sorts out the core root cause of the fault, dynamic propagation path, abnormal behavior, operation and maintenance optimization suggestions and other core information, and automatically generates a standardized and structured traceability report, providing an intuitive, standardized and accurate basis for operation and maintenance decision-making.
[0064] Example: In intelligent manufacturing production lines, industrial control computer (ICC) malfunctions exhibit a gradual, spreading pattern. The fault evolves from abnormal hard drive read / write parameters to system process lag, ultimately affecting business data transmission. The system monitors data changes dynamically in real time, continuously updates the weights of nodes in the fault graph, dynamically corrects the fault propagation path, and fully reconstructs the entire process of the fault's gradual spread. Simultaneously, based on multi-dimensional quantitative scoring, the fault is determined to be a medium-risk hardware-derived fault, clearly indicating that the fault only covers a single ICC and has no risk of spreading across the entire production line. A standardized source tracing report is automatically generated, including root cause identification, risk level, propagation path, and maintenance recommendations, guiding maintenance personnel to accurately and efficiently handle the fault.
[0065] The unique technical approach of this solution lies in its ability to dynamically and in real-time correct fault propagation analysis and source tracing mode without standardized output, unlike the static and fixed fault propagation analysis and source tracing mode of existing technologies. This solution enables more accurate source tracing analysis of complex chain faults and gradual faults, more standardized basis for operation and maintenance decisions, and completely solves the industry problems of subjective human judgment, inconsistent results, and insufficient targeted operation and maintenance.
[0066] In summary, current publicly available technologies in the field of industrial computer fault operation and maintenance monitoring generally suffer from four major shortcomings. First, at the data processing level, they typically employ a fixed threshold, single-dimensional monitoring model, lacking a systematic environmental interference removal mechanism, leading to significant false alarms and missed alarms under complex industrial conditions. Second, at the intelligent identification level, they heavily rely on massive amounts of labeled fault samples, exhibiting poor adaptability to small sample scenarios and failing to identify latent, gradual faults. Third, at the source tracing analysis level, they rely on static knowledge graphs and qualitative subjective judgments, lacking dynamic source tracing capabilities for cascading faults and lacking unified quantitative judgment standards. Fourth, at the operation and maintenance optimization level, monitoring, source tracing, and operation and maintenance processes are fragmented, operation and maintenance strategies are rigid and lack adaptive capabilities, and the system lacks an autonomous iterative optimization mechanism.
[0067] This technical solution is unique in that it addresses the problem of noise interference in complex operating conditions from the source through a four-layer, multi-source data purification architecture, overcoming the limitations of traditional single-dimensional monitoring. Through a small-sample identification mechanism based on normal sample comparison learning, it overcomes the long-standing industry bias of sample-dependent technology and solves the intelligent identification challenge in scenarios where industrial control computer fault samples are scarce. By employing a dynamic fault propagation map and a three-level quantitative attribution system, it achieves accurate tracing and standardized root cause localization of dynamic cascading faults, filling the gap in quantitative tracing technology. Through tracing-driven adaptive hierarchical operation and maintenance and a reverse closed-loop iterative architecture, it achieves integrated linkage across the entire process of monitoring, tracing, operation and maintenance, and optimization, constructing an intelligent operation and maintenance system that can autonomously evolve.
[0068] The core technologies in this solution work together and deeply empower each other, forming a technological synergy that cannot be achieved by existing publicly available technologies. It enables precise monitoring, quantitative tracing, adaptive operation and maintenance, and continuous optimization of industrial computer faults in harsh industrial scenarios with strong interference, small sample sizes, and frequent hidden chain failures. The overall technical solution breaks through the inherent bottlenecks of existing industrial computer fault operation and maintenance technologies, and realizes a comprehensive technological upgrade of the operation and maintenance mode from passive handling to proactive prediction, from manual experience to data quantification, and from static solidification to autonomous evolution. It has significant technological advancements and is fully adapted to the high-precision, unmanned, and adaptive intelligent operation and maintenance needs of industrial computer clusters in high-end intelligent manufacturing scenarios.
Claims
1. An intelligent operation and maintenance traceability monitoring system for industrial computers, characterized in that: It includes a multi-source hierarchical data acquisition and purification module, a dynamic fault tracing and analysis module, an adaptive closed-loop operation and maintenance decision-making module, and a visual early warning display module; The multi-source hierarchical data acquisition and purification module builds a four-layer synchronous acquisition architecture consisting of hardware, system, business, and environmental layers. It constructs a fixed environmental interference feature library and uses a residual comparison algorithm to remove abnormal data corresponding to environmental interference in the industrial field, thereby purifying the real operating data of the industrial computer itself. The dynamic fault tracing and analysis module builds a small-sample comparative learning model based on normal operation samples, mines hidden fault characteristics, dynamically constructs a fault propagation map, and matches a three-level fault hierarchy attribution system to complete the quantitative location of fault root causes. The adaptive closed-loop operation and maintenance decision module binds the fault quantification and tracing results to match the hierarchical operation and maintenance strategy, executes the corresponding fault handling operations, and simultaneously accumulates operation and maintenance data to iteratively optimize the front-end data purification rules and tracing model parameters. The visualization and early warning display module shows the results of fault tracing, fault propagation path and operation and maintenance execution status, and outputs fault risk early warning information.
2. The intelligent operation and maintenance traceability monitoring system for industrial computer faults as described in claim 1, characterized in that: The multi-source hierarchical data acquisition and purification module includes a hardware data acquisition unit, a system data acquisition unit, a business data acquisition unit, an environmental data acquisition unit, and a data interference stripping and purification unit. The hardware data acquisition unit collects real-time operating status data of the industrial computer's motherboard, network card, hard drive, and power supply; the system data acquisition unit collects process, port, driver, and system operation log data of the industrial computer; the business data acquisition unit collects industrial business interaction, data transmission, and instruction execution data of the industrial computer; the environmental data acquisition unit collects environmental data such as electromagnetic intensity, temperature, humidity, vibration, and dust concentration in the industrial site; and the data interference stripping and purification unit calls the environmental interference feature library and uses a residual comparison algorithm to remove false abnormal data caused by environmental factors, outputting a structured and clean fault dataset.
3. The intelligent operation and maintenance traceability monitoring system for industrial computer faults as described in claim 1, characterized in that: The dynamic fault tracing and analysis module includes a small sample comparison learning and identification unit. The small sample comparison learning and identification unit completes model training based on samples of normal operation of industrial computers, and reverses the identification of abnormal fault features based on the normal data distribution characteristics to complete the accurate identification of explicit and implicit faults.
4. The intelligent operation and maintenance traceability monitoring system for industrial computer faults as described in claim 1, characterized in that: The dynamic fault tracing and analysis module includes a dynamic fault propagation graph construction unit; the dynamic fault propagation graph construction unit associates the temporal change characteristics of hierarchical operation data in real time, dynamically updates the fault node weight and node association relationship, and restores the complete propagation path of the fault from the hardware layer to the business layer in real time.
5. The intelligent operation and maintenance traceability monitoring system for industrial computer faults as described in claim 1, characterized in that: The dynamic fault tracing and analysis module includes a three-level fault quantification and attribution unit; The three-level fault quantification attribution unit constructs a three-level fault hierarchy system consisting of surface phenomena, mid-level logic, and underlying root causes. It completes the differentiation of primary and secondary faults and the location of core root causes through weight scoring, correlation scoring, and risk value scoring.
6. The intelligent operation and maintenance traceability monitoring system for industrial computer faults as described in claim 1, characterized in that: The adaptive closed-loop operation and maintenance decision module includes a hierarchical adaptive operation and maintenance unit; the hierarchical adaptive operation and maintenance unit matches corresponding predictive maintenance, immediate repair, and emergency isolation and handling operation and maintenance strategies according to the fault quantification risk level and fault level.
7. The intelligent operation and maintenance traceability monitoring system for industrial computer faults as described in claim 1, characterized in that: The adaptive closed-loop operation and maintenance decision-making module includes a reverse iterative optimization unit; the reverse iterative optimization unit collects operation and maintenance process data and fault recurrence data, and iteratively updates the environmental interference feature library parameters, fault propagation map weights and small sample comparison learning model feature parameters.
8. An intelligent operation and maintenance traceability monitoring method for industrial computer faults, applied to the intelligent operation and maintenance traceability monitoring system for industrial computer faults as described in any one of claims 1-7, characterized in that: Includes the following steps: S1 collects multi-source operational data from industrial computers and industrial field environmental data in a layered manner, removes environmental interference data, and purifies the real fault operation data. S11 synchronously collects industrial computer operation data from the hardware layer, system layer, and business layer, as well as on-site operating condition data from the environmental layer. S12 uses an environmental interference feature library and a residual comparison algorithm to remove false anomalies caused by environmental interference and generate a structured, clean dataset. S2 uses a clean dataset to perform intelligent fault identification and dynamic source tracing, and quantitatively locates the core root cause of the fault. S21 uses a small-sample contrastive learning model to reverse-engineer abnormal data features and identify explicit and implicit faults. S22 dynamically constructs a fault propagation graph to reconstruct the fault time-series propagation path and node relationships; S23 completes the classification of fault levels and the location of core root causes through a three-level fault quantitative attribution system, and outputs quantitative source tracing results; S3 adaptively matches hierarchical operation and maintenance strategies based on quantitative traceability results and executes corresponding fault operation and maintenance handling operations. S4 accumulates data from the entire operation and maintenance process, and iteratively optimizes data purification rules and fault tracing models to achieve closed-loop intelligent operation and maintenance.
9. The intelligent operation and maintenance traceability monitoring method for industrial computer faults as described in claim 8, characterized in that: S1 specifically includes: S13 performs hierarchical labeling on the collected multi-source heterogeneous data and completes data classification and archiving according to hardware anomalies, system anomalies, business anomalies, and environmental anomalies. S14 filters effective fault feature data based on hierarchical classification data, providing structured data support for fault source tracing analysis.
10. The intelligent operation and maintenance traceability monitoring method for industrial computer faults as described in claim 8, characterized in that: S2 specifically includes: S24 updates the node association weights of the fault propagation graph in real time and dynamically corrects the fault propagation path; Based on the fault quantification score, S25 determines the fault risk level and the scope of fault impact, and generates a standardized source tracing report.