Industrial software cloud base intelligent operation and maintenance and fault self-healing method based on digital twinning
By constructing a hierarchical digital twin model and a fault knowledge graph, and combining real-time data with historical cases, intelligent operation and maintenance and fault self-healing of industrial software cloud infrastructure have been achieved. This solves the problems of inaccurate status monitoring and inadequate self-healing control in traditional operation and maintenance, and improves fault diagnosis efficiency and cloud infrastructure stability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- INNER MONGOLIA ACADEMY OF SCIENCE & TECHNOLOGY
- Filing Date
- 2025-12-22
- Publication Date
- 2026-05-05
AI Technical Summary
Traditional industrial software cloud infrastructure maintenance relies on manual experience, making it impossible to grasp the operational status in real time and comprehensively. It also lacks systematic knowledge accumulation and correlation analysis, resulting in inaccurate fault diagnosis and the self-healing control system is prone to causing business interruption.
A hierarchical digital twin model is constructed, which combines physical structure and real-time operation data to achieve accurate state mapping and anomaly monitoring. A fault knowledge graph is constructed for multi-dimensional analysis. A self-healing strategy library is built based on cause nodes and historical cases, and the self-healing strategy is verified through simulation using the digital twin model.
It has enabled intelligent and efficient operation and maintenance of industrial software cloud base, improved the accuracy of fault diagnosis and the reliability of self-healing, reduced labor costs, and ensured the stable operation of cloud base.
Smart Images

Figure CN121980302A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of digital twin technology, and in particular to a method for intelligent operation and maintenance and fault self-healing of industrial software cloud infrastructure based on digital twins. Background Technology
[0002] Traditional industrial software operation and maintenance methods rely on manual experience for troubleshooting, which cannot provide real-time and comprehensive understanding of the overall operational status of the cloud platform. Anomaly detection is delayed, and there is a lack of systematic knowledge accumulation and correlation analysis methods. Furthermore, there is a lack of adaptability verification for complex faults, resulting in limited self-healing success rates and a high risk of business interruption. For example, Chinese patent application CN111596604A discloses a digital twin-based intelligent fault diagnosis and self-healing control system for engineering equipment. This system includes a physical entity module, a data acquisition module, an information processing module, a fault diagnosis module, a self-healing control module, and a digital twin module. The data acquisition module collects real-time information data about the operation of the engineering equipment from the physical entity module and transmits the data to the digital twin module for digital twin simulation of the equipment. Simultaneously, after processing by the information processing module, the data undergoes intelligent diagnostic analysis in the fault diagnosis module, and the self-healing control module handles the resulting faults through self-healing control. The digital twin module interacts and provides feedback to other modules, achieving information exchange and closed-loop optimization. While this patent application can improve the accuracy of fault prediction, reduce the failure rate, lower equipment maintenance costs, and enhance the stability and robustness of equipment operation, it still faces the challenge of adapting to industrial software cloud-based maintenance scenarios. 1. Engineering equipment operation and maintenance focuses on the state simulation of physical entities, without covering the logic layer and application layer unique to industrial software cloud base, and cannot realize multi-level data linkage mapping, making it difficult to fully reflect the operation status of the complex architecture of the cloud base; 2. Fault diagnosis lacks systematic knowledge accumulation and correlation analysis capabilities. It has not introduced a fault knowledge graph and relies solely on a single diagnostic logic after data processing. When faced with the complex causal relationships of multiple types of faults in the cloud base, the accuracy and efficiency of cause location are insufficient. 3. The self-healing control lacks a strategy simulation and verification mechanism and only directly executes repair operations. When dealing with multi-level related faults in the cloud base, improper strategies can easily lead to the expansion of faults or business interruptions, which cannot meet the high stability and high complexity operation and maintenance requirements of industrial software cloud bases. Summary of the Invention
[0003] The purpose of this invention is to provide an intelligent operation and maintenance and fault self-healing method for industrial software cloud infrastructure based on digital twins. By dynamically integrating the physical structure of the cloud infrastructure with real-time operating data through a layered digital twin model, it achieves accurate mapping of operating status and real-time monitoring of anomalies. Hierarchical early warning enables faster operation and maintenance response. A self-healing strategy library is built based on cause nodes and historical cases to achieve intelligent fault self-healing. This significantly improves the intelligence, accuracy and efficiency of industrial software cloud infrastructure operation and maintenance, reduces labor costs, and ensures the stable and reliable operation of industrial software, thereby solving the problems mentioned in the background art.
[0004] To achieve the above objectives, the present invention provides the following technical solution: Intelligent operation and maintenance and fault self-healing methods for industrial software cloud infrastructure based on digital twins include: Acquire physical structure data and real-time operating status data of the industrial software cloud base, construct a hierarchical digital twin model of the cloud base, and dynamically update the hierarchical digital twin model by combining the geometric parameters, performance parameters and topological relationships of the physical base. Based on the constructed hierarchical digital twin model, the actual operating status of the industrial software cloud base is mapped in real time and simulated. A visual monitoring interface is constructed for visual display, abnormal states that deviate from the normal range are identified, and hierarchical early warning information is generated. Construct a fault knowledge graph, extract abnormal indicator data corresponding to abnormal states, perform multi-dimensional data association analysis, filter entity data associated with abnormal indicators, generate a fault analysis dataset, and locate the unique fault cause based on the fault knowledge graph and the fault analysis dataset. A self-healing strategy library is built based on the cause nodes of the fault knowledge graph and historical self-healing cases. The self-healing strategy is matched according to the located fault cause, and the self-healing strategy is input into the hierarchical digital twin model for simulation verification. Effective self-healing strategies are obtained and deployed to the cloud base for execution.
[0005] Furthermore, the process of constructing a layered digital twin model for the cloud infrastructure includes: Extract the physical structure data of the industrial software cloud base, including the physical structure parameters and hardware performance parameters of the hardware components, and combine it with the network topology data of the cloud base to construct a geometric simulation model of the cloud base and generate physical layer state data. Collect the virtualization resource configuration, network routing rules, data transmission protocols and software dependencies of the cloud base, construct the logical topology model of the cloud base, simulate the resource scheduling process, data transmission path and software interaction logic, and generate logical layer configuration data; Obtain the deployment architecture, service call chain, business process configuration, and performance index thresholds of industrial software, construct the application operation model of the cloud base, associate physical layer state data with logical layer configuration data, establish the mapping relationship between business performance indicators and the state data of geometric simulation model and the configuration data of logical topology model, and generate a layered digital twin model of the cloud base.
[0006] Furthermore, the layered digital twin model of the cloud base also includes: Construct accuracy calibration rules for the geometric simulation model, and obtain real-time hardware operating parameters from the physical layer state data based on a preset extraction time interval; The real-time operating parameters of the hardware are compared with the actual operating status of the cloud base hardware output by the geometric simulation model, and the mapping coefficients of the hardware performance parameters of the geometric simulation model are corrected based on the comparison results. Obtain the virtualization resource configuration and software dependencies of the cloud base, and build a configuration data verification library based on the dependencies; Extract the configuration data from the logical topology model, compare the configuration data with the rules of the configuration data verification library, and generate configuration correction suggestions; When the virtualized resources of the cloud base change, the resource scheduling simulation parameters and data transmission path parameters of the logical topology model are automatically updated. Based on the mapping relationship, a collaborative adjustment mechanism for performance index thresholds is established. When the hardware performance parameters of the physical layer model or the resource configuration of the logical topology model changes, the business performance index thresholds of the application layer model are recalculated through the collaborative adjustment mechanism for performance index thresholds, and the business performance monitoring rules of the application layer model are updated synchronously.
[0007] Furthermore, the comparison results are used to correct the hardware performance parameter mapping coefficients of the geometric simulation model, including: Extract the simulated values of the current hardware operating status output from the geometric simulation model; Retrieve the current hardware real-time operating parameters corresponding to the current hardware operating status simulation value, and use the current hardware real-time operating parameters and the current hardware operating status simulation value to perform absolute difference processing to obtain the current absolute error value; The error rate is obtained by comparing the absolute error value with the current real-time operating parameters of the hardware. Retrieve the preset basic correction step size α from the database, wherein the value of the basic correction step size is in the range of 0 < α ≤ 0.5; The error rate corresponding to the previous hardware performance parameter mapping coefficient correction process is retrieved as the preceding error rate; The preset basic correction step size α is adjusted using the error rate and the previous error rate to obtain the adjusted correction step size; Using the adjusted correction step size αt The mapping coefficients for hardware performance parameters are corrected.
[0008] Furthermore, using the adjusted correction step size α t Correcting the mapping coefficients of hardware performance parameters, including: Retrieve the simulated hardware operating status values and real-time hardware operating parameters output by the geometric simulation model during the previous hardware performance parameter mapping coefficient correction process; The simulated hardware operating status value and real-time hardware operating parameters output by the geometric simulation model during the previous hardware performance parameter mapping coefficient correction process are used as the reference simulated hardware operating status value and reference real-time hardware operating parameters. The error trend factor r is obtained by combining the simulated values of the reference hardware operating status and the real-time operating parameters of the reference hardware with the real-time operating parameters of the current hardware and the simulated values of the current hardware operating status. Retrieve the adjusted correction step size α t The coefficient value k0 of the hardware performance parameter mapping coefficient after the last correction; The error trend factor is combined with the adjusted correction step size α t The hardware performance parameter mapping coefficients are corrected based on the coefficient value k0 of the hardware performance parameter mapping coefficients.
[0009] Furthermore, a visual monitoring interface is constructed for visualization, including a physical layer, logical layer, and application layer architecture based on a layered digital twin model. The visual monitoring interface sets up three independent visualization modules, and the status data of the three visualization modules are updated synchronously based on the model accuracy calibration cycle.
[0010] Furthermore, identifying abnormal states that deviate from the normal range specifically includes: Based on the data characteristics of the physical, logical, and application layers of the hierarchical digital twin model, and combined with historical operational data and industry standards, a dynamic benchmark library for each layer is established. By associating the dynamic benchmark libraries of each layer with the physical layer hardware performance parameters, logical layer configuration data, and application layer business performance index thresholds of the hierarchical digital twin model, the benchmark values of each layer's dynamic benchmark library are calibrated in real time.
[0011] Furthermore, identifying abnormal states that deviate from the normal range also includes: Obtain the physical layer status data of the hierarchical digital twin model, compare the physical layer status data with the benchmark value of the corresponding physical layer dynamic benchmark library, and if the data exceeds the benchmark value for multiple consecutive acquisition cycles, it is marked as an abnormal physical layer parameter. The physical layer geometric simulation model based on the hierarchical digital twin model monitors the target operating status of cloud base hardware devices, including device offline, interface connection interruption or physical location offset. If the target operating status is detected, it is marked as an abnormal physical layer device status. Obtain the logical layer resource usage data and transmission link data of the hierarchical digital twin model, and compare them with the benchmark values of the corresponding logical layer dynamic benchmark library. If the resource usage data, transmission delay / packet loss rate exceed the benchmark values, or the transmission link path deviates from the preset route, it is marked as a logical layer resource abnormality or transmission abnormality. Extract the rules from the configuration data verification library, perform compliance verification on the configuration data of the logical layer, and mark the configuration of the logical layer as abnormal if the configuration does not conform to the rules or the dependent components are missing. Obtain application layer business performance data from the hierarchical digital twin model and compare it with the benchmark value in the application layer dynamic benchmark library. If the business performance data exceeds the benchmark value or the trend of change is abnormal, it is marked as an application layer performance anomaly. Monitor the application layer call chain status, including microservice call failures, call timeouts, and chain interruption status. Combine this with the status data of the physical and logical layers, which are free of anomalies. If any chain status problem exists, it is marked as an application layer chain anomaly.
[0012] Furthermore, when generating tiered early warning information, the process also includes determining the root cause of the anomaly based on labeled abnormal state data: Acquire marked application layer abnormal status data, including application layer performance abnormalities and application layer link abnormalities, and associate logical layer resource usage data with physical layer hardware status data; Based on the correlation results with the logical layer and physical layer, if the resource usage of the logical layer exceeds the baseline value of the logical layer dynamic benchmark library, and the hardware parameters of the physical layer exceed the baseline value of the physical layer dynamic benchmark library, it is determined that the application layer is abnormal due to the underlying resources. If the logical layer resource usage data and the physical layer hardware status data match the benchmark values of the corresponding dynamic benchmark library, it is determined that the application itself is abnormal. Obtain the marked logical layer abnormal status data, including logical layer resource abnormalities, transmission abnormalities, and configuration abnormalities, and associate them with physical layer network device status data; Based on the correlation results with the physical layer, if the parameters of the physical layer network device are lower than the baseline values of the physical layer dynamic baseline library, it is determined to be a logical layer anomaly caused by the physical layer hardware; if the physical layer network device data meets the baseline values of the physical layer dynamic baseline library, it is determined to be a logical layer protocol configuration anomaly.
[0013] Furthermore, a fault knowledge graph is constructed, specifically including: Collect historical fault data from industrial software cloud infrastructure to form a knowledge data source; Define target entities and relationships between entities in the knowledge graph template. Target entities include fault types, anomaly indicators, cause entities, hierarchical data nodes, and self-healing strategies. The knowledge data source is transformed into target entities and relationships in a knowledge graph, generating an initial fault knowledge graph, and a real-time update mechanism is established to automatically update the target entity attributes and relationships in the graph.
[0014] Furthermore, based on the fault knowledge graph and fault analysis dataset, the unique cause of the fault can be located, specifically including: Obtain the fault analysis dataset, and based on the abnormal indicators and hierarchical data node entities in the fault knowledge graph, perform semantic matching on the fault analysis dataset to determine the corresponding node of each entry in the dataset in the graph. Based on the correlation between the matched graph nodes and the abnormal indicators and causes, and the hierarchical data nodes and causes in the fault knowledge graph, generate at least one potential cause reasoning path. Based on the relationships between the physical layer, logical layer, and application layer of the hierarchical digital twin model, the rationality of each reasoning path is verified. Based on the reasoning path of potential causes after rationality verification, the corresponding fault model is matched in the fault knowledge graph, and the corresponding diagnostic knowledge is extracted from the graph based on the selected fault model. Based on the extracted diagnostic knowledge, supplementary data is retrieved, and fuzzy matching is performed between the supplementary data and the standard diagnostic data in the fault knowledge graph to calculate the matching similarity of the supplementary data. Based on the confidence propagation algorithm, combined with the similarity of the supplementary data and the weights of the abnormal indicators in the fault analysis dataset, the confidence of the current fault cause is calculated. If the confidence level is lower than the preset confidence threshold, the range of supplementary data retrieval will be adjusted and rematched. If the confidence level is higher than the preset confidence threshold, the search will continue to the next level along the current potential cause based on the linkage relationship between the physical layer, logical layer and application layer of the hierarchical digital twin model. Based on the fault analysis dataset, the matching degree of each potential cause after confidence verification is calculated, and the potential cause with the highest matching degree is selected. Obtain the corresponding layered data from the visualization monitoring interface, perform entity verification on the selected potential causes, confirm that the potential causes can explain all abnormal phenomena, and locate the unique cause of the failure.
[0015] Furthermore, identifying the sole cause of the failure also includes: The single cause of failure output is input into the hierarchical digital twin model to reproduce the failure chain. Based on the reproduction results, the causal relationship between the cause of failure and the abnormal phenomenon is verified. Based on the reproduced fault chain, combined with the hierarchical correlation data displayed on the visual monitoring interface, the impact of the fault cause on the physical layer, logic layer, and application layer and the transmission path are analyzed to generate a fault impact chain report. Obtain the correlation between causes and adaptive self-healing strategies in the fault knowledge graph, analyze the blocking coverage of each candidate self-healing strategy on the fault impact chain in conjunction with the fault impact chain report, calculate the adaptability of the candidate strategies, and synchronize the calculation results to the visualization monitoring interface.
[0016] Compared with the prior art, the beneficial effects of the present invention are: This invention constructs and dynamically updates a layered digital twin model, combining independent visualization modules for the physical, logical, and application layers to achieve precise mapping and intuitive display of the cloud infrastructure's operational status. By dynamically calibrating benchmark values in real time through a benchmark library, it quickly identifies anomalies at each layer and generates tiered early warnings, providing precise data support for operations and maintenance, improving the timeliness of anomaly response. Based on historical fault data, it constructs a dynamically updated fault knowledge graph, integrating multi-dimensional correlation analysis and confidence verification mechanisms to accurately pinpoint the unique cause of a fault, avoiding the blind troubleshooting of traditional operations and maintenance, significantly shortening the fault diagnosis cycle, and improving the accuracy of cause location. Relying on the fault knowledge graph and historical self-healing cases, it constructs a strategy library, simulating and verifying candidate strategies through the digital twin model to ensure the effectiveness of the self-healing strategy before issuing and executing it, achieving intelligent automatic fault repair, reducing manual intervention, lowering operations and maintenance costs, and ensuring the stable operation of the cloud infrastructure. Attached Figure Description
[0017] Figure 1 This is a flowchart of the intelligent operation and maintenance and fault self-healing method for industrial software cloud base of the present invention. Figure 2 This is a flowchart for accurately locating the cause of a fault in this invention. Detailed Implementation
[0018] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0019] To address the technical issues that existing technologies fail to cover both the logical and application layers of the cloud infrastructure, thus hindering multi-layer linkage mapping, and suffer from inefficient and inaccurate cause localization due to the lack of a fault knowledge graph, and the absence of self-healing strategies that can easily lead to fault escalation or business interruption during simulation verification, please refer to [link to relevant documentation]. Figures 1-2 This embodiment provides the following technical solution: Intelligent operation and maintenance and fault self-healing methods for industrial software cloud infrastructure based on digital twins include: Digital twin model construction: Obtain physical structure data and real-time operating status data of the industrial software cloud base. The physical structure data includes physical structure parameters of hardware components, network topology data, and software deployment configuration data. The real-time operating status data includes resource utilization, response latency, and log data. Construct a layered digital twin model of the cloud base and dynamically update the layered digital twin model by combining the geometric parameters, performance parameters, and topological relationships of the physical base. Real-time visual monitoring and early warning: Based on the constructed layered digital twin model, the actual operating status of the industrial software cloud base is mapped in real time and simulated. A visual monitoring interface is constructed for visualization display, including the physical layer, logical layer, and application layer architecture based on the layered digital twin model. The visual monitoring interface sets up corresponding three independent visualization modules, and the status data of the three visualization modules are updated synchronously based on the model accuracy calibration cycle. The physical layer geometric model changes with the hardware status, such as grayscale display when the equipment is offline. The logical layer topology links are dynamically adjusted with resource scheduling changes. The application layer link diagram updates node connections with business call relationships. Abnormal states that deviate from the normal range are identified, and graded early warning information is generated, including information such as severity level (fatal / serious / moderate / minor), abnormal location, scope of impact, and risk level. Fault diagnosis: Constructing a fault knowledge graph, specifically including: Knowledge data collection: Collect historical fault data of industrial software cloud base, including fault phenomena, abnormal indicators, causes, handling solutions, hardware / software operation and maintenance manuals, including equipment fault modes, software dependency conflict rules, abnormal status data and corresponding handling records, to form a knowledge data source; Entity and Relationship Definition: Define target entities and relationships between entities in the knowledge graph template. Target entities include fault types, anomaly indicators, cause entities, hierarchical data nodes, and self-healing strategies. The relationships include anomaly indicators and corresponding causes, causes and associated hierarchical data nodes, and causes and adaptive self-healing strategies. Knowledge graph construction and updating: The semantic mapping algorithm is used to transform the knowledge data source into target entities and relationships in the knowledge graph, generate the initial fault knowledge graph, and establish a real-time update mechanism. When new fault cases, abnormal states, or hierarchical data association rules change, the target entity attributes and relationships in the graph are automatically updated. Extract abnormal indicator data corresponding to abnormal states, perform multi-dimensional data correlation analysis, filter entity data associated with abnormal indicators, generate a fault analysis dataset, and locate the unique fault cause based on the fault knowledge graph and the fault analysis dataset. Intelligent fault self-healing: A self-healing strategy library is built based on the cause nodes of the fault knowledge graph and historical self-healing cases. Each self-healing strategy in the library includes parameter adjustment, service restart, resource scheduling, redundancy switching, and module isolation. The self-healing strategy is matched according to the located fault cause. The self-healing strategy is input into the hierarchical digital twin model for simulation verification. The effective self-healing strategy is obtained and distributed to the cloud base for execution.
[0020] In this embodiment, by integrating physical structure and real-time operational data, the hierarchical model is dynamically updated to achieve accurate mapping of the cloud base status, providing dynamic data support for operation and maintenance. Based on the real-time mapping of the hierarchical model, three independent visualization modules are designed and updated synchronously. The three modules are dynamically adjusted according to different states, intuitively presenting the status and promptly identifying anomalies to generate hierarchical warnings. This facilitates operation and maintenance personnel to quickly grasp the status of each layer of the cloud base and respond to anomalies in a timely manner. Multi-source data is transformed into a graph and dynamically updated to improve the accuracy of cause localization, providing accurate cause basis for fault handling and reducing blind investigation. A strategy library is established based on causes and historical cases, and after simulation verification, it is deployed and executed to achieve automatic fault handling. Combining the graph and simulation verification to match strategies avoids invalid operations, improves fault handling efficiency, and reduces manual intervention costs.
[0021] In this embodiment, the process of constructing a layered digital twin model of the cloud base includes: Physical layer modeling: Extract physical structure data of the industrial software cloud base, including physical structure parameters such as the size, installation location, and interface type of servers, storage devices, and network devices, as well as hardware performance parameters such as CPU frequency, memory capacity, and storage IOPS. Combined with the network topology data of the cloud base, construct a geometric simulation model of the cloud base to generate physical layer state data, including real-time hardware operating parameters and device physical location information, and realize the physical state mapping of hardware devices. Logical layer modeling: Collect virtualization resource configurations of the cloud base, such as virtual machine specifications, container cluster configurations, network routing rules, data transmission protocols and software dependencies, construct a logical topology model of the cloud base, simulate resource scheduling processes, data transmission paths and software interaction logic, and generate logical layer configuration data, including resource scheduling rules, data transmission path parameters and software interaction protocol data. Application layer modeling: This involves acquiring the deployment architecture, service call chains, business process configurations, and performance thresholds of industrial software. The deployment architecture includes microservice decomposition and deployment node structure. It also involves building an application runtime model for the cloud infrastructure, linking physical layer state data with logical layer configuration data, establishing a mapping relationship between business performance indicators and the state data of the geometric simulation model and the configuration data of the logical topology model, achieving a linked mapping between business runtime status and underlying resources, and generating a layered digital twin model of the cloud infrastructure. This also includes: Construct accuracy calibration rules for the geometric simulation model, and obtain real-time hardware operating parameters from the physical layer state data based on a preset extraction time interval; The real-time operating parameters of the hardware are compared with the actual operating status of the cloud base hardware output by the geometric simulation model. Based on the comparison results, the mapping coefficients of the hardware performance parameters of the geometric simulation model are corrected until the deviation between the model output and the actual state of the physical device is less than the preset deviation value, so as to ensure the consistency between the physical layer model and the actual state of the hardware. Obtain the virtualization resource configuration and software dependencies of the cloud base, and build a configuration data verification library based on the dependencies, including compliance rules for virtualization resource configuration and integrity rules for software dependencies; Extract the configuration data from the logical topology model, compare the configuration data with the rules of the configuration data verification library, and generate configuration correction suggestions; When the virtualized resources of the cloud base change, the resource scheduling simulation parameters and data transmission path parameters of the logical topology model are automatically updated. The adaptation of the logical topology model to the actual virtualized resource state is achieved through configuration updates. Based on the mapping relationship, a collaborative adjustment mechanism for performance index thresholds is established. When the hardware performance parameters of the physical layer model or the resource configuration of the logical topology model changes, the business performance index thresholds of the application layer model are recalculated through the collaborative adjustment mechanism for performance index thresholds, and the business performance monitoring rules of the application layer model are updated synchronously.
[0022] In this embodiment, a geometric simulation model is constructed by extracting hardware physical parameters, performance parameters, and network topology data. Precision calibration rules are added to achieve accurate mapping of the hardware physical state, ensuring consistency between the model and the actual hardware state. The model integrates dual-dimensional hardware parameters with network topology modeling, providing accurate model support for physical layer state monitoring and visualization. Dynamic comparison and correction address static model deviation issues. Configuration is corrected through a verification library, and parameters are automatically updated when resources change, avoiding the lag problem of manual adjustments. This provides an accurate model for logical layer resource scheduling analysis. An application operation model is constructed in conjunction with an industrial software deployment framework, establishing a mapping relationship between physical layer state and logical layer configuration data. A performance threshold collaborative adjustment mechanism is built to achieve linkage mapping between business and underlying resources. When the underlying layer changes, application layer rules are automatically and collaboratively adjusted to avoid monitoring disconnection, providing an accurate model for application layer business performance analysis.
[0023] In this embodiment, the comparison results are used to correct the hardware performance parameter mapping coefficients of the geometric simulation model, including: Extract the simulated values of the current hardware operating status output from the geometric simulation model; Retrieve the current hardware real-time operating parameters corresponding to the current hardware operating status simulation value, and use the current hardware real-time operating parameters and the current hardware operating status simulation value to perform absolute difference processing to obtain the current absolute error value; The error rate is obtained by comparing the absolute error value with the current real-time operating parameters of the hardware. The preset basic correction step size α is retrieved from the database, wherein the value range of the basic correction step size is 0 < α ≤ 0.5, and the default value is 0.38; The error rate corresponding to the previous hardware performance parameter mapping coefficient correction process is retrieved as the preceding error rate; The preset basic correction step size α is adjusted using the aforementioned error rate and the preceding error rate to obtain the adjusted correction step size α. t =(1+EE x )×α; where α represents the preset base correction step size; α t Indicates the adjusted correction step size; E represents the current error rate; E x Indicates the preceding error rate; Using the adjusted correction step size α t The mapping coefficients for hardware performance parameters are corrected.
[0024] The aforementioned technical solution combines absolute error and error rate as dual indicators to evaluate the deviation between simulated values and real-time parameters. Compared to relying solely on absolute error, this approach is more adaptable to hardware performance parameters of different magnitudes, avoiding distortion in error judgment caused by differences in parameter magnitudes and improving the accuracy of deviation assessment. Secondly, it dynamically adjusts the correction step size based on the current error rate; a higher error rate results in a larger adjustment step size, quickly reducing large deviations while avoiding coefficient oscillations caused by over-correction, achieving a balance between correction efficiency and stability. Furthermore, by referencing previous error rates and connecting them to historical correction data, the mapping coefficient correction is made continuous, avoiding coefficient fluctuations caused by isolated adjustments and enhancing the consistency of the correction process. In addition, the corrected mapping coefficients better reflect the actual operating state of the hardware, effectively reducing the output deviation of the geometric simulation model, improving the model's simulation accuracy of the hardware's operating state, and providing reliable support for the physical layer accuracy of the layered digital twin model of the cloud platform. Finally, the entire correction process requires no manual intervention. It relies on automated processes to dynamically optimize the mapping coefficients, adapt to real-time changes in hardware operating status, and ensure that the geometric simulation model maintains high adaptability over the long term. This provides accurate model data support for the intelligent operation and maintenance and fault self-healing of industrial software cloud infrastructure, indirectly improving operation and maintenance efficiency and the reliability of fault diagnosis.
[0025] In this embodiment, the adjusted correction step size α is used. t Correcting the mapping coefficients of hardware performance parameters, including: Retrieve the simulated hardware operating status values and real-time hardware operating parameters output by the geometric simulation model during the previous hardware performance parameter mapping coefficient correction process; The simulated hardware operating status value and real-time hardware operating parameters output by the geometric simulation model during the previous hardware performance parameter mapping coefficient correction process are used as the reference simulated hardware operating status value and reference real-time hardware operating parameters. The error trend factor r = 1 - 0.5 × |sign(x) is obtained by combining the simulated values of the reference hardware operating status and the real-time operating parameters of the reference hardware with the current real-time operating parameters and the simulated values of the current hardware operating status. dr -x r )-sign(x dm -x m )|, where r represents the error trend factor; x dr and x dm This represents the current real-time operating parameters and simulated values of the current hardware operating status; x r and x m This represents the simulated values of the reference hardware's operating status and the real-time operating parameters of the reference hardware. Retrieve the adjusted correction step size α t The coefficient value k0 of the hardware performance parameter mapping coefficient after the last correction; The error trend factor is combined with the adjusted correction step size α t The hardware performance parameter mapping coefficients are corrected based on the coefficient value k0 of the hardware performance parameter mapping coefficients.
[0026] The corrected hardware performance parameter mapping coefficients are obtained using the following formula: k t =k0×α t ×(1-r) Where, k t α represents the corrected hardware performance parameter mapping coefficient; k0 represents the coefficient value of the hardware performance parameter mapping coefficient after the last correction; t This indicates the adjusted correction step size; r represents the error trend factor.
[0027] The aforementioned technical solution introduces an error trend factor *r*, quantifying the error change trend based on the sign difference between the current and reference hardware real-time parameters and simulated values. This accurately determines whether the deviation is widening or converging, avoiding blind correction and improving the targeting and rationality of the correction. Secondly, the correction process connects the previous mapping coefficient *k0* with historical reference data, ensuring continuous correlation in coefficient correction. This avoids drastic fluctuations caused by isolated adjustments, enhancing the stability of the correction process and reducing oscillation risk. Furthermore, the adjusted correction step size α... tIn synergy with the error trend factor r, it effectively improves the dynamic balance between correction efficiency and stability. Furthermore, the corrected mapping coefficients better reflect the actual operating patterns of the hardware, effectively reducing simulation deviations in the geometric simulation model and significantly improving the model's accuracy in simulating the hardware's operating state. This provides solid support for the physical layer accuracy of the layered digital twin model of the cloud platform, thereby ensuring the accuracy of decision-making in subsequent intelligent operation and maintenance, fault self-healing, and other scenarios, indirectly improving the reliability and operational efficiency of the cloud platform.
[0028] In this embodiment, identifying abnormal states that deviate from the normal range specifically includes: Based on the data characteristics of the physical, logical, and application layers of the hierarchical digital twin model, and combined with historical operational data and industry standards, a dynamic benchmark library for each layer is established. In this embodiment, when constructing the dynamic benchmark library, it is necessary to set benchmarks for each level based on the historical operational data and hierarchical data characteristics of the cloud base, specifically including: Obtain hardware performance parameters from physical structure data and extract historical normal operation data of hardware performance parameters; based on the statistical quantiles of historical normal operation data, set physical layer static benchmarks, which include benchmarks set based on hardware resource utilization and hardware operating status parameters; Obtain the logical layer configuration data from the physical structure data, and combine the operational requirements of logical layer resource scheduling and data transmission to set the static benchmark of the logical layer. The static benchmark of the logical layer includes benchmarks based on resource scheduling response efficiency and data transmission quality parameters. Acquire business performance data from real-time operational status data and classify the business performance data according to business type; based on the service quality requirements of different business types, set application layer benchmarks, including benchmarks based on business response efficiency, business processing capacity, and business operation stability parameters; The above-mentioned physical layer static benchmarks, logic layer static benchmarks, and application layer benchmarks are integrated to form the basic benchmark framework of the dynamic benchmark library, providing a hierarchical benchmark basis for subsequent dynamic calibration. By associating the dynamic benchmark libraries of each layer with the physical layer hardware performance parameters, logical layer configuration data, and application layer business performance index thresholds of the hierarchical digital twin model, the benchmark values of each layer's dynamic benchmark library are calibrated in real time.
[0029] In this embodiment, identifying abnormal states that deviate from the normal range further includes: Obtain the physical layer status data of the hierarchical digital twin model, compare the physical layer status data with the benchmark value of the corresponding physical layer dynamic benchmark library, and if the data exceeds the benchmark value for multiple consecutive acquisition cycles, it is marked as an abnormal physical layer parameter. The physical layer geometric simulation model based on the hierarchical digital twin model monitors the target operating status of cloud base hardware devices, including device offline, interface connection interruption or physical location offset. If the target operating status is detected, it is marked as an abnormal physical layer device status. Obtain the logical layer resource usage data and transmission link data of the hierarchical digital twin model, and compare them with the benchmark values of the corresponding logical layer dynamic benchmark library. If the resource usage data, transmission delay / packet loss rate exceed the benchmark values, or the transmission link path deviates from the preset route, it is marked as a logical layer resource abnormality or transmission abnormality. Extract the rules from the configuration data verification library, perform compliance verification on the configuration data of the logical layer, and mark the configuration of the logical layer as abnormal if the configuration does not conform to the rules or the dependent components are missing. Obtain application layer business performance data from the hierarchical digital twin model and compare it with the benchmark value in the application layer dynamic benchmark library. If the business performance data exceeds the benchmark value or the trend of change is abnormal, it is marked as an application layer performance anomaly. Monitor the application layer call chain status, including microservice call failures, call timeouts, and chain interruption status. Combine this with the status data of the physical and logical layers, which are free of anomalies. If any chain status problem exists, it is marked as an application layer chain anomaly.
[0030] In this embodiment, static benchmarks are set for the physical, logical, and application layers by combining historical operational data of the cloud base, industry standards, and data characteristics of each layer. After integration into the basic framework, the data of the hierarchical digital twin model is calibrated in real time to form a hierarchical and dynamically adaptable benchmark basis. This avoids the disconnect between the fixed benchmark and the actual operating state and ensures that the benchmark adapts to the dynamic changes of the cloud base. Continuous periodic comparison is used to eliminate accidental interference, and the hardware entity status is directly monitored by combining the geometric model to achieve dual-dimensional anomaly identification of parameters and entity status. This helps maintenance personnel to quickly locate the source of physical layer faults and prevent the spread of hardware problems. The benchmark data comparison is combined with configuration rule verification to achieve multi-dimensional anomaly detection, comprehensively investigate potential problems in the logical layer, and provide complete anomaly information for subsequent fault diagnosis. Application layer anomaly identification distinguishes between anomalies caused by application layer faults and underlying problems, accurately pinpointing application layer faults and improving the accuracy of anomaly location.
[0031] In this embodiment, when generating graded early warning information, the method further includes determining the root cause of the anomaly based on the labeled abnormal state data: Acquire marked application layer abnormal status data, including application layer performance abnormalities and application layer link abnormalities, and associate logical layer resource usage data with physical layer hardware status data; Based on the correlation results with the logical layer and physical layer, if the resource usage of the logical layer exceeds the baseline value of the logical layer dynamic benchmark library, and the hardware parameters of the physical layer exceed the baseline value of the physical layer dynamic benchmark library, it is determined that the application layer is abnormal due to the underlying resources. If the logical layer resource usage data and the physical layer hardware status data match the benchmark values of the corresponding dynamic benchmark library, it is determined that the application itself is abnormal. Obtain the marked logical layer abnormal status data, including logical layer resource abnormalities, transmission abnormalities, and configuration abnormalities, and associate them with physical layer network device status data; Based on the correlation results with the physical layer, if the parameters of the physical layer network device are lower than the baseline values of the physical layer dynamic baseline library, it is determined to be a logical layer anomaly caused by the physical layer hardware; if the physical layer network device data meets the baseline values of the physical layer dynamic baseline library, it is determined to be a logical layer protocol configuration anomaly.
[0032] In this embodiment, root cause classification and determination are achieved through cross-layer data association, solving the problem of ambiguous root causes in traditional determination and helping operation and maintenance personnel quickly clarify the responsibility for application layer anomalies. A correlation determination mechanism between logical layer and physical layer network devices is established to break the disconnect between the two layers of anomaly determination, achieve accurate classification of root cause types, enable operation and maintenance personnel to carry out targeted hardware repair or configuration optimization, reduce ineffective operation and maintenance operations, and provide more specific root cause information support for hierarchical early warning.
[0033] In this embodiment, the unique cause of a fault is located based on a fault knowledge graph and a fault analysis dataset, specifically including: Obtain the fault analysis dataset, and based on the abnormal indicators and hierarchical data node entities in the fault knowledge graph, perform semantic matching on the fault analysis dataset to determine the corresponding node of each entry in the dataset in the graph. Based on the correlation between the matched graph nodes and the abnormal indicators and causes, and the hierarchical data nodes and causes in the fault knowledge graph, generate at least one potential cause reasoning path. Based on the relationship between the physical layer, logical layer, and application layer of the hierarchical digital twin model, the rationality of each reasoning path is verified, and paths that contradict the hierarchical linkage relationship are eliminated, such as paths that point to physical layer causes even though there are no abnormalities in the physical layer. Based on the reasoning path of potential causes after rationality verification, the corresponding fault model is matched in the fault knowledge graph. Based on the selected fault model, the corresponding diagnostic knowledge in the graph is extracted, including abnormal features, correlation parameters, and verification points. Based on the extracted diagnostic knowledge, supplementary data is retrieved, including historical calibration data of the hierarchical digital twin model, configuration change logs of the cloud base, and historical comparison data of real-time operating status data. The supplementary data is then fuzzily matched with the standard diagnostic data in the fault knowledge graph, and the matching similarity of the supplementary data is calculated. Based on the confidence propagation algorithm, combined with the similarity of supplementary data matching and the weights of abnormal indicators in the fault analysis dataset, the confidence level of the current fault cause is calculated. If the confidence level is lower than the preset confidence threshold, the scope of supplementary data retrieval will be adjusted, such as adding historical fault-related data from the past 3 months for rematching. If the confidence level is higher than the preset confidence threshold, the search will continue to the next lower level along the layer where the current potential cause is located, based on the linkage relationship between the physical layer, logical layer, and application layer of the hierarchical digital twin model. For example, if the current associated application layer is extended down to the logical layer, and if the current associated logical layer is extended down to the physical layer. Based on the fault analysis dataset, the matching degree of each potential cause after confidence verification is calculated, such as the consistency of abnormal indicators and the consistency of associated hierarchical data, and the potential causes with the highest matching degree are selected. Acquire the corresponding layered data from the visual monitoring interface, including physical layer hardware data, logical layer resource data, and application layer business data. Perform entity verification on the selected potential causes to confirm that the potential causes can explain all abnormal phenomena and locate the unique cause of the fault. The single fault cause output is input into the hierarchical digital twin model to simulate the hardware state changes, resource scheduling anomalies, and business indicator fluctuations after the cause is triggered, reproduce the fault occurrence chain, and verify the causal relationship between the fault cause and the abnormal phenomenon based on the reproduction results. Based on the reproduced fault chain, combined with the hierarchical correlation data displayed on the visual monitoring interface, the impact and transmission path of the fault cause on the physical layer related hardware load transmission, logical layer resource scheduling blockage propagation, and application layer business process interruption scope are analyzed, and a fault impact chain report is generated. The system obtains the correlation between causes and adaptive self-healing strategies in the fault knowledge graph, analyzes the blocking coverage of each candidate self-healing strategy on the fault impact chain in conjunction with the fault impact chain report, calculates the adaptability of the candidate strategies, and synchronizes the calculation results to the visualization monitoring interface to provide a basis for subsequent self-healing strategy matching.
[0034] In this embodiment, semantic matching is used to associate fault analysis datasets and fault knowledge graph nodes to generate potential cause reasoning paths, eliminate contradictory paths, and ensure that the paths conform to the layered operation logic of the cloud base. This reduces the time that maintenance personnel spend investigating invalid paths and improves the efficiency of cause search. The data range is dynamically adjusted and the confidence level is quantified, breaking through the limitations of traditional fixed data sources and lack of confidence assessment. This provides maintenance personnel with clear evidence of whether the cause is credible and provides quantitative support for cause judgment. When the confidence level is high, the search is extended downward along the layers. The potential cause entities are verified by real-time data of each layer through the visual monitoring interface, ensuring that the cause can explain all anomalies. This enables the tracing of causes from the surface to the depths, achieving cross-layer deep cause mining, helping maintenance personnel to pinpoint the real core cause rather than the surface anomaly, and providing precise targeting for subsequent self-healing. The simulation capabilities of digital twins are used to realize fault reproduction and strengthen the reliability of causal relationships. Cause location is pre-linked with the adaptability of self-healing strategies, providing data support for maintenance personnel to select self-healing strategies and avoiding handling failures caused by strategy mismatch.
[0035] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
Claims
1. A method for intelligent operation and maintenance and fault self-healing of industrial software cloud infrastructure based on digital twins, characterized in that: include: Acquire physical structure data and real-time operating status data of the industrial software cloud base, construct a hierarchical digital twin model of the cloud base, and dynamically update the hierarchical digital twin model by combining the geometric parameters, performance parameters and topological relationships of the physical base. Based on the constructed hierarchical digital twin model, the actual operating status of the industrial software cloud base is mapped in real time and simulated. A visualization monitoring interface is constructed for visualization display, including the physical layer, logical layer and application layer architecture based on the hierarchical digital twin model. The corresponding three independent visualization modules are set in the visualization monitoring interface, and the status data of the three visualization modules are updated synchronously based on the model accuracy calibration cycle. Abnormal states that deviate from the normal range are identified, and hierarchical early warning information is generated. Construct a fault knowledge graph, extract abnormal indicator data corresponding to abnormal states, perform multi-dimensional data association analysis, filter entity data associated with abnormal indicators, generate a fault analysis dataset, and locate the unique fault cause based on the fault knowledge graph and the fault analysis dataset. A self-healing strategy library is built based on the cause nodes of the fault knowledge graph and historical self-healing cases. The self-healing strategy is matched according to the located fault cause, and the self-healing strategy is input into the hierarchical digital twin model for simulation verification. Effective self-healing strategies are obtained and deployed to the cloud base for execution.
2. The intelligent operation and maintenance and fault self-healing method for industrial software cloud infrastructure based on digital twins as described in claim 1, characterized in that, The process of building a layered digital twin model for the cloud infrastructure includes: Extract the physical structure data of the industrial software cloud base, including the physical structure parameters and hardware performance parameters of the hardware components, and combine it with the network topology data of the cloud base to construct a geometric simulation model of the cloud base and generate physical layer state data. Collect the virtualization resource configuration, network routing rules, data transmission protocols and software dependencies of the cloud base, construct the logical topology model of the cloud base, simulate the resource scheduling process, data transmission path and software interaction logic, and generate logical layer configuration data; Obtain the deployment architecture, service call chain, business process configuration and performance indicator thresholds of industrial software, build the application operation model of the cloud base, and associate physical layer status data with logical layer configuration data; Establish a mapping relationship between business performance indicators and the state data of the geometric simulation model and the configuration data of the logical topology model to generate a layered digital twin model of the cloud base.
3. The intelligent operation and maintenance and fault self-healing method for industrial software cloud base based on digital twin as described in claim 2, characterized in that, The layered digital twin model of the cloud base also includes: Construct accuracy calibration rules for the geometric simulation model, and obtain real-time hardware operating parameters from the physical layer state data based on a preset extraction time interval; The real-time operating parameters of the hardware are compared with the actual operating status of the cloud base hardware output by the geometric simulation model, and the mapping coefficients of the hardware performance parameters of the geometric simulation model are corrected based on the comparison results. Obtain the virtualization resource configuration and software dependencies of the cloud base, and build a configuration data verification library based on the dependencies; Extract the configuration data from the logical topology model, compare the configuration data with the rules of the configuration data verification library, and generate configuration correction suggestions; When the virtualized resources of the cloud base change, the resource scheduling simulation parameters and data transmission path parameters of the logical topology model are automatically updated. Based on the mapping relationship, a collaborative adjustment mechanism for performance index thresholds is established. When the hardware performance parameters of the physical layer model or the resource configuration of the logical topology model changes, the business performance index thresholds of the application layer model are recalculated through the collaborative adjustment mechanism for performance index thresholds, and the business performance monitoring rules of the application layer model are updated synchronously.
4. The intelligent operation and maintenance and fault self-healing method for industrial software cloud base based on digital twin as described in claim 3, characterized in that, The comparison results correct the hardware performance parameter mapping coefficients of the geometric simulation model, including: Extract the simulated values of the current hardware operating status output from the geometric simulation model; Retrieve the current hardware real-time operating parameters corresponding to the current hardware operating status simulation value, and use the current hardware real-time operating parameters and the current hardware operating status simulation value to perform absolute difference processing to obtain the current absolute error value; The error rate is obtained by comparing the absolute error value with the current real-time operating parameters of the hardware. Retrieve the preset basic correction step size α from the database, wherein the value of the basic correction step size is in the range of 0 < α ≤ 0.5; The error rate corresponding to the previous hardware performance parameter mapping coefficient correction process is retrieved as the preceding error rate; The preset basic correction step size α is adjusted using the aforementioned error rate and the preceding error rate to obtain the adjusted correction step size α. t ; Using the adjusted correction step size α t The mapping coefficients for hardware performance parameters are corrected.
5. The intelligent operation and maintenance and fault self-healing method for industrial software cloud infrastructure based on digital twins as described in claim 4, characterized in that, Using the adjusted correction step size α t Correcting the mapping coefficients of hardware performance parameters, including: Retrieve the simulated hardware operating status values and real-time hardware operating parameters output by the geometric simulation model during the previous hardware performance parameter mapping coefficient correction process; The simulated hardware operating status value and real-time hardware operating parameters output by the geometric simulation model during the previous hardware performance parameter mapping coefficient correction process are used as the reference simulated hardware operating status value and reference real-time hardware operating parameters. The error trend factor r is obtained by combining the simulated values of the reference hardware operating status and the real-time operating parameters of the reference hardware with the real-time operating parameters of the current hardware and the simulated values of the current hardware operating status. Retrieve the adjusted correction step size α t The coefficient value k0 of the hardware performance parameter mapping coefficient after the last correction; The error trend factor is combined with the adjusted correction step size α t The hardware performance parameter mapping coefficients are corrected based on the coefficient value k0 of the hardware performance parameter mapping coefficients.
6. The intelligent operation and maintenance and fault self-healing method for industrial software cloud infrastructure based on digital twins as described in claim 1, characterized in that, Identify abnormal states that deviate from the normal range, specifically including: Based on the data characteristics of the physical, logical, and application layers of the hierarchical digital twin model, and combined with historical operational data and industry standards, a dynamic benchmark library for each layer is established. By associating the dynamic benchmark libraries of each layer with the physical layer hardware performance parameters, logical layer configuration data, and application layer business performance index thresholds of the hierarchical digital twin model, the benchmark values of each layer's dynamic benchmark library are calibrated in real time.
7. The intelligent operation and maintenance and fault self-healing method for industrial software cloud base based on digital twin as described in claim 6, characterized in that, Identifying abnormal states that deviate from the normal range also includes: Obtain the physical layer status data of the hierarchical digital twin model, compare the physical layer status data with the benchmark value of the corresponding physical layer dynamic benchmark library, and if the data exceeds the benchmark value for multiple consecutive acquisition cycles, it is marked as an abnormal physical layer parameter. The physical layer geometric simulation model based on the hierarchical digital twin model monitors the target operating status of cloud base hardware devices, including device offline, interface connection interruption or physical location offset. If the target operating status is detected, it is marked as an abnormal physical layer device status. Obtain the logical layer resource usage data and transmission link data of the hierarchical digital twin model, and compare them with the benchmark values of the corresponding logical layer dynamic benchmark library. If the resource usage data, transmission delay / packet loss rate exceed the benchmark values, or the transmission link path deviates from the preset route, it is marked as a logical layer resource abnormality or transmission abnormality. Extract the rules from the configuration data verification library, perform compliance verification on the configuration data of the logical layer, and mark the configuration of the logical layer as abnormal if the configuration does not conform to the rules or the dependent components are missing. Obtain application layer business performance data from the hierarchical digital twin model and compare it with the benchmark value in the application layer dynamic benchmark library. If the business performance data exceeds the benchmark value or the trend of change is abnormal, it is marked as an application layer performance anomaly. Monitor the application layer call chain status, including microservice call failures, call timeouts, and chain interruption status. Combine this with the status data of the physical and logical layers, which are free of anomalies. If any chain status problem exists, it is marked as an application layer chain anomaly.
8. The intelligent operation and maintenance and fault self-healing method for industrial software cloud base based on digital twin as described in claim 1, characterized in that, When generating tiered early warning information, the process also includes determining the root cause of the anomaly based on labeled abnormal state data. Acquire marked application layer abnormal status data, including application layer performance abnormalities and application layer link abnormalities, and associate logical layer resource usage data with physical layer hardware status data; Based on the correlation results with the logical layer and physical layer, if the resource usage of the logical layer exceeds the baseline value of the logical layer dynamic benchmark library, and the hardware parameters of the physical layer exceed the baseline value of the physical layer dynamic benchmark library, it is determined that the application layer is abnormal due to the underlying resources. If the logical layer resource usage data and the physical layer hardware status data match the benchmark values of the corresponding dynamic benchmark library, it is determined that the application itself is abnormal. Obtain the marked logical layer abnormal status data, including logical layer resource abnormalities, transmission abnormalities, and configuration abnormalities, and associate them with physical layer network device status data; Based on the correlation results with the physical layer, if the parameters of the physical layer network device are lower than the baseline values of the physical layer dynamic baseline library, it is determined to be a logical layer anomaly caused by the physical layer hardware. If the data from the physical layer network device matches the baseline value of the physical layer dynamic baseline library, it is determined that the logical layer protocol configuration is abnormal.
9. The intelligent operation and maintenance and fault self-healing method for industrial software cloud base based on digital twin as described in claim 1, characterized in that, Constructing a fault knowledge graph, specifically including: Collect historical fault data from industrial software cloud infrastructure to form a knowledge data source; Define target entities and relationships between entities in the knowledge graph template. Target entities include fault types, anomaly indicators, cause entities, hierarchical data nodes, and self-healing strategies. The knowledge data source is transformed into target entities and relationships in a knowledge graph, generating an initial fault knowledge graph, and a real-time update mechanism is established to automatically update the target entity attributes and relationships in the graph.
10. The intelligent operation and maintenance and fault self-healing method for industrial software cloud infrastructure based on digital twins as described in claim 1, characterized in that, Based on fault knowledge graphs and fault analysis datasets, unique fault causes can be located, specifically including: Obtain the fault analysis dataset, and based on the abnormal indicators and hierarchical data node entities in the fault knowledge graph, perform semantic matching on the fault analysis dataset to determine the corresponding node of each entry in the dataset in the graph. Based on the correlation between the matched graph nodes and the abnormal indicators and causes, and the hierarchical data nodes and causes in the fault knowledge graph, generate at least one potential cause reasoning path. Based on the relationships between the physical layer, logical layer, and application layer of the hierarchical digital twin model, the rationality of each reasoning path is verified. Based on the reasoning path of potential causes after rationality verification, the corresponding fault model is matched in the fault knowledge graph, and the corresponding diagnostic knowledge is extracted from the graph based on the selected fault model. Based on the extracted diagnostic knowledge, supplementary data is retrieved, and fuzzy matching is performed between the supplementary data and the standard diagnostic data in the fault knowledge graph to calculate the matching similarity of the supplementary data. Based on the confidence propagation algorithm, combined with the similarity of the supplementary data and the weights of the abnormal indicators in the fault analysis dataset, the confidence of the current fault cause is calculated. If the confidence level is lower than the preset confidence threshold, the range of supplementary data retrieval will be adjusted and rematched. If the confidence level is higher than the preset confidence threshold, the search will continue to the next level along the current potential cause based on the linkage relationship between the physical layer, logical layer and application layer of the hierarchical digital twin model. Based on the fault analysis dataset, the matching degree of each potential cause after confidence verification is calculated, and the potential cause with the highest matching degree is selected. Obtain the corresponding layered data from the visual monitoring interface, perform entity verification on the screened potential causes, confirm that the potential causes can explain all abnormal phenomena, and locate the unique cause of the failure. The single cause of failure output is input into the hierarchical digital twin model to reproduce the failure chain. Based on the reproduction results, the causal relationship between the cause of failure and the abnormal phenomenon is verified. Based on the reproduced fault chain, combined with the hierarchical correlation data displayed on the visual monitoring interface, the impact of the fault cause on the physical layer, logic layer, and application layer and the transmission path are analyzed to generate a fault impact chain report. Obtain the correlation between causes and adaptive self-healing strategies in the fault knowledge graph, analyze the blocking coverage of each candidate self-healing strategy on the fault impact chain in conjunction with the fault impact chain report, calculate the adaptability of the candidate strategies, and synchronize the calculation results to the visualization monitoring interface.
Citation Information
Patent Citations
Engineering equipment fault intelligent diagnosis and self-healing control system and method based on digital twinning
CN111596604A
Fault root cause analysis method and device and network equipment
CN116866149A
Simulation deduction system based on digital twinning
CN117150757A
Intelligent operation and maintenance monitoring method and system for data center
CN120602308A
Network equipment dynamic operation and maintenance system and method based on artificial intelligence
CN120856557A
Cited By
Unmanned measurement and control communication system autonomous decision optimization method based on digital twinning
CN122239493A
A smart cloud warehouse risk evolution evaluation system and method based on digital twinning
CN122248042A
A smart cloud warehouse risk evolution evaluation system and method based on digital twinning
CN122248042B