Cloud-side collaborative digital platform based on integrated management and control of regional production and operation

By constructing a cloud-edge collaborative digital platform and utilizing dynamic knowledge graphs and infection source propagation simulation models, the problem of insufficient regional-level intelligent diagnosis and collaborative training capabilities in existing technologies has been solved. This enables accurate assessment and propagation simulation of fault risks, thereby improving the safety and operational efficiency of regional production and operations.

CN121967492APending Publication Date: 2026-05-01YUNNAN HUADIAN FUXIN ENERGY POWER GENERATION CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
YUNNAN HUADIAN FUXIN ENERGY POWER GENERATION CO LTD
Filing Date
2025-12-27
Publication Date
2026-05-01

Smart Images

  • Figure CN121967492A_ABST
    Figure CN121967492A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of digital management, and discloses a cloud-side collaborative digital platform based on regional production and operation integrated management and control, which comprises an infrastructure layer used for constructing a unified hardware resource pool and an automatic operation and maintenance base for supporting a cloud-side collaborative architecture, global unified management and automatic operation and maintenance of a bottom server, storage and network hardware equipment are realized; the platform layer is used for converging global data and integrating an application development module and an AI module to form a common tool chain and data sharing service covering low-code development, micro-service and model full-life-cycle management; and the application layer is used for calling the common technical capability and the data sharing service provided by the platform layer, and constructing and deploying a business application system serving for integrated management and control of regional production and operation. According to the invention, a single-point operation and maintenance mode is converted into a new integrated intelligent management and control mode with cross-domain cooperation, active early warning and accurate decision making, and the safety, operation efficiency and risk resistance of regional production and operation are improved.
Need to check novelty before this filing date? Find Prior Art

Description

A cloud-edge collaborative digital platform based on integrated regional production and operation management. Technical Field

[0001] This invention relates to the field of digital management technology, and more specifically, to a cloud-edge collaborative digital platform based on integrated regional production and operation management. Background Technology

[0002] Integrated regional production and operation management refers to the use of a unified digital platform to aggregate data, coordinate status, and optimize decisions for dispersed production units, equipment systems, and operational processes within a specific geographical or logical region. This upgrades the management model from isolated operation to overall coordination, and from local optimization to global optimization. The core purpose of building such a cloud-edge collaborative digital platform is to break down data silos, unify the management view, and leverage cloud computing power and model capabilities combined with real-time edge response capabilities to achieve real-time perception, intelligent analysis, collaborative scheduling, and proactive optimization of the regional production and operation status.

[0003] However, existing technological solutions, while achieving the above objectives, suffer from the drawback of failing to effectively integrate truly regional-level intelligent diagnostic and collaborative training capabilities. Existing technologies largely remain at the device or subsystem level in terms of monitoring and simple alarms, and their analytical models are often static, isolated, and generalized. They lack a dynamic knowledge graph built from a regional global perspective, failing to depict the risk transmission relationships between devices, and thus cannot achieve cross-device, cross-level source tracing and propagation prediction of fault risks. Furthermore, model training and updates rely on centralized historical data, failing to fully utilize real-time operating data generated at the edge for rapid, personalized model optimization. This results in lagging diagnostic capabilities, poor adaptability, and difficulty in supporting accurate regional collaborative decision-making and preventative maintenance.

[0004] No effective solutions have yet been proposed to address the problems in the relevant technologies. Summary of the Invention

[0005] To address the problems in related technologies, this invention proposes a cloud-edge collaborative digital platform based on integrated regional production and operation management, in order to overcome the aforementioned technical problems existing in the existing related technologies.

[0006] Therefore, the specific technical solution adopted by the present invention is as follows:

[0007] According to a first aspect of the present invention, a cloud-edge collaborative digital platform based on integrated regional production and operation management is provided, the platform comprising:

[0008] The infrastructure layer is used to build a unified hardware resource pool and automated operation and maintenance base to support the cloud-edge collaborative architecture, and to achieve unified management and automated operation and maintenance of the underlying server, storage and network hardware devices.

[0009] The platform layer connects with the infrastructure layer. As the capability hub of the cloud-edge collaborative architecture, the platform layer is used to aggregate full-domain data and integrate application development modules and AI modules to form a common toolchain and data sharing services covering low-code development, microservices, and model lifecycle management.

[0010] The application layer interacts and connects with the platform layer to call the common technical capabilities and data sharing services provided by the platform layer, build and deploy a business application system that serves the integrated management and control of regional production and operation, and realize a closed loop from global optimization decision-making to edge execution through the cloud-edge collaboration mechanism.

[0011] Preferably, the application layer includes:

[0012] The regional-level hierarchical diagnostic unit is used to extract the operating status data of all production equipment from the platform layer, and based on the infection source transmission simulation model trained by the AI ​​module, to conduct hierarchical assessment of the fault risk of each production equipment, and generate fault risk assessment results to be transmitted to the intelligent inspection unit.

[0013] The intelligent inspection unit is used to prioritize inspection tasks based on the results of fault risk assessment, plan the optimal inspection path under the conditions of full coverage, resource constraints and cost constraints, and generate the optimal inspection strategy by combining the inspection task priorities and the optimal inspection path.

[0014] The regional power prediction unit is used to call the long short-term memory network model trained by the AI ​​module to realize short-term and ultra-short-term power prediction of distributed power sources and loads within the region, and transmit the power prediction results to the regional power market management and control unit.

[0015] The regional-level power market management and control unit is used to make centralized optimization decisions on the declaration, clearing and settlement of adjustable resources participating in market transactions within the region, based on power forecast results and predefined power market management and control rules, and to generate transaction plans and dispatch instructions that are transmitted to the regional collaborative control unit.

[0016] The regional collaborative control unit is used to aggregate transaction plans and scheduling instructions, fault risk assessment results and optimal inspection strategies, formulate collaborative control instructions for edge production units, and issue collaborative control instructions to edge production units for execution through cloud-edge collaborative mechanisms.

[0017] Preferably, the regional-level hierarchical diagnostic unit, when extracting operational status data of all production equipment from the platform layer and, based on the infection source transmission simulation model trained by the AI ​​module, performs a hierarchical assessment of the fault risk of each production equipment and generates a fault risk assessment result which is then transmitted to the intelligent inspection unit, includes:

[0018] The operating status data, historical fault records, and related operating condition data of each production equipment are extracted from the data base of the platform layer and preprocessed to obtain the structured features of the production equipment.

[0019] Based on predefined production equipment topology and physical connections, a knowledge graph is constructed with each production equipment as a node.

[0020] Using the structured features of production equipment as input, the node core degree hierarchical stripping algorithm is called to calculate the importance of each node. Based on the node importance, risk source nodes, primary susceptible nodes, and several secondary susceptible nodes used for fault risk propagation simulation in the associated knowledge graph are identified.

[0021] Using the risk source node as the starting point of fault propagation, the infection source propagation simulation model is invoked to simulate the fault risk transmission link from the risk source node to the primary susceptible node and each secondary susceptible node in the associated knowledge graph.

[0022] Calculate the failure risk value of each path node in the failure risk propagation link under the propagation resistance of the associated node, and use it as the failure risk value of the corresponding production equipment; evaluate the failure risk level of each production equipment based on the failure risk value of each production equipment, and generate a failure risk assessment report including the failure risk level.

[0023] Preferably, the importance of each node is calculated using a node core degree hierarchical stripping algorithm. Based on the node importance, risk source nodes, primary susceptible nodes, and several secondary susceptible nodes used for fault risk propagation simulation in the associated knowledge graph are identified, including:

[0024] The calculation includes risk propagation characteristic parameters such as the diffusion right of each node in the associated knowledge graph that triggers the ability of downstream nodes, the exposure right triggered by upstream nodes, and the weight balance factor that comprehensively reflects the balance and activity of nodes in the propagation process.

[0025] The initial importance of each node is calculated based on the risk propagation characteristic parameters, and the initial importance is used as the initial importance function for the node iterative decomposition of the node core degree hierarchy stripping algorithm.

[0026] Based on the initial importance function, the nodes with the smallest initial importance value and all their associated edges in the associated knowledge graph are iteratively removed. The nodes to be removed are assigned to different levels according to the removal order, until all nodes in the associated knowledge graph have been removed and the level assignment is completed, forming the initial node level sequence.

[0027] Select nodes at the same level and introduce an iteration factor. Using the relative order in which the current node is removed in the same level, the total number of nodes in the level, and the importance values ​​of adjacent levels as parameters, calculate the intra-level calibration value of the current node. Sum the level number of the current node with the intra-level calibration value to obtain the node importance that comprehensively reflects the global and local importance of the node.

[0028] The nodes are sorted in descending order of importance, and the nodes in the associated knowledge graph are classified and identified by combining the diffusion weight and exposure weight of each node, so as to obtain the risk source node, the primary susceptible node and several secondary susceptible nodes for fault risk propagation simulation.

[0029] Preferably, the nodes are sorted in descending order of importance, and the nodes in the associated knowledge graph are classified and identified by combining the diffusion weight and exposure weight of each node, resulting in risk source nodes, primary susceptible nodes, and several secondary susceptible nodes for fault risk propagation simulation, including:

[0030] The nodes are sorted in descending order of absolute importance. The node with the highest importance in the sorted list and whose diffusion weight is greater than its exposure weight is used as the risk source node for fault risk propagation simulation.

[0031] The node with the highest node importance in the sorted list and whose exposure weight is greater than its diffusion weight is selected as the primary susceptible node for fault risk propagation simulation.

[0032] The remaining nodes in the sorted list are used as secondary susceptible nodes for fault risk propagation simulation.

[0033] Preferably, taking the risk source node as the starting point of fault propagation, the infection source propagation simulation model is invoked to simulate the fault risk transmission link from the risk source node to the primary susceptible node and each secondary susceptible node in the associated knowledge graph, including:

[0034] The initial state of the risk source node in the associated knowledge graph is marked as the infected state and used as the initial infection source for the fault risk propagation simulation. The initial state of the remaining nodes is marked as the susceptible state.

[0035] The fault risk propagation evolution process is executed iteratively based on discrete time steps, and in each iteration, the state of each node is traversed step by step starting from the initial source of infection.

[0036] If the current node is in a susceptible state, the risk probability of the current node being infected is calculated by combining the directed weights of the current node, all upstream nodes in the infected state, and the connecting edges; if the current node is in an infected state, the recovery probability of the infected state being converted to the recovered state is determined according to the preset recovery probability; if the infected state is converted to the recovered state, the current node exits the propagation chain.

[0037] The status of all nodes is updated synchronously. Susceptible nodes that reach the risk probability are updated to the infected state; infected nodes that reach the preset recovery probability are updated to the recovery state, so as to drive the failure risk to spread from the infected nodes to the downstream nodes in each step, realizing the chain propagation of failure risk from the risk source node to the primary susceptible node, and then from the newly infected primary susceptible node to the secondary susceptible node.

[0038] In each iteration, the state transition event between the corresponding source node and the infected node is recorded, and the process stops when the number of iterations is satisfied or there are no infected state nodes in the associated knowledge graph.

[0039] The state transition events recorded in each iteration are summarized to obtain the propagation event log, and the fault risk transmission chain is extracted based on the propagation event log.

[0040] Preferably, the state transition events recorded in each iteration are summarized to obtain a propagation event log, and the fault risk transmission chain is extracted based on the propagation event log, including:

[0041] The state transition events recorded in each iteration are summarized to obtain the propagation event log. All initial propagation events that are directly caused by the initial risk source node as the source of infection are filtered out from the propagation event log.

[0042] The initial propagation events are integrated with the associated nodes according to the time step of occurrence to form a primary transmission link of failure risk from the risk source node to the main susceptible node;

[0043] Using the primary susceptible nodes that have already been infected in the primary transmission chain as new sources of infection, the secondary transmission events triggered by the new sources of infection are further filtered out in the transmission event log;

[0044] Secondary propagation events are integrated according to their time steps and associated nodes to form a secondary transmission link of fault risk from the primary susceptible node to the secondary susceptible node;

[0045] By integrating the primary and secondary transmission links of fault risk, a comprehensive fault risk transmission link is obtained.

[0046] Preferably, when the intelligent inspection unit prioritizes inspection tasks based on fault risk assessment results, plans the optimal inspection path under the conditions of full coverage, resource constraints, and cost constraints, and generates the optimal inspection strategy by combining the inspection task priorities and the optimal inspection path, it includes:

[0047] Based on the priority of generating initial inspection tasks;

[0048] Analyze the influence relationship between each production equipment and related equipment in the fault risk transmission chain, and use the betweenness centrality algorithm to identify attenuating equipment that blocks the transmission of fault risk and amplifying equipment that expands the scope of risk;

[0049] When a sudden change signal in the status of production equipment is detected, the scope of the change is analyzed based on the fault risk transmission link corresponding to the production equipment, and the topological importance of the attenuating and amplifying equipment is integrated into the scope of the change, and the priority of the inspection task is updated.

[0050] The updated inspection task priority is passed to the preset path planning and optimization component through the microservice provided by the application development module. The preset path planning and optimization component transforms the updated inspection task priority into a benefit item in the path planning. It uses the geographical location of production equipment, the status of available inspection resources, the cost of various resources, and full coverage as multi-level constraints, and takes the lowest total execution cost or the highest total efficiency as the optimization objective to build a path optimization model.

[0051] Solve the path optimization model to obtain the optimal inspection strategy that prioritizes the attenuation and amplification devices on the fault risk propagation link; distribute the optimal inspection strategy to the edge inspection terminal for execution, and transmit the real-time data generated during the execution back to the data base of the platform layer.

[0052] According to a second aspect of the present invention, a computer device is provided. The computer device includes a memory and a processor, the memory storing a computer program, the computer program being executed by the processor, which is a cloud-edge collaborative digital platform based on integrated regional production and operation management.

[0053] According to a third aspect of the present invention, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, which is executed by a processor, representing a cloud-edge collaborative digital platform based on integrated regional production and operation management.

[0054] The beneficial effects of this invention are as follows:

[0055] 1. This invention addresses the fundamental shortcomings of existing technologies in achieving regional-level collaboration and predictive management through a cloud-edge collaborative, data-driven, and intelligent closed-loop architecture. Specifically, the platform provides a stable foundation for collaboration by achieving unified management and elastic supply of resources at the infrastructure layer; it aggregates data from across the entire domain and integrates AI and development toolchains at the platform layer to form a reusable intelligent capability hub; and finally, at the application layer, it achieves hierarchical diagnosis, propagation simulation, and impact assessment of regional-level faults based on dynamic knowledge graphs and networked risk models, driving the generation of optimal inspection strategies that integrate global risk situations and resource constraints. This transforms the traditional discrete, passive, single-point operation and maintenance model into a new integrated intelligent management and control model that features cross-domain collaboration, proactive early warning, and precise decision-making, significantly improving the overall security, operational efficiency, and risk resistance capabilities of regional production and operation.

[0056] 2. This invention constructs a dynamic device association knowledge graph and integrates an improved node core degree hierarchical stripping algorithm. This enables the automatic and accurate identification of critical devices as risk sources, primary susceptible devices as transmission hubs, and secondary susceptible devices constituting the propagation path from a global perspective of network topology. Furthermore, by invoking an infection source simulation model based on propagation dynamics, the invention dynamically deduces and visualizes the critical path of fault propagation from the risk source node to the entire network in the knowledge graph. This not only achieves quantitative classification of individual device risks but also systematically reveals the risk propagation chain and global impact. Ultimately, it provides decision-making insights and execution basis for regional integrated intelligent management and control, from predictive maintenance and precise emergency intervention to resource optimization and scheduling.

[0057] 3. This invention deeply integrates network topology analysis into operation and maintenance decision-making, achieving a fundamental shift from passive response to proactive intervention and from single-point inspection to systemic prevention and control. It utilizes algorithms such as betweenness centrality to identify strategically valuable attenuation and amplification devices in the transmission chain, enabling precise identification of key nodes that can block risk propagation or exacerbate risk spread when formulating inspection strategies. Furthermore, by integrating a path optimization model that incorporates task benefits, geographic information, resource status, and cost constraints, it generates feasible inspection strategies that ensure priority and efficient execution of critical equipment. This maximizes the overall effectiveness of risk prevention and control with limited operation and maintenance resources, significantly improving the initiative and intelligence level of regional operation and maintenance work. Attached Figure Description

[0058] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0059] Figure 1 is an architecture diagram of a cloud-edge collaborative digital platform based on integrated regional production and operation management according to an embodiment of the present invention;

[0060] Figure 2 is a schematic diagram illustrating the specific implementation of generating fault risk assessment results in a regional-level hierarchical diagnostic unit of a cloud-edge collaborative digital platform based on integrated regional production and operation management according to an embodiment of the present invention.

[0061] Figure 3 is a schematic diagram of the computer equipment;

[0062] Figure 4 is a schematic diagram of the platform layer in a cloud-edge collaborative digital platform based on integrated regional production and operation management according to an embodiment of the present invention.

[0063] Figure 5 is a schematic diagram of the application layer in a cloud-edge collaborative digital platform based on integrated regional production and operation management according to an embodiment of the present invention.

[0064] In the picture:

[0065] 1. Infrastructure Layer; 2. Platform Layer; 201. Data Foundation; 202. Application Development Module; 203. AI Module; 3. Application Layer; 301. Regional-level Hierarchical Diagnostic Unit; 302. Intelligent Inspection Unit; 303. Regional-level Power Prediction Unit; 304. Regional-level Electricity Market Management and Control Unit; 305. Regional Collaborative Control Unit. Detailed Implementation

[0066] To further illustrate the various embodiments, the present invention provides accompanying drawings, which are part of the disclosure of the present invention. These drawings are mainly used to illustrate the embodiments and can be used in conjunction with the relevant descriptions in the specification to explain the operating principles of the embodiments. With reference to these drawings, those skilled in the art should be able to understand other possible implementation methods and the advantages of the present invention.

[0067] According to an embodiment of the present invention, a cloud-edge collaborative digital platform based on integrated regional production and operation management is provided.

[0068] The present invention will now be further described with reference to the accompanying drawings and specific embodiments. As shown in Figure 1, according to an embodiment of the present invention, a cloud-edge collaborative digital platform based on integrated regional production and operation management includes:

[0069] Infrastructure Layer 1 is used to build a unified hardware resource pool and automated operation and maintenance base to support the cloud-edge collaborative architecture, and to achieve unified management and automated operation and maintenance of the underlying server, storage and network hardware devices.

[0070] It should be noted that Infrastructure Layer 1 includes modules for device management, network device management, storage device management, resource management, portal management, overview management, robust operation management, maintenance management, and comprehensive management.

[0071] Server management; resource allocation and planning: based on demand and resource management strategies, allocate and plan resources for physical servers, determining the processing power, memory capacity, storage space, and network bandwidth of each server. Automated deployment and configuration: utilize automation tools and technologies to achieve rapid deployment and configuration of physical servers, including automatically creating and configuring servers using predefined images or templates, reducing manual setup time. Monitoring and performance management: monitor the performance and availability of physical servers. Use monitoring tools to monitor server operating status, load, network traffic, and other key metrics in real time. This allows for timely detection and resolution of potential problems and ensures the normal operation of servers.

[0072] Security Management: Physical server security is a crucial aspect of cloud platform management. Various security measures are implemented to protect servers from unauthorized access, data breaches, and other threats, including access control, firewalls, encryption technologies, and security auditing. Maintenance and Updates: Regular maintenance and updates to physical servers are performed, including operating system and software patch updates, hardware replacement, performance optimization, and other necessary maintenance tasks. Disaster Recovery and Backup: Disaster recovery and backup strategies are implemented to ensure the reliability and durability of physical servers and hosted data, including redundancy setups, backup replication, and disaster recovery plans. Troubleshooting and Technical Support: Responsible for troubleshooting and providing technical support when physical servers malfunction or experience problems. Providing responses and solutions to minimize server downtime and service interruptions.

[0073] Network device management; network topology planning, designing and planning the network topology of the cloud platform. This includes determining the location, connection methods, and network layout of network devices to meet requirements and performance needs. Automated configuration and management, using automated tools and technologies to configure and manage network devices. This includes using network configuration management tools to automate device configuration, port management, and routing settings. Network monitoring and troubleshooting, implementing a network monitoring system to monitor network devices in real time. This includes monitoring device status, bandwidth utilization, network latency, and packet loss. If a fault or problem occurs, performing troubleshooting and repair operations to restore normal network operation.

[0074] Security management, particularly the security of network devices, is a crucial aspect of cloud platform management. Security measures are implemented to protect network devices from unauthorized access, intrusion, and other cybersecurity threats. These measures include access control, firewalls, intrusion detection and prevention systems, and more.

[0075] Network performance optimization involves optimizing network equipment performance to ensure high-speed, reliable network connectivity. This includes strategies such as bandwidth management, load balancing, network optimization, and traffic control. Network expansion and resilience are implemented based on demand. This involves adding, removing, or reconfiguring network equipment as needed to meet the requirements of different scales and workloads. Network policy and security auditing involves developing network policies and conducting security audits to ensure that network equipment meets security standards and compliance requirements, including reviewing network configurations, access permissions, and logging.

[0076] Storage device management; storage resource allocation, allocating resources to storage devices based on demand and resource management strategies. This includes determining the capacity, performance requirements, and access control policies for each storage device. Storage management and configuration, using storage management tools and technologies to configure and manage storage devices. This includes creating storage volumes, allocating storage space, setting access permissions, and storage policies. Data backup and recovery, implementing data backup and recovery strategies to ensure the reliability and durability of data on storage devices, including regular data backups, snapshots, data redundancy, and disaster recovery plans. Storage performance optimization, optimizing storage device performance to provide efficient and fast data storage and retrieval. Utilizing technologies such as high-speed storage media, storage caching, data compression, and tiered storage. Storage monitoring and troubleshooting, monitoring the performance and availability of storage devices. Using monitoring tools to monitor the status, capacity utilization, I / O load, etc., of storage devices in real time. Troubleshooting and repairing any failures or problems to ensure the normal operation of storage devices.

[0077] Storage security management is a crucial aspect of cloud platform management. Various security measures are implemented to protect storage devices from unauthorized access, data breaches, and other threats. These include access control, encryption, authentication, and security auditing.

[0078] Storage capacity expansion and elasticity allow for scaling and adjusting storage capacity as needed. Cloud services can add, remove, or reconfigure storage devices as required to meet varying scales and storage demands.

[0079] The cloud management platform simplifies the overall platform management model and enables elastic scaling of business applications. Simultaneously, it needs to support planned and on-demand backups, enabling rapid recovery of backed-up virtual machine states and storage resources. This is because, in a cloud-edge collaborative architecture, when edge sites simultaneously handle local persistent writes and the cloud-edge link experiences minute-level network partitions or intermittent jitter, existing planned or on-demand backups combined with snapshots for rapid recovery of virtual machine states and storage resources will face typical cross-domain consistency issues. Specifically, this manifests as: the snapshot object coverage having unclosed consistency domains. Virtual machine-level snapshots typically only guarantee the consistency of a virtual disk image at a specific point in time, while the persistent states of business-side databases, object storage, message queues, or edge local caches belong to different write domains. During recovery, multi-source state time point drift is easily formed; that is, the VM image rollback point, database commit point, and object or block storage snapshot point are not equal, leading to logical inconsistencies in the recovered business. To address this, this invention provides a cloud-edge collaborative digital platform based on integrated regional production and operation management, which also includes:

[0080] The cross-cloud-edge consistency check unit is used to establish a consistency checkpoint triggering mechanism for write resources at the platform layer in a cloud-edge collaborative architecture. When a backup is triggered, a checkpoint identifier is generated and a checkpoint triggering command is issued.

[0081] Specifically, in the cloud-edge collaborative architecture, a consistency checkpoint triggering mechanism is established at the platform layer for all resource objects participating in persistent writing, upgrading a backup or snapshot operation from a single resource behavior to a joint action across resources. When a planned backup or on-demand backup is triggered, a checkpoint identifier is uniformly generated by the cloud, and a checkpoint triggering command is issued to all associated write units, requiring each write unit to enter the checkpoint preparation state.

[0082] The write freeze and sequential processing unit is used to divide the write resources before and after the checkpoint marker within the freeze window in response to the checkpoint trigger command, and generate a freeze confirmation set.

[0083] Specifically, upon receiving the checkpoint identifier, the edge site and the cloud-side write unit simultaneously enter the write freeze and sequentialization phase, strictly defining the boundaries of writes before and after the checkpoint and generating a freeze confirmation set. During the boundary definition process, within the freeze window, all newly generated business writes need to be redirected to the local sequential write log, and a forced flush operation is performed on the persistent media to ensure that the write state before the checkpoint is completely written to disk, while the writes after the checkpoint identifier are explicitly marked as the post-checkpoint sequence.

[0084] The cross-resource consistency snapshot unit is used to perform snapshot operations on write resources associated with the frozen confirmation set, and bind the federated snapshot set for which the snapshot operation is performed to a unified checkpoint identifier.

[0085] Specifically, snapshot operations such as virtual machine snapshots, block storage snapshots, and object storage snapshots are only performed after all relevant write units have returned a frozen confirmation set with associated checkpoint identifiers, and these snapshots are uniformly bound to the same checkpoint identifier. During this process, various snapshots no longer exist as independent resource snapshots, but rather as a joint snapshot set under the same consistency domain. Their validity is determined by whether they completely cover all writes before CP_ID, and includes the following steps:

[0086] Step S1: Confirm the aggregation freeze and generate a consistency domain list and checkpoint parameter package.

[0087] Specifically, the freeze confirmation records returned by each write unit are normalized and aggregated. The freeze confirmation set includes at least: the write unit identifier, the freeze timestamp, the local write sequence number at the freeze time, and a freeze success flag (a value of 1 indicates a successful freeze, and a value of 0 indicates a failed freeze). If any write resource has a freeze success flag of 0, the consistency domain cannot be closed, the output is empty, and refreezing is required; if all freezes are successful, a consistency domain list is constructed, and the freeze boundary sequence number of each write unit is packaged into a checkpoint parameter package.

[0088] To avoid checkpoint boundary drift caused by "some units freezing early and some freezing late", it is also necessary to calculate the freeze synchronization deviation for subsequent snapshot window control. The freeze synchronization deviation is equal to the difference between the maximum and minimum timestamps of a certain write unit completing freeze confirmation. The smaller the freeze synchronization deviation, the more synchronized the freeze is. The value range is 0.5-10 seconds, which is determined according to the on-site link jitter and write load.

[0089] Step S2: Lock the snapshot trigger window for each resource and calculate the snapshot order constraint.

[0090] Specifically, after freeze confirmation, the platform needs to define a snapshot trigger window to prevent an excessively long freeze period due to prolonged periods without snapshots. The snapshot trigger window starts at the latest freeze time (e.g., the maximum timestamp of a write unit completing freeze confirmation), and ends at the snapshot trigger window, which is the sum of the snapshot window start and the snapshot window width (a constant configuration item, ranging from 1 to 60 seconds, proportional to the resource size). Then, to ensure that write domain boundaries are solidified before the compute state is solidified, a snapshot order constraint needs to be generated. This constraint follows the rule: snapshots are executed on all persistent media (block storage, object storage, edge local persistent volumes) first, followed by snapshots on virtual machines or container compute states; if multiple virtual machines are mounted on shared storage, the shared storage snapshot must precede any virtual machine snapshot.

[0091] Step S3: Perform cross-resource snapshots and bind a unified checkpoint identifier to form a joint snapshot set.

[0092] Specifically, within the snapshot trigger window associated with the checkpoint identifier, block storage snapshots, object storage snapshots, edge local persistent volume snapshots, and virtual machine snapshots or container consistency snapshots are triggered sequentially according to the snapshot order constraint. Each type of snapshot returns a snapshot handle and snapshot time. Finally, the platform binds all snapshots to the same checkpoint identifier, thus forming a joint snapshot set. Furthermore, to ensure cross-resource temporal consistency of the joint snapshot set, the snapshot time dispersion needs to be calculated. This snapshot time dispersion is equal to the difference between the maximum and minimum values ​​of the actual completion timestamps of each resource snapshot, with a value range of 1-30 seconds. If this snapshot time dispersion exceeds a set threshold, it indicates that the current joint snapshot set is a weakly consistent snapshot, only allowed for emergency recovery and not for normal rollback.

[0093] Step S4: Before verifying the full coverage checkpoint identifier, write and output the valid marker for the consistent snapshot.

[0094] Specifically, the core of verifying whether the joint snapshot completely covers all writes before the checkpoint is to compare the freeze boundary sequence number (representing the last write sequence number allowed to fall into the pre-checkpoint domain at the freeze time) with the persistent boundary visible in the snapshot (the persistent sequence number visible within the snapshot). Both are non-negative integers, ranging from 0 to 10. 15 The coverage ratio is determined by the log system and its operational lifespan. First, for each write unit, its checkpoint alignment markers (e.g., the persistent sequence number of the write log, the disk flush sequence number of the database, and the commit sequence number of the object storage metadata) are read from the persistent media after the snapshot. Then, a coverage ratio is constructed, which is equal to the ratio of the sum of all coverage indicators to the total number of write units. The confirmation of the coverage indicator requires comparing the freeze boundary sequence number with the persistent boundary visible in the snapshot. If the latter is greater than or equal to the former, the coverage indicator is set to 1; otherwise, it is set to 0. When the constructed coverage ratio equals 1, it indicates that all write units have covered the writes before the freeze boundary in the snapshot, and a consistent snapshot valid flag equal to 1 is output. If the constructed coverage ratio is less than 1, a consistent snapshot valid flag equal to 0 is output, and a list of uncovered units is recorded. Finally, the coverage proof record is integrated and output for subsequent auditing and replay boundary selection in the recovery phase.

[0095] The strategic processing unit is used to load a federated snapshot set when cloud edge disconnection, node failure, or manual recovery operation occurs, and to replay or discard written resources after the checkpoint mark, using the checkpoint mark as the boundary.

[0096] Specifically, when a cloud-edge connection failure, node failure, or manual triggering of a recovery operation occurs, the most recent snapshot is not directly restored. Instead, the aforementioned combined snapshot set is first loaded, and written records after the checkpoint are processed strategically, using the checkpoint identifier as the boundary. For example, only write sequences that meet the complete commit conditions are replayed, or incomplete writes after the checkpoint identifier are directly discarded when a missing write context is detected, thereby ensuring that the recovered system state is logically consistent.

[0097] The resource management module mainly provides management functions such as resource application and modification, and provides convenient and comprehensive resource management functions through a unified management method; the operation center module realizes management functions such as resource quota; the operation and maintenance center collects operation and maintenance information in real time and realizes comprehensive analysis and presentation; the management center module provides management functions such as organization, role, resource pool, and process, and provides users with management functions such as user permissions.

[0098] The cloud management platform connects to different resource pools and cloud services, providing unified service delivery and service assurance functions; it provides open APIs to the northbound side, which can be called by third-party operation / maintenance / application systems.

[0099] Portal Management: The cloud management platform provides a service portal interface for operations and maintenance personnel. The service portal has functions such as browsing service catalogs, submitting service requests, querying instances, making changes and terminating requests, controlling and statistically analyzing deployed resources. The cloud management platform also provides a management portal interface for operations and maintenance personnel. The management portal enables functions such as global system management, service design and changes, service content review and approval, and service resource opening and auditing.

[0100] Overall Management: The cloud management platform provides a comprehensive overview of resources and services, including real-time overview and management of resource application information, system operation information, and resource object details. Resource application information includes the application process, pending tasks, and application-related statistical overviews. System operation information includes the usage and unused status of virtual CPUs, virtual memory, and storage, as well as the power-on and power-off status of physical hosts and virtual machines. Specific resource objects include resource pools, clusters, hosts, storage, networks, virtual machines, virtual machine images, virtual machine cloud disks, and virtual machine clusters.

[0101] Strong operational management; the cloud management platform possesses the capability to manage the overall cloud service resources and permissions, ensuring the secure and efficient operation of the cloud platform. This includes virtual machine permission configuration, cloud disk configuration, network configuration, virtual machine cluster configuration, orchestration service configuration, and workflow configuration functions, reducing the overall operational threshold for operators and improving operational efficiency.

[0102] Resource Management: The cloud management platform manages hardware such as servers, storage, and networks based on mainstream domestically produced chips. It supports functions such as host cluster management, physical host management, virtual machine cluster management, virtual machine management, image management, key management, instance type management, storage management, cloud disk management, and network management. It achieves unified management of all resources and real-time monitoring of hardware status.

[0103] Operations and maintenance management; the cloud management platform provides overall resource operations and maintenance management capabilities, including real-time monitoring of computing resource information, storage resource information, network resource information, etc. It also allows for the review of historical resource usage, provides comprehensive information filtering functions, and reduces the difficulty of troubleshooting.

[0104] The operations and maintenance management module has the capability to customize the configuration of various events and event sources on the search platform, configure event thresholds, manage tasks, and perform information statistical analysis. It ensures customized configuration monitoring conditions under different operating scenarios, achieving multi-dimensional and reliable overall operations and maintenance functions.

[0105] Comprehensive management; the cloud management platform's specific comprehensive management capabilities include user management, account management, user information management, organizational structure management, role management, environment management, resource status management, professional title and position management, etc.

[0106] Virtual Resources: Resource virtualization runs on top of Infrastructure Layer 1, enabling virtualized management of hardware resources below and shielding the differences between heterogeneous underlying hardware; it provides API interfaces for unified management and scheduling by the cloud platform above, eliminating the upper-layer GuestOS's dependence on hardware devices and drivers, and enhancing hardware compatibility, high reliability, high availability, scalability, and performance optimization in the virtualized runtime environment.

[0107] Computing virtualization, based on open-source kernel virtualization technologies such as KVM, meets the requirements of independent control, with a code self-reliance rate exceeding 80%. By integrating infrastructure resources such as data center servers, it improves the utilization of physical resources and enhances management efficiency. The cloud host operating system can run directly in KVM virtual machines without modification, and each virtual machine can enjoy independent virtual hardware resources, such as network cards, disks, and graphics adapters.

[0108] Storage virtualization logically maps underlying storage resources into a single, transparent storage solution with a single view, eliminating the need for users to worry about the underlying physical storage devices. It supports elastic scaling of distributed storage, improving overall efficiency, resolving storage space waste, and consolidating fragmented storage into a contiguously addressed logical storage space. This overcomes the capacity limitations of a single physical disk, automatically reallocating data and utilizing efficient snapshot technology to reduce capacity requirements during storage pool expansion, thereby improving storage resource utilization.

[0109] An automated, reliable, distributed data persistence layer provides an infinitely scalable storage cluster. All user data stored in the system is ultimately stored by this layer. High reliability, high scalability, high performance, and high automation are essentially provided by this layer.

[0110] Network virtualization; supports high-quality virtual switches (OVS) and the Data Plane Development Kit (DPDK), improving forwarding performance and reliability, enhancing data processing performance and throughput, and supporting high-performance packet processing capabilities in network applications. Virtualization, combined with DPDK technology, performs functional modifications and performance acceleration optimizations on OVS at the virtualization layer. Packets received from a network port connected to OVS no longer need to undergo kernel-level processing; they directly reach user space via the DPDKPMD driver. Through lock-free ring queues, memory large pages, flow classification, and multi-core affinity technologies, data path performance is optimized, accelerating packet processing speed on both physical and virtual network interfaces.

[0111] Platform Layer 2, connected to Infrastructure Layer 1, serves as the capability hub of the cloud-edge collaborative architecture. It is used to aggregate full-domain data and integrate application development module 202 and AI module 203 to form a common toolchain and data sharing service covering low-code development, microservices, and model lifecycle management.

[0112] In one embodiment, as shown in FIG4, platform layer 2 includes:

[0113] The data foundation 201 is used for data aggregation, data storage, and IoT management of multi-source heterogeneous data from infrastructure layer 1 and application layer 3, while providing standardized data sharing services and access interfaces; the application development module 202 is used to provide low-code development tools and microservice frameworks for application layer 3, supporting visual orchestration, service governance, and deployment; the AI ​​module 203 is used to provide full lifecycle management tools covering model development, model training, and model management, providing common technical support for application layer 3.

[0114] The cloud platform services include: computing services; the cloud platform is fully compatible with mainstream domestic chip architectures such as x86 and ARM, and supports the deployment of resource pools for multiple architectures within the same region to meet diverse computing needs. The computing resource pools mainly provide multi-engine computing services such as cloud hosts, containers, and bare metal, supporting different architecture types of business applications and achieving unified scheduling and management of computing resources. Through a unified visual management interface, it enables flexible and on-demand allocation of computing resources, rapid construction of stable and reliable computing resource blocks, and full lifecycle management of computing resources, while simplifying the complexity of operation and maintenance management and improving the utilization rate of computing resources.

[0115] Cloud server service; provides cloud server management functions to meet the computing resource application needs of various business applications and quickly build computing resources to meet application business requirements.

[0116] Bare metal services; the cloud platform provides two types of integration capabilities: elastic bare metal and ordinary bare metal, and implements different levels of integration methods according to different scenarios.

[0117] Container services; the cloud platform supports container services through unified account and permission management. Container cluster management nodes and compute nodes support deployment on physical machines as well as virtual machines. By integrating basic computing components such as physical servers, smart network interface cards, operating systems, and cloud disks, it achieves high-performance bare-metal cloud server resource management.

[0118] Storage services; the cloud platform provides elastic block storage resources for cloud server instances through cloud disks, achieving high reliability, flexibility, ease of use, and elastic scaling, while also supporting snapshot functionality. It offers complete cloud storage lifecycle management capabilities. Cloud disk management provides capabilities such as creating, deleting, mounting to cloud hosts, unmounting, querying, and scaling cloud disks. Administrators can define different cloud disk types and identify different backend resource pools through tags, such as SSD high-performance storage, SATA large-capacity storage, and centralized storage. Different types of cloud disks can be selected through the service interface to meet the storage needs of different business scenarios.

[0119] Cloud disk snapshots can quickly save data copies at any point in time, and cloud disk data can be restored from any snapshot point. It supports multi-level snapshots, allowing for rapid saving of data storage status in the cloud disk and quick data rollback in case of configuration errors. It also supports full snapshots and incremental snapshots, reducing the risk of data loss through backups. It supports creating cloud disk snapshot policies, specifying the snapshot execution cycle and snapshot save time, enabling scheduled saving of cloud disk data. It provides cloud disk sharing functionality, allowing cloud disks to be mounted to multiple cloud servers simultaneously. It offers cloud disk configuration change functionality, allowing users to change the cloud disk type or perform expansion operations. It provides cloud disk cloning functionality, allowing data to be copied from a cloud disk to a new cloud disk. It provides data disk image creation functionality, allowing data to be exported from a data disk and placed into a custom image. It also provides system disk image creation functionality, allowing data to be exported from the system disk and placed into a custom image.

[0120] Network services; the cloud platform achieves high reliability, high speed, high stability and other high availability characteristics by building a variety of internal networks to balance network resources, while supporting basic cloud collaborative network capabilities.

[0121] Virtual Private Clouds (VPCs) build isolated and private virtual network environments for cloud resources such as cloud servers, providing network functions including subnet creation, security group and network ACL configuration, routing table management, and requesting elastic public IP addresses and bandwidth. They construct virtualized network infrastructure based on various network elements that simulate real-world data centers, meeting the need for rapid planning of cloud network environments. VPC networks have the ability to autonomously manage various network configurations, including private IP address ranges, subnet segments, DNS, and inter-network routing rules to achieve network access control. Furthermore, VPCs can utilize various types of services such as cloud hosts, cloud disks, and virtual load balancers.

[0122] It provides load balancing services to achieve balanced distribution of network traffic for application loads. By distributing received traffic to a designated group of cloud hosts according to load balancing strategies, it meets the scalability requirements of business application systems and effectively copes with high-concurrency load pressure through horizontal scaling of resources. It hides the actual service ports, enhances internal system security, and meets the business requirements of high concurrency and high availability.

[0123] It supports both IPv4 and IPv6 protocol stacks for virtual machines, enabling communication with nodes supporting both IPv4 and IPv6 protocols. IPv4 and IPv6 network addresses within the virtual machine operating system can be manually specified or configured and distributed via the cloud platform. Combined with business access control rules, it can be applied to the IPv4 and IPv6 flow tables in the virtualization host kernel for virtual machine traffic forwarding. When a virtual machine migrates, the IPv4 and IPv6 flow tables can be synchronized, enabling on-demand network updates.

[0124] Application layer 3 interacts and connects with platform layer 2. It is used to call the common technical capabilities and data sharing services provided by platform layer 2, build and deploy a business application system that serves the integrated management and control of regional production and operation, and realize a closed loop from global optimization decision-making to edge execution through cloud-edge collaboration mechanism.

[0125] In one embodiment, as shown in FIG5, application layer 3 includes:

[0126] The regional-level hierarchical diagnostic unit 301 is used to extract the operating status data of all production equipment from platform layer 2, and based on the infection source transmission simulation model trained by the AI ​​module, to conduct hierarchical assessment of the fault risk of each production equipment, and generate fault risk assessment results which are transmitted to the intelligent inspection unit; the intelligent inspection unit 302 is used to prioritize inspection tasks based on the fault risk assessment results, plan the optimal inspection path under the conditions of full coverage, resource constraints and cost constraints, and generate the optimal inspection strategy by combining the inspection task priority and the optimal inspection path; the regional-level power prediction unit 303 is used to call the long short-term memory network model trained by the AI ​​module to realize short-term and ultra-short-term power prediction of distributed power sources and loads in the region, and transmit the power prediction results to the regional-level power market management and control unit.

[0127] It should be noted that the Long Short-Term Memory (LSTM) network model, through its unique gated recurrent unit structure, can effectively capture the complex temporal dependencies between distributed power output and load changes within a region, thereby achieving high-precision short-term and ultra-short-term power prediction.

[0128] The model takes historical power time series data gathered from the platform layer 2 data base as its core and integrates multi-dimensional features such as weather forecasts and date types as inputs. The architecture typically consists of an input layer, one or more stacked LSTM layers, and a fully connected output layer. The LSTM layer selectively memorizes long-term historical states and filters irrelevant information through the collaborative operation of forget gates, input gates, and output gates, thereby accurately modeling the nonlinear dynamic changes of the power sequence.

[0129] The model is trained by AI module 203 in platform layer 2. It divides historical data into training and validation sets, uses the actual power value as the label, takes the multidimensional feature sequence of the previous time step as input, and performs supervised training in a time-step rolling manner. The network weights are optimized by minimizing the prediction error through the backpropagation algorithm.

[0130] The regional-level power market control unit 304 is used to make centralized optimization decisions on the declaration, clearing and settlement of adjustable resources in the region participating in market transactions based on power forecast results and predefined power market control rules, and to generate transaction plans and dispatch instructions that are transmitted to the regional collaborative control unit.

[0131] The regional collaborative control unit 305 is used to aggregate transaction plans and scheduling instructions, fault risk assessment results and optimal inspection strategies, formulate collaborative control instructions for edge production units, and issue collaborative control instructions to edge production units for execution through cloud-edge collaborative mechanisms.

[0132] In one embodiment, when the regional-level hierarchical diagnostic unit 301 extracts the operating status data of all production equipment from the platform layer 2, and performs a hierarchical assessment of the fault risk of each production equipment based on the infection source transmission simulation model trained by the AI ​​module, and generates a fault risk assessment result which is then transmitted to the intelligent inspection unit 302, the following steps are included:

[0133] The operating status data, historical fault records, and related operating condition data of each production equipment are extracted from the data base 201 of platform layer 2 and preprocessed to obtain the structured features of the production equipment.

[0134] Based on predefined production equipment topology and physical connections, a knowledge graph is constructed with each production equipment as a node.

[0135] Specifically, the construction of the knowledge graph is a data-driven process that deeply relies on and invokes the core capabilities of Platform Layer 2. Its construction essentially involves: Application Layer 3 initiating a composite request to Platform Layer 2, which then coordinates its data foundation 201 and AI module 203 to jointly complete the real-time construction and updating of the graph. The specific steps are as follows:

[0136] Application layer 3 sends a request to data base 201 of platform layer 2 to obtain real-time aggregated multi-source heterogeneous data such as the operating status of all production equipment, historical fault records, and related operating conditions. Data base 201 provides this raw data, which is usually cleaned, aligned, and feature-extracted by the AI ​​module or data processing service of platform layer 2 to generate structured features that can be used for graph construction.

[0137] Application layer 3, based on predefined business rules such as equipment topology, physical connections, and process logic, calls the graph construction engine through microservices or APIs provided by application development module 202. This engine uses extracted structured features as node attributes and predefined connections as edges to dynamically generate or update a weighted directed graph, i.e., a relational knowledge graph. The edge weights, such as propagation probabilities, can be initialized and dynamically adjusted by the model trained by platform layer 2 based on historical data.

[0138] Using the structured features of production equipment as input, the node core degree hierarchical stripping algorithm is called to calculate the importance of each node. Based on the node importance, risk source nodes, primary susceptible nodes, and several secondary susceptible nodes used for fault risk propagation simulation in the associated knowledge graph are identified.

[0139] Using the risk source node as the starting point of fault propagation, the infection source propagation simulation model is invoked to simulate the fault risk transmission link from the risk source node to the primary susceptible node and each secondary susceptible node in the associated knowledge graph.

[0140] Specifically, by generating fault risk transmission links, it is possible to accurately locate a few but crucial risk source nodes in the network, namely the fault initiator, the primary susceptible node (i.e., the key transmission hub), and the secondary susceptible node (i.e., the main scope of impact), thereby shifting limited inspection and maintenance resources from broad coverage to precise deployment of defenses for the critical few.

[0141] By simulating the complete transmission chain of a fault from its source through critical nodes to its endpoint, it enables predictive maintenance. Maintenance strategies can shift from post-fault repair to preventative maintenance at upstream nodes along the critical transmission path, thus blocking or mitigating risk spread and preventing problems before they occur. When a device experiences a sudden failure, the transmission chain can instantly predict the scope of impact and critical affected equipment, providing clear priorities and impact isolation plans for emergency response, thereby significantly improving the accuracy and efficiency of emergency decision-making.

[0142] Calculate the failure risk value of each path node in the failure risk propagation link under the propagation resistance of the associated node, and use it as the failure risk value of the corresponding production equipment; evaluate the failure risk level of each production equipment based on the failure risk value of each production equipment, and generate a failure risk assessment report including the failure risk level.

[0143] Specifically, this invention transforms traditional isolated equipment status monitoring into networked and systematic risk analysis by dynamically aggregating multi-source data from the data base 201 and constructing a real-time updated device association knowledge graph. Furthermore, based on the knowledge graph, it uses a node core degree hierarchical stripping algorithm to intelligently identify risk sources, transmission hubs, and critical paths. It also utilizes an infection source propagation model to dynamically simulate the chain reaction of faults across the entire network of devices, ultimately accurately quantifying the global risk value of each node in the transmission network. This achieves a breakthrough from single-point early warning to network-wide risk situation awareness and propagation path prediction, providing core decision-making basis for proactive and precise regional-level preventative maintenance and collaborative scheduling.

[0144] In one embodiment, a node core degree hierarchical stripping algorithm is invoked to calculate the importance of each node. Based on the node importance, risk source nodes, primary susceptible nodes, and several secondary susceptible nodes used for fault risk propagation simulation in the associated knowledge graph are identified, including:

[0145] The calculation includes risk propagation characteristic parameters such as the diffusion right of each node in the associated knowledge graph that triggers the ability of downstream nodes, the exposure right triggered by upstream nodes, and the weight balance factor that comprehensively reflects the balance and activity of nodes in the propagation process.

[0146] The initial importance of each node is calculated based on the risk propagation characteristic parameters, and the initial importance is used as the initial importance function for the node iterative decomposition of the node core degree hierarchy stripping algorithm.

[0147] Based on the initial importance function, the nodes with the smallest initial importance value and all their associated edges in the associated knowledge graph are iteratively removed. The nodes to be removed are assigned to different levels according to the removal order, until all nodes in the associated knowledge graph have been removed and the level assignment is completed, forming the initial node level sequence.

[0148] Select nodes at the same level and introduce an iteration factor. Using the relative order in which the current node is removed in the same level, the total number of nodes in the level, and the importance values ​​of adjacent levels as parameters, calculate the intra-level calibration value of the current node. Sum the level number of the current node with the intra-level calibration value to obtain the node importance that comprehensively reflects the global and local importance of the node.

[0149] The nodes are sorted in descending order of importance, and the nodes in the associated knowledge graph are classified and identified by combining the diffusion weight and exposure weight of each node, so as to obtain the risk source node, the primary susceptible node and several secondary susceptible nodes for fault risk propagation simulation.

[0150] Specifically, the node core degree hierarchical stripping algorithm, namely the improved K-shell decomposition algorithm, is the core algorithm for the regional hierarchical diagnostic unit 301 in application layer 3 to implement its business logic, but it is heavily dependent on two key services provided by platform layer 2:

[0151] Data input: The object association knowledge graph it analyzes is aggregated and maintained by data base 201 from multiple sources.

[0152] Parameter support: The edge weights on which the calculation of characteristic parameters such as diffusion weight and exposure weight depends are obtained by AI module 203 based on historical data.

[0153] Here, diffusion is the ability of a quantified node to trigger a downstream node, and its value is the sum of the weights of all outgoing edges of the node; the calculation formula is: diffusion weight = Σ (the weight of the directed edge from the node to each of its downstream nodes). The higher the weight, the greater the potential strength of the node to trigger a fault in the downstream node.

[0154] Exposure weight quantifies the likelihood of a node being triggered by an upstream node, and its value is the sum of the weights of all incoming edges to that node; the calculation formula is: Exposure weight = Σ (the weight of each directed edge pointing to the upstream node). The higher the weight, the more vulnerable the node is to the failure of an upstream node.

[0155] Specifically, upstream and downstream nodes are relative concepts defined based on the risk transmission direction indicated by the directed edge: Specifically, for any directed edge (A→B) in the graph from node A to node B, it means that the fault or risk can be transmitted from A to B. In this relationship, node A is called the upstream node of node B (i.e. the risk source), while node B is called the downstream node of node A (i.e. the risk bearer).

[0156] The weight balance factor comprehensively reflects the balance between input and output and the activity level of a node during the propagation process. Its core calculation is the relative ratio of diffusion weight to exposure weight. The calculation formula is: Weight balance factor = 1 - (diffusion weight / (diffusion weight + exposure weight) - 0.5) 2 The value reaches its maximum value of 1 when the diffusion right and exposure right are equal, indicating that the node is in a perfect active balance state. When the difference between the two is extremely large, the value approaches 0, indicating that the node is highly biased towards a pure diffusion source or absorption sink. The product of this factor and (diffusion right + exposure right) can further characterize the overall propagation activity of the node.

[0157] The initial importance of a node is a comprehensive function of the three parameters mentioned above. This step ensures that an initially important node must simultaneously possess strong diffusion capabilities, high exposure risk, and the potential to play an active balancing role in the propagation network. This calculation process is entirely completed by the application layer business logic calling the basic graph data and weight parameters provided by platform layer 2.

[0158] As shown in Table 1, this embodiment takes the fault association knowledge graph of core equipment in a thermal power generating unit as the research object, selecting six key nodes: boiler system A, turbine system B, generator system C, feedwater pump system D, condenser system E, and high-pressure heater system F. Based on the fault propagation logic between equipment (e.g., boiler tube rupture can trigger excessive turbine vibration, and feedwater pump failure can lead to insufficient boiler water supply), a diffusion weight is set, specifically in the range of 0-1, representing the node's ability to trigger downstream faults, and an exposure weight, specifically in the range of 0-1, representing the degree to which the node is triggered by upstream faults. The weight balance factor, which comprehensively reflects the propagation balance and activity, is calculated using the formula "weight balance factor = 2 × diffusion weight × exposure weight / (diffusion weight + exposure weight)". The initial importance value is taken as weight balance factor × 10. The node core degree hierarchical stripping algorithm is based on the initial importance from... The nodes and associated edges are removed iteratively from smallest to largest importance. The removal order is as follows: F (initial importance 4.9), E (5.7), D (6.4), C (7.3), B (7.9), and A (8.6). The resulting hierarchy is: F belongs to level 1, E to level 2, D and C to level 3, B to level 4, and A to level 5. The iteration factor is set to 0.1. The intra-level calibration value is calculated using the formula: Intra-level calibration value = (intra-level removal order - 1) / (total number of nodes in the layer - 1) × iteration factor. The single-node hierarchical calibration value is 0. The node importance is the sum of the hierarchical number and the intra-level calibration value. Finally, nodes are sorted in descending order of importance and classified according to diffusion weight and exposure weight: nodes with diffusion weight ≥ 0.75 and node importance ≥ 5.0 are identified as risk source nodes; nodes with exposure weight ≥ 0.7 and node importance ≥ 4.0 are identified as primary susceptible nodes; and the rest are secondary susceptible nodes.

[0159] Table 1 Fault Association Table

[0160] In one embodiment, nodes are sorted in descending order of importance, and the nodes in the associated knowledge graph are classified and identified by combining the diffusion weight and exposure weight of each node, resulting in risk source nodes, primary susceptible nodes, and several secondary susceptible nodes for fault risk propagation simulation, including:

[0161] The nodes are sorted in descending order of absolute importance. The node with the highest importance in the sorted list and whose diffusion weight is greater than its exposure weight is used as the risk source node for fault risk propagation simulation.

[0162] The node with the highest node importance in the sorted list and whose exposure weight is greater than its diffusion weight is selected as the primary susceptible node for fault risk propagation simulation.

[0163] The remaining nodes in the sorted list are used as secondary susceptible nodes for fault risk propagation simulation.

[0164] Specifically, this invention combines the absolute importance ranking of nodes with their inherent diffusion and exposure attributes, thereby achieving automated and precise classification of key roles in complex networks. This overcomes the inherent defect that a single importance indicator cannot distinguish the specific role of a node in risk propagation (whether it is a source, hub, or path), enabling the simulation process to be efficiently started from the predicted risk source node, and thus providing the most critical structured input for generating inspection and maintenance strategies with clear prevention and control targets.

[0165] In one embodiment, taking the risk source node as the starting point of fault propagation, the infection source propagation simulation model is invoked to simulate the fault risk transmission chain from the risk source node to the primary susceptible node and each secondary susceptible node in the associated knowledge graph, including:

[0166] The initial state of the risk source node in the associated knowledge graph is marked as the infected state and used as the initial infection source for the fault risk propagation simulation. The initial state of the remaining nodes is marked as the susceptible state.

[0167] The fault risk propagation evolution process is executed iteratively based on discrete time steps, and in each iteration, the state of each node is traversed step by step starting from the initial source of infection.

[0168] If the current node is in a susceptible state, the risk probability of the current node being infected is calculated by combining the directed weights of the current node, all upstream nodes in the infected state, and the connecting edges. If the current node is in an infected state, the recovery probability of converting the infected state to a recovered state is determined according to the preset recovery probability. If the infected state is converted to a recovered state, the current node exits the propagation chain.

[0169] Specifically, the core objective of this step is to simulate the complete dynamic evolution of fault risks in a device network, from propagation and diffusion to eventual repair, using computable probability rules on a dynamically weighted, relational knowledge graph. By defining the susceptible-infected-recovery state for each node and its probability transfer rules based on network connections and weights, the model can quantify the uncertainty and temporality of risk propagation along a preset transmission path. This allows for the statistical analysis of the transmission frequency, critical path, and impact range of risks from source to end through numerous simulations, providing a data-driven basis for assessing system vulnerability, identifying key bottlenecks, and developing preventative intervention strategies.

[0170] The status of all nodes is updated synchronously. Susceptible nodes that reach the risk probability are updated to the infected state; infected nodes that reach the preset recovery probability are updated to the recovery state, so as to drive the failure risk to spread from the infected nodes to the downstream nodes in each step, realizing the chain propagation of failure risk from the risk source node to the primary susceptible node, and then from the newly infected primary susceptible node to the secondary susceptible node.

[0171] In each iteration, the state transition event between the corresponding source node and the infected node is recorded, and the process stops when the number of iterations is satisfied or there are no infected state nodes in the associated knowledge graph.

[0172] The state transition events recorded in each iteration are summarized to obtain the propagation event log, and the fault risk transmission chain is extracted based on the propagation event log.

[0173] Specifically, the construction and training of the infection source transmission simulation model in the AI ​​module 203 of platform layer 2 is a closed-loop process that deeply integrates domain knowledge, graph neural networks and time-series data learning. Its core architecture is a graph neural network simulator that takes a dynamically weighted knowledge graph as input, differentiable propagation dynamics as the core, and historical fault data as the supervision signal. Its operating principle is to automatically learn and optimize the quantization model of edge weights in the knowledge graph and the state transition function of nodes through a data-driven approach, thereby simulating the dynamic propagation process of fault risk with high fidelity.

[0174] Specifically, the construction of the infection source transmission simulation model begins with multi-source historical data provided by the data foundation, including the time-series operating status of equipment, alarm logs, maintenance records, and equipment topology relationships. Based on this, a historical knowledge graph with initial weights is constructed, where nodes represent equipment, edges represent potential fault transmission relationships, and the initial weights of the edges can be predefined based on physical principles or statistical correlations. The core architecture of the model adopts a joint structure of a graph neural network encoder and a differentiable propagation dynamics layer: the graph neural network encoder is responsible for learning high-order representations from complex node and edge features and outputting a dynamic, context-dependent edge weight adjustment parameter; at the same time, the differentiable propagation dynamics layer uses this real-time probability to simulate the continuous or discrete temporal evolution process of infection and recovery states on the graph.

[0175] The model is trained by maximizing the likelihood probability between its predicted trajectory and historical real-world fault propagation records. This is achieved by continuously adjusting the parameters in the graph neural network encoder and propagation dynamics layer through the backpropagation algorithm, so that the fault initiation point, propagation path, and propagation speed simulated by the model are as consistent as possible with historical data. After training, the model is encapsulated as a service that can be called via API. When the application layer 3 inputs a real-time or hypothetical knowledge graph and an initial risk source, the model can quickly deduce the risk propagation range, critical path, and infection probability of each node in the future.

[0176] In one embodiment, the state transition events recorded in each iteration are aggregated to obtain a propagation event log. The fault risk propagation chain is extracted based on the propagation event log, including:

[0177] The state transition events recorded in each iteration are summarized to obtain the propagation event log. All initial propagation events that are directly caused by the initial risk source node as the source of infection are filtered out from the propagation event log.

[0178] The initial propagation events are integrated with the associated nodes according to the time step of occurrence to form a primary transmission link of failure risk from the risk source node to the main susceptible node;

[0179] Using the primary susceptible nodes that have already been infected in the primary transmission chain as new sources of infection, the secondary transmission events triggered by the new sources of infection are further filtered out in the transmission event log;

[0180] Secondary propagation events are integrated according to their time steps and associated nodes to form a secondary transmission link of fault risk from the primary susceptible node to the secondary susceptible node;

[0181] By integrating the primary and secondary transmission links of fault risk, a comprehensive fault risk transmission link is obtained.

[0182] Specifically, the infection probability of susceptible nodes is calculated by weighting upstream infected nodes and connecting edges, and the state transition of infected nodes is determined by combining the recovery probability. This quantifies the uncertainty of fault propagation, making the simulation results more consistent with the random triggering characteristics of actual equipment failures. Synchronous node state updates drive the chain propagation of each step through state transitions. The implementation principle is to recreate the chain reaction logic of equipment failures, fully replicating the actual propagation sequence from the risk source to the primary susceptible node and then to the secondary susceptible node. State transition events are recorded and stopped when there are no infected nodes, achieving full-link tracing of the propagation process and providing complete data support for subsequent extraction of the transmission chain. Initial / secondary propagation events are filtered from the event log and integrated into the transmission chain. The path is split and integrated according to the propagation source level, thus presenting the scope of the chain impact of the failure, clarifying the role of each related device in the propagation, and facilitating the development of targeted emergency response plans after individual device failures.

[0183] As shown in Table 2, this embodiment uses the knowledge graph associated with thermal power generating units as the simulation object. Node A corresponds to the boiler system and is set as the risk source node. Node B corresponds to the steam turbine system as the primary susceptible node. Nodes C, D, and E correspond to the generator system, feedwater pump system, and condenser system as secondary susceptible nodes, respectively. Initially, node A is marked as infected, and nodes B, C, D, and E are marked as susceptible. The simulation then iteratively executes discrete-time step simulations. At time step t=0, only node A is infected, and no state transition event occurs. At time step t=2, the calculated probability that node B is infected by the source node A reaches the threshold, and node B's state changes from susceptible to infected. The system records this state transition event, marking node A as the source of infection and node B as the target node. At time step t=3, the calculated probability that node D is infected by the current source node B reaches the threshold, and node D's state changes from susceptible to infected. The system records this transition event, marking node B as the source of infection and node D as the target node. At time step t=5, the calculated probability that node C is infected... When the probability of infection of source node B reaches the target, its state changes from susceptible to infected. At the same time, the initial source node A reaches the preset recovery probability, and its state changes from infected to recovered. The system records the events of node B infecting node C and node A recovering. At time step t=7, it is calculated that the probability of node E being infected by source node D reaches the target, and its state changes from susceptible to infected. At the same time, node B reaches the recovery probability, and its state changes from infected to recovered. The system records the events of node D infecting node E and node B recovering. Subsequently, at time step t=8, node C reaches the recovery probability and becomes recovered; at time step t=10, node D reaches the recovery probability and becomes recovered; finally, at time step t=12, node E reaches the recovery probability and becomes recovered, and the entire simulation process stops.

[0184] After the simulation, the system aggregates the event logs recorded at all time steps. First, it filters out propagation events directly originating from the initial risk source node A, integrating them to form the primary propagation link A→B from A to B. Next, using the already infected primary susceptible node B in the primary link as a new source, it filters out the secondary propagation events triggered by B, integrating them to form secondary links B→D and B→C. Then, using the already infected secondary susceptible node D as the source, it filters out the subsequent propagation events triggered by D, integrating them to form the secondary link D→E. By integrating the primary propagation link with each level of secondary propagation link, the comprehensive propagation links A→B→D→E and A→B→C, describing the path of fault risk propagation throughout the network, are obtained.

[0185] Table 2 Simulation Event Table for Fault Risk Propagation (Taking Risk Source A as an Example)

[0186] In one embodiment, the intelligent inspection unit 302, when prioritizing inspection tasks based on fault risk assessment results, planning the optimal inspection path under the conditions of full coverage, resource constraints, and cost constraints, and generating the optimal inspection strategy by combining the inspection task priorities and the optimal inspection path, includes:

[0187] Based on the failure risk assessment results, the initial inspection task priority is generated; the influence relationship between each production equipment and related equipment in the failure risk transmission link is analyzed, and the betweenness centrality algorithm is called to identify the attenuating equipment that blocks the transmission of failure risk and the amplifying equipment that expands the risk range;

[0188] When a sudden change in the status of production equipment is detected, the scope of the change is analyzed based on the fault risk transmission link corresponding to the production equipment. The topological importance of attenuating and amplifying equipment is then integrated into the scope of the change's impact, and the priority of the inspection task is updated.

[0189] This invention upgrades traditional static inspections based on single-point status or fixed cycles to a dynamic and precise operation and maintenance model based on system vulnerability analysis and real-time risk prediction by introducing network topology analysis and dynamic risk simulation. It uses a betweenness centrality algorithm to identify attenuation devices that critically block risk propagation and amplification devices that significantly amplify the impact from the global transmission network structure, focusing inspection targets from all devices to key global nodes. When any device experiences a sudden state change, the system can immediately simulate its potential impact range based on pre-built transmission links and dynamically adjust inspection priorities by integrating the topological importance of key nodes. This ensures that, under limited resource conditions, inspection resources are always prioritized for key locations most likely to trigger systemic risks or most effectively suppress risk spread. Ultimately, this achieves a fundamental shift from passively responding to existing faults to proactively preventing systemic risks, significantly improving the efficiency of operation and maintenance resource utilization and the overall system's risk resilience.

[0190] The updated inspection task priority is passed to the preset path planning and optimization component through the microservice provided by the application development module. The preset path planning and optimization component transforms the updated inspection task priority into a benefit item in path planning. It uses the geographical location of production equipment, the status of available inspection resources, the cost of various resources, and full coverage as multi-level constraints, and takes the lowest total execution cost or the highest total efficiency as the optimization objective to build a path optimization model.

[0191] It should be noted that the architecture of the path optimization model is essentially a multi-constraint combinatorial optimization system. Its core components include: input parameters, namely, task benefits derived from inspection priorities, equipment geographical location, real-time available resource status, various cost coefficients, and constraints such as full coverage that must be met; decision variables, usually a series of binary variables, used to determine "which resources and in what order which tasks are executed"; objective function, formalized as maximizing total benefits or minimizing total costs, thereby ensuring that high-priority decaying and amplifying equipment is accessed first; and a set of constraint equations, used to strictly express business rules such as resource capacity, task time windows, and path continuity. Its operating principle lies in encoding dynamically updated inspection business needs, such as sudden high-priority tasks and real-time physical world constraints, such as resource location and status, into a computable mathematical programming problem. That is, by modeling the decision-making process of accessing key equipment at the right time, with the right resources, and in the optimal order, it is transformed into a search in a vast solution space for resource-task matching and path sequences that satisfy all constraints and achieve the optimal objective.

[0192] The parameters of the path optimization model are not obtained through traditional machine learning training, but are continuously optimized through inverse optimization analysis and parameter calibration of historical operation and maintenance data and expert strategies by AI module 203. This process can be regarded as model training, aiming to align the value judgments of the mathematical model with actual business preferences.

[0193] Solve the path optimization model to obtain the optimal inspection strategy that prioritizes the attenuation and amplification devices on the fault risk transmission link; send the optimal inspection strategy to the edge inspection terminal for execution, and transmit the real-time data generated during the execution back to the data base of platform layer 2.

[0194] It should be noted that solving the path optimization model relies on specialized mathematical programming solvers, which are mainly divided into two categories: one is solvers based on precise algorithms such as branch and bound and cutting plane (e.g., CPLEX, Gurobi solvers), which systematically enumerate and eliminate infeasible solution spaces to find mathematically optimal solutions for small- to medium-sized problems within an acceptable time; the other is solvers based on metaheuristic algorithms such as genetic algorithms, simulated annealing, and large-scale neighborhood search, which work by simulating heuristic rules such as natural evolution or physical processes to efficiently search for feasible solutions that approximate the optimal solution in ultra-large-scale or real-time scenarios.

[0195] It should be noted that the betweenness centrality algorithm is a node importance measurement algorithm based on the global network topology. Its core principle is to quantify the control capability of a node as a network hub or bridge by calculating the shortest (or optimal) path between all pairs of nodes in the network and counting the frequency of each node appearing on these shortest paths. Specifically, the algorithm first calculates the optimal transmission path between all pairs of devices (usually based on the principle of minimum transmission resistance), and then counts the number of times each device appears on the optimal path between any other pair of devices. This occurrence count, after normalization, is the betweenness centrality value of the device. The higher the value, the more critical transmission paths the device is on, and the stronger its control or influence on the overall network connectivity and risk transmission process.

[0196] In one embodiment, the analysis of the influence relationship between each production device and related devices in the fault risk propagation chain, and the invocation of the betweenness centrality algorithm to identify attenuating devices that block the propagation of the fault risk chain and amplifying devices that expand the scope of risk, include:

[0197] Extract the risk transmission path with the least resistance between each pair of production equipment in the fault risk transmission chain, and form a set of risk transmission paths;

[0198] The total number of times each production device appears in the risk transmission path between any other production device pair in the statistical risk transmission path set, and the total number of different target devices that can be reached by all risk transmission paths starting from each production device, are counted and stored in the global path hub frequency table.

[0199] Based on the global path hub frequency table, structural attenuation and structural amplification indices are calculated for each production device. The structural attenuation and structural amplification indices are compared with preset thresholds, and a list of attenuation devices and a list of amplification devices are generated based on the comparison results.

[0200] It should be noted that the structural attenuation potential index calculated by this invention can be directly mapped to the ability of a device to act as a potential risk blockage point (because blocking high betweenness centrality devices can maximally disrupt network connectivity), while the structural amplification potential index can be combined with the reach range of a node to reflect its potential as a source of risk diffusion. This provides a solid data-driven basis for accurately locating attenuation and amplification devices at the network topology level, enabling subsequent inspection priority setting and resource allocation to focus on these key structural points, thereby achieving the goal of obtaining the maximum overall risk prevention and control effectiveness with minimal intervention cost.

[0201] Furthermore, the cloud-edge collaborative digital platform based on integrated regional production and operation management provided by this invention also includes a cloud security system: a robust security system is built upon the cloud platform to ensure its overall security. The construction of the cloud platform security system mainly includes the development of five basic technical systems: physical security, network security, host security, application security, and data security.

[0202] The cloud platform security construction mainly includes three parts: a cloud security operation platform, which centrally manages the cloud security situation and coordinates various security atomic modules for joint prevention and control; a compliance module, which is a supporting module closely built around cloud computing, cloud storage, and cloud networks, mainly including cloud host security, cloud bastion host, log auditing, cloud firewall, cloud WAF, database auditing, etc.; and a national cryptography module, which has transformed the cloud platform with national cryptography capabilities, combined with products such as KMS, national cryptography suites, and cryptographic machines, to ensure system security and cryptographic compliance requirements and improve security protection capabilities.

[0203] Cloud security operations; the cloud management platform provides unified cloud security operations management capabilities, with security capabilities presented in a service catalog, and has the ability to automatically complete deployment, activation, and use; through the operations management portal, it provides platform administrators with comprehensive management functions, such as user management, resource management, and capability management, to achieve efficient and simple operations.

[0204] The cloud security compliance module is divided into two parts: platform-side security and user-side security. Platform-side security mainly focuses on the business characteristics of the cloud platform, combining platform boundary security protection, host machine and virtualization security, and platform security to achieve continuous protection of the cloud platform while meeting compliance requirements. This is achieved through the deployment of hardware and software facilities such as bastion hosts, log auditing, database auditing, next-generation cloud firewalls, and host security to ensure the overall security of the cloud platform.

[0205] The National Cryptography Module comprises three main parts: cloud cryptography service operation and management, cryptography resource pool, and cryptography application layer. Specifically, cloud cryptography service operation and management provides unified cryptography services to all business application systems on the cloud platform through a unified cryptography service bus. It offers multi-user cryptography service capabilities, manages various cryptography service interfaces, service subscriptions, application calls, application authentication, platform operation, and provides multi-user management and cryptography resource management.

[0206] Cryptographic Resource Pool: Provides unified basic cryptographic computing services, consisting of various cryptographic devices and basic cryptographic service units, including: cloud server cryptographic machines, signature verification servers, CA certificate authentication systems, key management systems, client and collaborative signature systems, etc., and can provide various cryptographic atomic service capabilities.

[0207] The cryptographic application layer comprises four parts: terminal security cryptographic applications, network access security cryptographic applications, cloud platform and business cryptographic applications, and cloud security management cryptographic applications. It primarily considers the needs of cloud users and the cloud platform, providing end-to-end cryptographic protection, including cryptographic application services from the perspectives of terminals, communication networks, cloud platforms and business systems, and operations and maintenance. Upper-layer applications based on the cryptographic infrastructure service platform typically consist of client and server components. They are adapted and interfaced with the cryptographic infrastructure service platform, utilizing a cloud cryptographic service operation and management platform to provide a cryptographic service bus. This provides unified cryptographic service API interfaces, protocols, and SDKs to offer cryptographic services to cloud-based business applications, enabling the application of domestically produced commercial cryptography in areas such as identity authentication, transmission encryption, and storage encryption.

[0208] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram is shown in Figure 3. The computer device includes a processor, memory, and a network interface connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The database stores static and dynamic information data. The network interface communicates with external terminals via a network connection. When the computer program is executed by the processor, it implements the steps in the above method embodiment.

[0209] Those skilled in the art will understand that the structure shown in Figure 3 is merely a block diagram of a portion of the structure related to the present invention and does not constitute a limitation on the computer device to which the present invention is applied. A specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0210] In addition, the present invention also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above method embodiments.

[0211] In addition, the present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps in the above method embodiments.

[0212] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the methods described above. Any references to memory, storage, databases, or other media used in the embodiments provided by this invention can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical storage, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.

[0213] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A cloud-edge collaborative digital platform based on integrated regional production and operation management, characterized in that, The platform comprises: an infrastructure layer, used to build a unified hardware resource pool and automated operation and maintenance base to support the cloud-edge collaborative architecture, enabling unified management and automated operation and maintenance of underlying servers, storage, and network hardware devices; a platform layer, connected to the infrastructure layer, serving as the capability hub of the cloud-edge collaborative architecture, used to aggregate full-domain data and integrate application development modules and AI modules, forming a common toolchain and data sharing services covering low-code development, microservices, and the entire lifecycle management of models; and an application layer, interacting with the platform layer, used to call upon the common technical capabilities and data sharing services provided by the platform layer, building and deploying services for integrated regional production and operation management. The system comprises a business application framework and a cloud-edge collaboration mechanism to achieve a closed loop from global optimization decision-making to edge execution. The application layer includes: a regional-level hierarchical diagnostic unit, which extracts operational status data of all production equipment from the platform layer and, based on an infection source transmission simulation model trained by an AI module, performs hierarchical assessment of the fault risk of each production equipment, generating fault risk assessment results that are transmitted to the intelligent inspection unit; and an intelligent inspection unit, which prioritizes inspection tasks based on the fault risk assessment results, plans the optimal inspection path under the conditions of full coverage, resource constraints, and cost constraints, and generates the optimal inspection strategy by combining the inspection task priorities and the optimal inspection path.

2. The cloud-edge collaborative digital platform based on integrated regional production and operation management as described in claim 1, characterized in that, The application layer also includes: a regional power prediction unit, which calls the long short-term memory network model trained by the AI ​​module to realize short-term and ultra-short-term power prediction of distributed power sources and loads within the region, and transmits the power prediction results to the regional power market management and control unit; a regional power market management and control unit, which, based on the power prediction results and predefined power market management and control rules, makes centralized optimization decisions on the application, clearing and settlement of adjustable resources participating in market transactions within the region, and forms transaction plans and dispatch instructions that are transmitted to the regional collaborative control unit; and a regional collaborative control unit, which aggregates transaction plans and dispatch instructions, fault risk assessment results and optimal inspection strategies, formulates collaborative control instructions for edge production units, and issues collaborative control instructions to edge production units for execution through a cloud-edge collaborative mechanism.

3. The cloud-edge collaborative digital platform based on integrated regional production and operation management as described in claim 1, characterized in that, The regional-level hierarchical diagnostic unit extracts operational status data of all production equipment from the platform layer and, based on the infection source transmission simulation model trained by the AI ​​module, performs a hierarchical assessment of the fault risk of each production equipment. When transmitting the fault risk assessment results to the intelligent inspection unit, the process includes: extracting operational status data, historical fault records, and related operating condition data of each production equipment from the platform layer's data base and preprocessing them to obtain the structured features of the production equipment; constructing a knowledge graph with each production equipment as a node based on predefined production equipment topology and physical connections; and using the structured features of the production equipment as input to call the node core degree hierarchical stripping algorithm. The method calculates the importance of each node, and based on the node importance, identifies the risk source node, primary susceptible node, and several secondary susceptible nodes in the associated knowledge graph used for fault risk propagation simulation. Taking the risk source node as the starting point of fault propagation, the infection source propagation simulation model is invoked to simulate the fault risk transmission link from the risk source node to the primary susceptible node and each secondary susceptible node in the associated knowledge graph. The fault risk value of each path node in the fault risk transmission link under the transmission resistance of the associated nodes is calculated as the fault risk value of the corresponding production equipment. Based on the fault risk value of each production equipment, the fault risk level of each production equipment is evaluated, and a fault risk assessment report including the fault risk level is generated.

4. The cloud-edge collaborative digital platform based on integrated regional production and operation management as described in claim 3, characterized in that, The node core degree hierarchical stripping algorithm is used to calculate the importance of each node. Based on this importance, the algorithm identifies risk source nodes, primary susceptible nodes, and several secondary susceptible nodes in the associated knowledge graph used for fault risk propagation simulation. This includes: calculating risk propagation characteristic parameters, which include the diffusion weight of each node's ability to trigger downstream nodes, the exposure weight triggered by upstream nodes, and a weight balance factor comprehensively reflecting the balance and activity of nodes during propagation; calculating the initial importance of each node based on these risk propagation characteristic parameters, and using this initial importance as the initial importance function for the node core degree hierarchical stripping algorithm to perform iterative node decomposition; and iteratively removing nodes with the minimum initial importance value and all their associated edges from the associated knowledge graph based on the initial importance function. The nodes to be removed are assigned to different levels according to the removal order, until all nodes in the associated knowledge graph have been removed and their level assignments are completed, forming an initial node level sequence. Nodes in the same level are selected and an iteration factor is introduced. The relative order in which the current node is removed in the same level, the total number of nodes in the level, and the importance values ​​of adjacent levels are used as parameters to calculate the intra-level calibration value of the current node. The level number of the current node is summed with the intra-level calibration value to obtain the node importance that comprehensively reflects the global and local importance of the node. The node importance is sorted in descending order, and the nodes in the associated knowledge graph are classified and identified in combination with the diffusion weight and exposure weight of each node to obtain the risk source nodes, primary susceptible nodes, and several secondary susceptible nodes for fault risk propagation simulation.

5. A cloud-edge collaborative digital platform based on integrated regional production and operation management as described in claim 4, characterized in that, The process of sorting nodes by importance in descending order and classifying and identifying nodes in the associated knowledge graph based on their diffusion and exposure weights results in risk source nodes, primary susceptible nodes, and several secondary susceptible nodes for fault risk propagation simulation. This includes: sorting nodes by absolute importance in descending order; identifying the node with the highest importance and diffusion weight greater than exposure weight in the sorted list as the risk source node for fault risk propagation simulation; identifying the node with the highest importance and exposure weight greater than diffusion weight in the sorted list as the primary susceptible node for fault risk propagation simulation; and identifying the remaining nodes in the sorted list as secondary susceptible nodes for fault risk propagation simulation.

6. The cloud-edge collaborative digital platform based on integrated regional production and operation management as described in claim 4, characterized in that, The method of using the risk source node as the starting point for fault propagation and calling the infection source propagation simulation model to simulate the fault risk transmission link from the risk source node to the primary susceptible node and each secondary susceptible node in the associated knowledge graph includes: marking the initial state of the risk source node in the associated knowledge graph as an infected state and using it as the initial infection source for fault risk propagation simulation; marking the initial state of the remaining nodes as susceptible states; iteratively executing the fault risk propagation evolution process based on discrete time steps, and in each iteration, gradually traversing the state of each node starting from the initial infection source; if the current node is in a susceptible state, calculating the risk probability of the current node being infected by combining the directed weights of the current node and the upstream nodes and connecting edges of all infected states; if the current node is in an infected state, determining the infection state according to a preset recovery probability. If the infected state is converted to the recovered state, the current node exits the propagation chain; the state of all nodes is updated synchronously, and susceptible nodes that have reached the risk probability are updated to the infected state; infected nodes that have reached the preset recovery probability are updated to the recovered state, so as to drive the failure risk to propagate from the infected node to the downstream node in each step, realizing the chain propagation of failure risk from the risk source node to the primary susceptible node, and then from the newly infected primary susceptible node to the secondary susceptible node; in each iteration, the state transition event formed by the corresponding infection source node and the infected node is recorded, and the propagation stops when the number of iterations is satisfied or there are no infected state nodes in the associated knowledge graph; the state transition events recorded in each iteration are summarized to obtain the propagation event log, and the failure risk transmission chain is extracted based on the propagation event log.

7. A cloud-edge collaborative digital platform based on integrated regional production and operation management as described in claim 6, characterized in that, The process of summarizing the state transition events recorded in each iteration to obtain a propagation event log, and extracting the fault risk transmission link based on the propagation event log, includes: summarizing the state transition events recorded in each iteration to obtain a propagation event log; filtering out all initial propagation events directly caused by the initial risk source node as the infection source from the propagation event log; integrating the initial propagation events with associated nodes according to their occurrence time steps to form a primary fault risk transmission link from the risk source node to the primary susceptible node; using the primary susceptible node already infected in the primary transmission link as a new infection source, continuing to filter out secondary propagation events triggered by the new infection source from the propagation event log; integrating the secondary propagation events with associated nodes according to their occurrence time steps to form a secondary fault risk transmission link from the primary susceptible node to the secondary susceptible node; and integrating the primary fault risk transmission link and the secondary fault risk transmission link to obtain a comprehensive fault risk transmission link.

8. The cloud-edge collaborative digital platform based on integrated regional production and operation management as described in claim 1, characterized in that, The intelligent inspection unit, when prioritizing inspection tasks based on fault risk assessment results and planning the optimal inspection path under the conditions of full coverage, resource constraints, and cost constraints, generates the optimal inspection strategy by combining the inspection task priorities and the optimal inspection path. This includes: generating initial inspection task priorities based on fault risk assessment results; analyzing the influence relationship between each production device and related devices in the fault risk propagation chain, and using the betweenness centrality algorithm to identify attenuating devices that block the transmission of fault risk and amplifying devices that expand the risk range; when a sudden change in the state of a production device is detected, analyzing the impact range of the sudden change based on the fault risk propagation chain corresponding to the production device, and integrating the impact range of the sudden change with the topology of attenuating and amplifying devices. The importance of the inspection task is assessed, and the updated priority is then passed to the preset path planning and optimization component via microservices provided by the application development module. This component transforms the updated priority into a benefit item in the path planning, using the geographical location of production equipment, the status of available inspection resources, the costs of various resources, and full coverage as multi-level constraints. The optimization objective is to minimize total execution cost or maximize total efficiency, thus constructing a path optimization model. Solving the path optimization model yields the optimal inspection strategy that prioritizes the execution of attenuating and amplifying devices along the fault risk propagation path. This optimal inspection strategy is then distributed to the edge-side inspection terminal for execution, and the real-time data generated during execution is transmitted back to the platform layer's data base.

9. A cloud-edge collaborative digital platform based on integrated regional production and operation management as described in claim 8, characterized in that, The analysis of the influence relationship between each production device and related devices in the fault risk transmission chain, and the use of the betweenness centrality algorithm to identify attenuating devices that block the transmission of fault risk and amplifying devices that expand the risk range, includes: extracting the risk transmission path with the least resistance between each pair of production devices in the fault risk transmission chain, forming a set of risk transmission paths; counting the total number of times each production device appears in the risk transmission path between any other pair of production devices in the risk transmission path set, and the total number of different target devices that can be reached by all risk transmission paths starting from each production device, and storing the statistical results in a global path hub frequency table; based on the global path hub frequency table, calculating structural attenuation and structural amplification indices for each production device, comparing the structural attenuation and structural amplification indices with preset thresholds respectively, and generating a list of attenuating devices and a list of amplifying devices based on the comparison results.

10. A cloud-edge collaborative digital platform based on integrated regional production and operation management as described in claim 1, characterized in that, The platform layer includes: a data foundation for aggregating, storing, and managing multi-source heterogeneous data from the infrastructure and application layers via IoT, while providing standardized data sharing services and access interfaces; an application development module for providing low-code development tools and microservice frameworks for the application layer, supporting visual orchestration, service governance, and deployment; and an AI module for providing full lifecycle management tools covering model development, model training, and model management, providing common technical support for the application layer.

11. A cloud-edge collaborative digital platform based on integrated regional production and operation management as described in claim 1, characterized in that, The platform also includes: a cross-cloud-edge consistency check unit, used to establish a consistency checkpoint triggering mechanism for write resources at the platform layer in the cloud-edge collaborative architecture, generating a checkpoint identifier and issuing a checkpoint triggering command when a backup is triggered; a write freeze and sequential processing unit, used to perform boundary division of write resources before and after the checkpoint identifier within the freeze window in response to the checkpoint triggering command, and generate a freeze confirmation set; a cross-resource consistency snapshot unit, used to perform snapshot operations on write resources associated with the freeze confirmation set, and bind the joint snapshot set of the snapshot operation to a unified checkpoint identifier; and a policy-based processing unit, used to load the joint snapshot set when a cloud-edge disconnection, node failure, or manual triggering recovery operation occurs, and to replay or discard write resources after the checkpoint identifier, using the checkpoint identifier as the boundary.