Cloud ASON controller deployment method, system and device for electric power Internet of Things, and medium
By deploying a distributed ASON controller cluster in the power Internet of Things, alarm information is automatically parsed and intelligent routing pre-calculation is performed, which solves the problems of centralized controllers becoming network bottlenecks and fault handling relying on manual intervention, and realizes on-demand resource supply and efficient fault recovery.
Patent Information
- Application Number
- CN202511709058.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-20
- Publication Date
- 2026-02-10
AI Technical Summary
Centralized controllers are prone to becoming network bottlenecks, and a failure could paralyze the control functions of the entire network. Traditional deployment methods use static allocation of time slot resources, which cannot dynamically adapt to sudden surges in video conferencing traffic, resulting in low resource utilization. Fault handling processes heavily rely on manual intervention, which is time-consuming.
Deploy a distributed ASON controller cluster consisting of a resource scheduling engine, a routing decision engine, and a fault recovery engine to monitor network status, automatically parse alarm information, build a virtual topology model, perform intelligent route pre-calculation, and, when a fault occurs, collaboratively execute an automated recovery process through the fault recovery engine and a lightweight agent.
It solves the problems of centralized controllers becoming network bottlenecks and single-point failures causing the entire network control function to be paralyzed. It realizes on-demand supply and intelligent optimization of resources, reduces the time required for manual intervention in fault handling, and improves resource utilization and fault recovery efficiency.
Smart Images

Figure CN121509856A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of power communication network operation and maintenance technology, specifically to a cloud-based ASON controller deployment method, system, device, and medium for the power Internet of Things. Background Technology
[0002] With the rapid development of the power Internet of Things (IoT), power communication networks carry many critical services such as video conferencing. Automatically Switched Optical Networks (ASON), as the core architecture of power communication networks, are responsible for core functions such as routing calculation, connection management, and fault recovery. The deployment and performance of its controllers directly affect the stability of the entire network and the quality of service transmission. However, facing the increasingly large and complex network environment of the power IoT, traditional controller deployment models have shown many limitations.
[0003] Traditional ASON controller deployments often employ a centralized architecture, which presents several shortcomings when facing the increasingly large and complex network environment of the power Internet of Things (IoT). Firstly, the centralized controller, as the sole control core of the entire network, is prone to becoming a performance bottleneck and a single point of failure. If this node experiences hardware or software failures or network congestion, the control functions of the entire communication network will be paralyzed, affecting the continuity and stability of critical services such as video conferencing. Secondly, traditional solutions generally use static allocation to pre-allocate time slots for circuits, failing to adjust time slots according to sudden traffic surges in services like video conferencing. This results in wasted bandwidth resources or congestion, low resource utilization, and difficulty in meeting the demands of highly elastic services. Furthermore, fault handling heavily relies on manual intervention for analyzing alarms and locating fault points, and on performing route switching step-by-step through the network management system. This process is time-consuming, with a long average recovery time for a single circuit, making it unsuitable for the stringent real-time and reliability requirements of video conferencing. Therefore, there is an urgent need for a cloud-based ASON controller deployment method for the power Internet of Things, in order to overcome the bottlenecks of existing technologies in terms of reliability, resource efficiency and operation and maintenance automation, and comprehensively improve the intelligence level and service assurance capabilities of the power communication network. Summary of the Invention
[0004] In view of the above-mentioned problems, the present invention is proposed.
[0005] Therefore, the technical problem solved by this invention is that centralized controllers are prone to becoming network bottlenecks, and once a failure occurs, the control function of the entire network may be paralyzed; traditional deployment methods use static allocation of time slot resources, which cannot dynamically adapt to sudden traffic surges in video conferencing, resulting in low resource utilization; and the fault handling process still heavily relies on manual intervention, which is time-consuming.
[0006] To address the aforementioned technical problems, this invention provides the following technical solution: a cloud-based ASON controller deployment method for the power Internet of Things, comprising the following steps: Deploy a distributed ASON controller cluster consisting of a resource scheduling engine, a routing decision engine, and a fault recovery engine; Monitor the operating status of the power communication network, issue alarm information when a fault occurs in the power communication network, and generate a list of circuits to be protected based on the parsing of the alarm information; The routing decision engine obtains the circuit operation data of the circuits in the optical transmission network that are in the list of circuits to be protected based on the list of circuits to be protected, and constructs a virtual topology model. The routing policy is obtained by performing route pre-calculation based on the virtual topology model through the routing decision engine; Deploy lightweight proxies at edge nodes in various cities and distribute routing decisions to the lightweight proxies; When a fault is detected in the circuit to be protected, the fault recovery process is executed by the fault recovery engine and the lightweight agent according to the routing policy.
[0007] As a preferred embodiment of the cloud-based ASON controller deployment method for the power Internet of Things described in this invention, the step of deploying a distributed ASON controller cluster consisting of a resource scheduling engine, a routing decision engine, and a fault recovery engine includes: Deploy the resource scheduling engine, routing decision engine, and fault recovery engine to the power cloud platform to build a distributed ASON controller cluster; The resource scheduling engine allocates communication resources to communication tasks. The routing decision engine performs route pre-calculation, generates routing policies, and publishes the routing policies. The fault recovery engine subscribes to the routing policies published by the routing decision engine and loads the routing policies into the local cache.
[0008] The beneficial effects of this preferred technical solution are as follows: By deploying a distributed ASON controller cluster, which includes a resource scheduling engine, a routing decision engine, and a fault recovery engine, on a power private cloud platform, the problems of centralized controllers easily becoming network bottlenecks and single-point failures causing the entire network control function to be paralyzed are solved. As a preferred embodiment of the cloud-based ASON controller deployment method for the power Internet of Things described in this invention, the steps of monitoring the operating status of the power communication network, issuing alarm information when a fault occurs in the power communication network, and generating a list of circuits to be protected based on the alarm information include: Monitor the operating status of the power communication network and issue alarm information when a fault is detected in the power communication network; Semantic analysis and keyword extraction are performed on the text description of each alarm message to obtain parsed information; Based on the parsed information, the circuits in the pre-stored circuit profile database are filtered to obtain a list of circuits to be protected.
[0009] As a preferred embodiment of the cloud-based ASON controller deployment method for the power Internet of Things described in this invention, the step of obtaining circuit operation data of circuits in the optical transmission network that are in the list of circuits to be protected based on the list of circuits to be protected by the routing decision engine, and constructing a virtual topology model includes: Obtain the circuit operation data for each circuit in the list of circuits to be protected; The integrated circuit operation data is obtained by integrating the circuit operation data; A virtual topology model is built using the routing decision engine based on integrated operational data.
[0010] As a preferred embodiment of the cloud-based ASON controller deployment method for the power Internet of Things described in this invention, the step of obtaining the routing strategy by performing route pre-calculation based on the virtual topology model through the routing decision engine includes: Historical failure rate information and current optical path bit error rate are collected and combined with a virtual topology model as input features; Perform route pre-computation based on input features to generate candidate detour paths; The reliability of all candidate detour paths is scored using a random forest model through a routing decision engine. The routes are sorted in descending order based on reliability scores, and a set number of candidate detour paths with the highest scores are selected as the routing strategy, while a number of dedicated time slots are reserved.
[0011] The beneficial effects of this preferred technical solution are as follows: by allocating dedicated resources to video conferencing services through a resource scheduling engine, and by dynamically calculating and reserving a dedicated time slot pool based on the importance of the services through a routing decision engine, on-demand supply and intelligent optimization of resources are achieved, solving the problem of low resource utilization caused by static allocation of time slots in the traditional method and inability to adapt to sudden traffic.
[0012] As a preferred embodiment of the cloud-based ASON controller deployment method for the power Internet of Things described in this invention, the steps of deploying lightweight agents at edge nodes in various cities and distributing routing decisions to the lightweight agents include: Deploy lightweight proxies at edge nodes in various cities; The system sends routing policy update information to the lightweight agent at set intervals and preloads the routing policy update information into local memory.
[0013] As a preferred embodiment of the cloud-based ASON controller deployment method for the power Internet of Things described in this invention, the steps of executing the fault recovery process according to the routing policy through the fault recovery engine and lightweight agent when a fault is detected in the circuit to be protected include: Detect the operating status of all circuits in the list of circuits to be protected; When a fault is detected in the circuit to be protected, the fault recovery engine releases the resource occupation of all structures in the circuit corresponding to the fault. Set a delay window period, and after the delay window period ends, the lightweight agent executes the fault recovery process according to the routing policy.
[0014] The beneficial effects of this preferred technical solution are as follows: when a fault occurs, the fault recovery engine and the lightweight agent work together to execute an automated recovery process of first deleting and then rebuilding. At the same time, combined with the cloud-edge collaborative architecture, the lightweight agent can preload policies and perform local autonomous recovery, which solves the problem that fault handling relies on manual intervention and recovery takes a long time.
[0015] This invention provides a cloud-based ASON controller deployment system for the power Internet of Things.
[0016] To address the aforementioned technical problems, the present invention further provides the following technical solution: a cloud-based ASON controller deployment system for the power Internet of Things, comprising: Controller deployment module: Deploys a distributed ASON controller cluster consisting of a resource scheduling engine, a routing decision engine, and a fault recovery engine; Resource allocation module: Allocates communication resources to communication tasks through the resource scheduling engine; Alarm information parsing module: monitors the operating status of the power communication network, issues alarm information when a fault occurs in the power communication network, and generates a list of circuits to be protected based on the parsing of the alarm information; Topology model construction module: Based on the list of circuits to be protected, the routing decision engine obtains the circuit operation data of the circuits in the optical transmission network that are in the list of circuits to be protected, and constructs a virtual topology model; Data computation module: The routing decision engine performs intelligent route pre-calculation based on the virtual topology model to obtain routing policies; Data delivery module: Deploy lightweight agents at edge nodes in various cities and distribute routing decisions to the lightweight agents; Fault recovery module: When a fault is detected in the circuit to be protected, the fault recovery engine and lightweight agent execute the fault recovery process according to the routing policy.
[0017] The present invention provides a computer device, including a memory and a processor, wherein the memory stores a computer program, characterized in that the processor executes the computer program to implement the steps of the cloud-based ASON controller deployment method for the power Internet of Things.
[0018] The present invention provides a computer-readable storage medium having a computer program stored thereon, characterized in that, when the computer program is executed by a processor, it implements the steps of the cloud-based ASON controller deployment method for the power Internet of Things.
[0019] The beneficial effects of this invention are as follows: By deploying a distributed ASON controller cluster, including a resource scheduling engine, a routing decision engine, and a fault recovery engine, on a power private cloud platform, this invention solves the problems of centralized controllers easily becoming network bottlenecks and single-point failures causing the entire network control function to be paralyzed. Through the resource scheduling engine, dedicated resources are allocated to video conferencing services. Combined with the routing decision engine, dedicated time slot pools are dynamically calculated and reserved based on service importance, achieving on-demand resource supply and intelligent optimization. This solves the problems of low resource utilization caused by the static allocation of time slots in traditional methods, which cannot adapt to sudden traffic surges. By automatically parsing alarm information and utilizing multiple pre-calculated detour routes based on the circuit to be protected, an automated recovery process of deletion followed by reconstruction is executed collaboratively by the fault recovery engine and a lightweight agent when a fault occurs. Simultaneously, the cloud-edge collaborative architecture enables the lightweight agent to pre-load policies and perform local autonomous recovery, solving the problems of fault handling relying on manual intervention and long recovery times. Attached Figure Description
[0020] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0021] Figure 1 This is a general flowchart of a cloud-based ASON controller deployment method for the power Internet of Things, provided as an embodiment of the present invention. Detailed Implementation
[0022] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the protection scope of the present invention.
[0023] Example 1, referring to Figure 1 This is the first embodiment of the present invention, which provides a cloud-based ASON controller deployment method for the power Internet of Things, including: S100: Deploy a distributed ASON controller cluster consisting of a resource scheduling engine, a routing decision engine, and a fault recovery engine; S200: Monitors the operating status of the power communication network, issues alarm information when a fault occurs in the power communication network, and generates a list of circuits to be protected based on the alarm information; S300: The routing decision engine obtains the circuit operation data of the circuits in the optical transmission network that are in the list of circuits to be protected based on the list of circuits to be protected, and constructs a virtual topology model. S400: Obtains routing policies by performing intelligent route pre-calculation based on the virtual topology model through the routing decision engine; S500: Deploy lightweight agents at edge nodes in various cities and distribute routing decisions to the lightweight agents; S600: When a fault is detected in the circuit to be protected, the fault recovery engine and lightweight agent execute the fault recovery process according to the routing policy.
[0024] It should be noted that, as the core architecture of power communication networks, automated switched optical networks (AS / ORNs) are responsible for core functions such as routing calculation, connection management, and fault recovery. The deployment and performance of their controllers directly affect the stability of the entire network and the quality of service transmission. The centralized controllers in existing AS / ORNs are prone to becoming network bottlenecks, and once a failure occurs, it may paralyze the control functions of the entire network. Traditional deployment methods use static allocation of time slot resources, which cannot dynamically adapt to sudden traffic surges in video conferencing, resulting in low resource utilization. The fault handling process still heavily relies on manual intervention, and the entire process is time-consuming. Therefore, the controller of AS / ORNs is very important for AS / ORNs.
[0025] Therefore, to address the aforementioned centralized controller operation and maintenance issues, a cloud-based ASON controller deployment method for the power Internet of Things (IoT) is constructed through steps S100~S700. By deploying a distributed ASON controller cluster including a resource scheduling engine, a routing decision engine, and a fault recovery engine on a power private cloud platform, the problems of centralized controllers easily becoming network bottlenecks and single-point failures causing the entire network control function to be paralyzed are solved. The resource scheduling engine allocates dedicated resources for video conferencing services, and the routing decision engine dynamically calculates and reserves dedicated time slot pools based on the importance of services, realizing on-demand supply and intelligent optimization of resources, solving the problems of low resource utilization caused by static allocation of time slots in traditional methods and inability to adapt to sudden traffic. By automatically parsing alarm information and using multiple pre-calculated detour routes based on the circuit to be protected, the fault recovery engine and lightweight agent jointly execute an automated recovery process of deletion and reconstruction when a fault occurs. At the same time, combined with the cloud-edge collaborative architecture, the lightweight agent can preload policies and perform local autonomous recovery, solving the problems of fault handling relying on manual intervention and long recovery time.
[0026] Example 2, refer to Figure 1 This is the second embodiment of the present invention, which provides a cloud-based ASON controller deployment method for the power Internet of Things.
[0027] In this embodiment of the application, the deployment of the distributed ASON controller cluster consisting of a resource scheduling engine, a routing decision engine, and a fault recovery engine in step S100 includes the following steps A1-A4: A1: Deploy the resource scheduling engine, routing decision engine, and fault recovery engine to the power cloud platform to build a distributed ASON controller cluster; In this embodiment of the application, in the power private cloud platform, the resource scheduling engine, routing decision engine and fault recovery engine are packaged into Docker container images respectively. The container images of the resource scheduling engine, routing decision engine and fault recovery engine are deployed as independent microservices through the Kubernetes open source system. The three microservices communicate with each other through message queues and each has its own independent namespace and network policy, thereby building a distributed ASON controller cluster.
[0028] A2: Allocate communication resources for communication tasks through the resource scheduling engine; In this embodiment, the communication task takes video conferencing service as an example. When a new video conferencing service is identified, the resource scheduling engine sends a request to the cluster management node through the Kubernetes open-source system to allocate a dedicated resource quota of 2 CPU cores and 8GB of memory to the video conferencing service. This ensures that other non-critical services cannot preempt these dedicated resource quotas, thereby guaranteeing the stable operation of the routing calculation task. When the video conferencing service ends and the CPU utilization rate remains below 30% for more than ten minutes, the resource scheduling engine releases the computing resources occupied by the video conferencing service back to the resource pool.
[0029] A3: Perform route pre-calculation through the route decision engine, generate route policies, and publish the route policies; In this embodiment, a routing decision engine performs route pre-calculation to generate a routing strategy containing multiple detour paths. The routing strategy is then serialized into a JSON message and published to a specific topic with a set name in a message queue. This ensures that the routing strategy can be reliably obtained by the fault recovery engine.
[0030] A4: The fault recovery engine subscribes to the routing policies published by the routing decision engine and loads the routing policies into the local cache.
[0031] In this embodiment, the fault recovery engine registers as a consumer of a specific topic with a set name in the message queue and continuously listens to the message stream of that specific topic with the set name. When the routing decision engine publishes a new routing policy message, the fault recovery engine will immediately receive the new routing policy message, parse it, and load it into its own local memory cache. In this way, the fault recovery engine always holds the latest routing policy.
[0032] In one alternative implementation, the distributed ASON controller cluster can also be constructed using a hybrid cloud deployment model. The core routing decision engine and resource scheduling engine are deployed in the power company's private cloud data center to ensure that critical decision-making logic and sensitive data remain on the internal network. At the same time, lightweight fault recovery engine instances are deployed on public clouds close to the edge nodes of various cities. An encrypted secure channel is established between the private cloud and the public cloud edge nodes through an IPsec or TLS tunnel. When the provincial center detects a fault, the command can quickly reach the city node through the low-latency public cloud edge network, and the local fault recovery engine instance will perform fault recovery.
[0033] In another alternative implementation, the distributed ASON controller cluster can also be constructed using the OpenStack virtual machine deployment mode. The resource scheduling engine, routing decision engine, and fault recovery engine are each encapsulated as independent virtual machine images and deployed in the OpenStack environment of the power private cloud platform. Each virtual machine is isolated through VLANs to ensure secure and stable communication between modules. The three modules, namely the resource scheduling engine, routing decision engine, and fault recovery engine, communicate asynchronously through RabbitMQ message queues to decouple system components and control the execution of fault recovery.
[0034] In this embodiment of the application, step S200, which involves monitoring the operating status of the power communication network, issuing alarm information when a fault occurs in the power communication network, and generating a list of circuits to be protected based on the alarm information, includes the following steps B1-B3: B1: Monitors the operating status of the power communication network and issues an alarm message when a fault is detected in the power communication network; In this embodiment of the application, the signal status of the physical port is continuously monitored by the optical transmission equipment in the power communication network. When the received optical power of a certain VC4 level SDH circuit is detected to be lower than -28dBm or the continuous bit error seconds exceed 5, it is determined to be a link failure. This link failure is reported to the network management system, and the network management system generates a structured alarm message and pushes it to the distributed ASON controller cluster.
[0035] B2: Perform semantic analysis and keyword extraction on the text description of each alarm message to obtain parsed information; In this embodiment of the application, after the distributed ASON controller cluster receives the alarm information generated in step B1, it uses natural language processing technology and a keyword matching algorithm to parse out keywords directly related to the communication task from the alarm information, and outputs these keywords as core parsing information. In this application, the communication task takes video conferencing service as an example. The keywords directly related to video conferencing service are video conferencing terminal and codec.
[0036] In one alternative implementation, semantic analysis and keyword extraction of the text description in each alarm message can also be achieved through machine learning. A training dataset consisting of historical alarm messages is pre-constructed, and each alarm message is manually labeled to determine whether it belongs to the video conferencing category. The text is converted into a high-dimensional feature vector using the TF-IDF algorithm, and then a binary classification model is trained using a support vector machine. When running online, new alarm messages are vectorized using the same TF-IDF algorithm and then fed into the trained SVM model for prediction, outputting the business category to which the alarm message belongs.
[0037] In another optional implementation, semantic analysis and keyword extraction of the text description in each alarm message can also be achieved through a BERT pre-trained language model. A BERT model finely tuned on massive amounts of Chinese power industry text is loaded. When an alarm message is received, the alarm message is segmented and vectorized before being input into the BERT model. Through the contextual understanding capability of the BERT model, it can not only identify explicit keywords, but also understand the semantic expression of implicit video conferencing services. The BERT model will output a classification probability, and then compare the classification probability with a preset threshold to determine the specific business category based on the comparison result.
[0038] B3: Based on the parsed information, the circuits in the pre-stored circuit profile database are filtered to obtain a list of circuits to be protected.
[0039] In this embodiment of the application, the keywords parsed in step B2 are matched with the circuit attribute information pre-stored in the database to filter out the VC4-level SDH circuits corresponding to the alarm information in step B2 and generate a list of circuits to be protected.
[0040] In this embodiment of the application, step S300, which involves obtaining circuit operation data of circuits in the optical transmission network that are in the list of circuits to be protected based on the list of circuits to be protected by the routing decision engine, and constructing a virtual topology model, includes the following steps C1-C3: C1: Obtain the circuit operation data for each circuit in the list of circuits to be protected; In this embodiment of the application, after receiving the list of circuits to be protected generated in step B3, the routing decision engine sends a data acquisition command to the network management system to collect all the operating data along the paths of all circuits in the list of circuits to be protected, including the network element information of each node, port connection status, VC4 time slot occupancy status, and the status of the protection subnet configured for the circuit. C2: Integrate the circuit operation data to obtain integrated operation data; In this embodiment, the network management system assigns a unique global identifier to each circuit to be protected. The routing decision engine obtains data from the network management system, including network element information of each node, port connection status, VC4 time slot occupancy, and protection subnet status data of the circuit configuration. Then, using the global identifier as the key, it matches and integrates the data from the network management system, including network element information of each node, port connection status, VC4 time slot occupancy, and protection subnet status data of the circuit configuration, to generate integrated operation data containing multiple circuit operation data. C3: Constructs a virtual topology model based on integrated runtime data through the routing decision engine.
[0041] In this embodiment of the application, the routing decision engine, based on the integrated operation data generated in step C2, integrates the network element information, port connection status, and VC4 timeslot occupancy status of the occupied node into a working route for each circuit in the list of circuits to be protected, with the global identifier as the root node; integrates the protection subnet status data of each node into a protection route; and integrates the network element information, port connection status, and VC4 timeslot occupancy status of the unoccupied node into an idle resource pool. The working route, the protection route, and the idle resource pool are then combined into an independent virtual topology model.
[0042] In this embodiment of the application, step S400, in which the routing decision engine performs route pre-calculation based on the virtual topology model to obtain the routing policy, includes the following steps D1-D4: D1: Collect historical failure rate information and current optical path bit error rate, and combine them with a virtual topology model as input features; In this embodiment of the application, collecting historical failure rate information includes obtaining the average failure frequency of each optical path over a period of time from a long-term maintained historical alarm database to form a historical failure rate matrix; obtaining the current optical path bit error rate by monitoring the bit error rate of each optical path; and combining the current optical path bit error rate and the historical failure rate matrix with the virtual topology model constructed in step C3 as input features. D2: Perform route pre-computation based on input features to generate candidate detour paths; In this embodiment of the application, the route pre-calculation based on the input features includes using a depth-first search to traverse all feasible path combinations that can connect the start and end points of the circuit to be protected based on the idle resource pool in the virtual topology model of the input features; by searching and traversing in the idle resource pool, the fault node is avoided, and the feasible path combinations obtained by the traversal are output as candidate detour paths. D3: The reliability of all candidate detour paths is scored using a random forest model through the routing decision engine; In this embodiment, the input features obtained in step D1 are input into a pre-trained random forest model through a routing decision engine. The positive class probability output by the random forest model is then used as the reliability score for each path among the candidate detour paths obtained in step D2. The random forest model consists of T decision trees, and each decision tree outputs an initial score for a certain path. The formula for calculating the reliability score using the initial score is as follows: ; in, Score the reliability of path l. Let T be the initial score of path l obtained through the T-th decision tree, where T is the number of decision trees and l is the number of paths in the candidate detour paths; the reliability score, which does not depend on a single decision tree, is obtained through the reliability score calculation formula.
[0043] In an alternative implementation, reliability scoring of all candidate detour paths can also be achieved through a weighted linear combination. Weights are assigned to multiple evaluation dimensions of each candidate path, including historical failure rate, current bit error rate, and path hop count. The original values of each dimension are normalized and then the comprehensive score is calculated by weighted summation, which can be used as the reliability score of each candidate path.
[0044] In another alternative implementation, reliability scoring of all candidate detour paths can also be achieved through graph neural networks. The virtual topology model is transformed into a graph structure, where network elements are nodes and fiber optic links are edges. Features such as historical failure rate and current bit error rate are embedded as attributes of nodes and edges. The entire graph is aggregated in multiple layers using a graph convolutional network, and finally a reliability probability value is output for each candidate path. This reliability probability value can be used as the reliability score for each candidate path.
[0045] D4: Sort the candidate detour paths in descending order based on the reliability scores, and select the candidate detour paths with the highest scores as the routing strategy, while reserving a number of dedicated time slots.
[0046] In this embodiment of the application, the number is set to 3. Based on the reliability score calculated in step D3, the detour paths in the candidate detour paths are sorted in descending order, and the three detour paths with the highest reliability scores are selected as the routing strategy. These three detour paths will be used as the main path, the first backup path, and the second backup path, respectively. When the main path fails, the first backup path and the second backup path will be activated in sequence according to their respective times. In this embodiment of the application, the formula for the number of reserved dedicated time slots is: ; in, The number of dedicated time slots, To reserve a proportional coefficient, adjustments will be made based on business importance; S represents the total number of currently available time slots. It should be noted that when calculating the number of reserved dedicated time slots, the reservation ratio coefficient is adjusted to ensure that critical services have sufficient resource guarantees. This process of reserving the number of dedicated time slots provides resource redundancy for video conferencing services, avoids the inability to establish new paths due to sudden traffic or resource competition, and ensures that the routing strategy can be successfully executed during the fault recovery phase.
[0047] In this embodiment of the application, step S500, which involves deploying lightweight agents at edge nodes in various cities and distributing routing decisions to the lightweight agents, includes the following steps E1-E2: E1: Deploy lightweight agents at edge nodes in various cities; In this embodiment of the application, a lightweight agent program with simplified functions and low resource consumption is deployed in the power communication equipment rooms of various cities. This lightweight agent acts as a local extension of the ASON controller cluster and is responsible for receiving and executing control commands from the ASON controller cluster.
[0048] E2: At set intervals, send routing policy update information to the lightweight agent and preload the routing policy update information into local memory.
[0049] In this embodiment of the application, the period is set to thirty minutes. Every thirty minutes, the distributed ASON controller cluster generates a policy update package based on the routing policy generated in step D4 and sends it to the lightweight agents in various cities. The policy update package only contains the routing entries that have changed compared with the routing policy sent in the previous period. After receiving the update package, the lightweight agent preloads the policy update package into its local memory for caching, ensuring that it can be directly invoked in the event of a failure, so as to achieve self-recovery quickly.
[0050] In one alternative implementation, the distribution of routing policy update information to the lightweight agent can also be achieved through event-driven real-time push. When the lightweight agent starts up, it subscribes to the message bus in the cloud and registers the city and region it governs. When the routing decision engine generates a new routing policy, it is not distributed periodically, but is immediately published to the bus as a high-priority event message. The message bus accurately pushes the event to the corresponding lightweight agent in the city based on the regional label. After receiving the message, the lightweight agent immediately parses and loads the new policy into memory, and at the same time confirms successful reception to the cloud.
[0051] In another alternative implementation, routing policy update information can be distributed to lightweight agents securely and reliably using blockchain. The provincial cloud center is set as the authoritative node, and the lightweight agents in various cities are set as ordinary nodes, forming a private consortium blockchain. Each generated routing policy update package is hashed and recorded on the blockchain, forming an immutable operation log. The lightweight agent not only receives the update package from the cloud, but also obtains the hash value of the update from the blockchain network for verification. Only when the hash value of the locally decompressed update package is completely consistent with the record on the blockchain is it allowed to be preloaded into local memory.
[0052] In this embodiment of the application, when a fault is detected in the circuit to be protected in step S600, the fault recovery process is executed by the fault recovery engine and the lightweight agent according to the routing policy, including the following steps F1-F2: F1: Detects the operating status of all circuits in the list of circuits to be protected; when a fault is detected in a circuit to be protected, the fault recovery engine releases the resource occupation of all structures in the circuit corresponding to the fault. In this embodiment, the signal status of physical ports is continuously monitored by optical transmission equipment in the power communication network. When the received optical power of a certain VC4-level SDH circuit is detected to be lower than -28dBm or the number of consecutive bit error seconds exceeds 5, it is determined to be a link failure. The fault recovery engine promptly releases the resource occupation of all occupied network elements, ports and VC4 time slots on the circuit corresponding to the faulty link. By releasing the resource occupation of the circuit corresponding to the faulty link, it is ensured that network resources will not be locked for a long time due to the fault, thus clearing the way for the establishment of new paths.
[0053] F2: Set a delay window period, and after the delay window period ends, the lightweight agent performs the fault recovery process according to the routing policy.
[0054] In this embodiment, after the resource occupancy is released in step F1, an adjustable delay window is entered to avoid signaling storms. The length of the delay window is 5 to 1000 seconds, and its value is set by the operation and maintenance personnel according to the network scale and experience. After the delay window ends, the lightweight agent in each city acts as a local execution unit, activates an optimal detour route according to the preloaded routing policy, and allocates the number of time slots calculated in step D4 to establish a new path, thereby completing the entire fault recovery process.
[0055] Example 3, referring to Figure 1 This is a third embodiment of the present invention, which provides a cloud-based ASON controller deployment system for the power Internet of Things, comprising: Controller deployment module: Deploys a distributed ASON controller cluster consisting of a resource scheduling engine, a routing decision engine, and a fault recovery engine; Resource allocation module: Allocates communication resources to communication tasks through the resource scheduling engine; Alarm information parsing module: monitors the operating status of the power communication network, issues alarm information when a fault occurs in the power communication network, and generates a list of circuits to be protected based on the parsing of the alarm information; Topology model construction module: Based on the list of circuits to be protected, the routing decision engine obtains the circuit operation data of the circuits in the optical transmission network that are in the list of circuits to be protected, and constructs a virtual topology model; Data computation module: The routing decision engine performs intelligent route pre-calculation based on the virtual topology model to obtain routing policies; Data delivery module: Deploy lightweight agents at edge nodes in various cities and distribute routing decisions to the lightweight agents; Fault recovery module: When a fault is detected in the circuit to be protected, the fault recovery engine and lightweight agent execute the fault recovery process according to the routing policy.
[0056] Example 4, the fourth embodiment of the present invention, differs from the previous three embodiments in that: if the function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0057] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-including system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device.
[0058] More specific examples of computer-readable media (a non-exhaustive list) include: electrical connections (electronic devices) having one or more wires, portable computer disk drives (magnetic devices), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Furthermore, computer-readable media can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in computer memory.
[0059] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0060] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A cloud-based ASON controller deployment method for the power Internet of Things, characterized in that, include, Deploy a distributed ASON controller cluster consisting of a resource scheduling engine, a routing decision engine, and a fault recovery engine; Monitor the operating status of the power communication network, issue alarm information when a fault occurs in the power communication network, and generate a list of circuits to be protected based on the parsing of the alarm information; The routing decision engine obtains the circuit operation data of the circuits in the optical transmission network that are in the list of circuits to be protected based on the list of circuits to be protected, and constructs a virtual topology model. The routing policy is obtained by performing route pre-calculation based on the virtual topology model through the routing decision engine; Deploy lightweight proxies at edge nodes in various cities and distribute routing decisions to the lightweight proxies; When a fault is detected in the circuit to be protected, the fault recovery process is executed by the fault recovery engine and the lightweight agent according to the routing policy.
2. The cloud-based ASON controller deployment method for the power Internet of Things as described in claim 1, characterized in that, The steps for deploying a distributed ASON controller cluster consisting of a resource scheduling engine, a routing decision engine, and a fault recovery engine include: Deploy the resource scheduling engine, routing decision engine, and fault recovery engine to the power cloud platform to build a distributed ASON controller cluster; The resource scheduling engine allocates communication resources to communication tasks. The routing decision engine performs route pre-calculation, generates routing policies, and publishes the routing policies. The fault recovery engine subscribes to the routing policies published by the routing decision engine and loads the routing policies into the local cache.
3. The cloud-based ASON controller deployment method for the power Internet of Things as described in claim 2, characterized in that, The steps for monitoring the operating status of the power communication network, issuing alarm information when a fault occurs in the power communication network, and generating a list of circuits to be protected based on the parsing of the alarm information include: Monitor the operating status of the power communication network and issue alarm information when a fault is detected in the power communication network; Semantic analysis and keyword extraction are performed on the text description of each alarm message to obtain parsed information; Based on the parsed information, the circuits in the pre-stored circuit profile database are filtered to obtain a list of circuits to be protected.
4. The cloud-based ASON controller deployment method for the power Internet of Things as described in claim 3, characterized in that, The steps of using a routing decision engine to obtain circuit operation data of circuits in the optical transmission network that match those in the list of circuits to be protected, and to construct a virtual topology model, include: Obtain the circuit operation data for each circuit in the list of circuits to be protected; The integrated circuit operation data is obtained by integrating the circuit operation data; A virtual topology model is built using the routing decision engine based on integrated operational data.
5. The cloud-based ASON controller deployment method for the power Internet of Things as described in claim 4, characterized in that, The steps involved in obtaining routing policies through route pre-calculation using a routing decision engine based on a virtual topology model include: Historical failure rate information and current optical path bit error rate are collected and combined with a virtual topology model as input features; Perform route pre-computation based on input features to generate candidate detour paths; The reliability of all candidate detour paths is scored using a random forest model through a routing decision engine. The routes are sorted in descending order based on reliability scores, and a set number of candidate detour paths with the highest scores are selected as the routing strategy, while a number of dedicated time slots are reserved.
6. The cloud-based ASON controller deployment method for the power Internet of Things as described in claim 5, characterized in that, The steps for deploying lightweight proxies at edge nodes in various cities and distributing routing decisions to the lightweight proxies include: Deploy lightweight proxies at edge nodes in various cities; The system sends routing policy update information to the lightweight agent at set intervals and preloads the routing policy update information into local memory.
7. A cloud-based ASON controller deployment method for the power Internet of Things as described in claim 6, characterized in that, When a fault is detected in the circuit to be protected, the steps for executing the fault recovery process according to the routing policy through the fault recovery engine and lightweight agent include: Detect the operating status of all circuits in the list of circuits to be protected; When a fault is detected in the circuit to be protected, the fault recovery engine releases the resource occupation of all structures in the circuit corresponding to the fault. Set a delay window period, and after the delay window period ends, the lightweight agent executes the fault recovery process according to the routing policy.
8. A cloud-based ASON controller deployment system for the power Internet of Things, employing the cloud-based ASON controller deployment method for the power Internet of Things as described in any one of claims 1 to 7, characterized in that, include: Controller deployment module: Deploys a distributed ASON controller cluster consisting of a resource scheduling engine, a routing decision engine, and a fault recovery engine; Resource allocation module: Allocates communication resources to communication tasks through the resource scheduling engine; Alarm information parsing module: monitors the operating status of the power communication network, issues alarm information when a fault occurs in the power communication network, and generates a list of circuits to be protected based on the parsing of the alarm information; Topology model construction module: Based on the list of circuits to be protected, the routing decision engine obtains the circuit operation data of the circuits in the optical transmission network that are in the list of circuits to be protected, and constructs a virtual topology model; Data computation module: The routing decision engine performs intelligent route pre-calculation based on the virtual topology model to obtain routing policies; Data delivery module: Deploy lightweight agents at edge nodes in various cities and distribute routing decisions to the lightweight agents; Fault recovery module: When a fault is detected in the circuit to be protected, the fault recovery engine and lightweight agent execute the fault recovery process according to the routing policy.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the cloud-based ASON controller deployment method for the power Internet of Things as described in any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the cloud-based ASON controller deployment method for the power Internet of Things as described in any one of claims 1 to 7.