On-chip optical network node fault self-healing method and system, electronic equipment and storage medium
By introducing data monitoring, fault detection and self-healing strategy operation modules into the on-chip optical network, self-healing strategy adjustment is performed for different topological structures, solving the problem of rapid response and adaptive adjustment in node failures, and achieving efficient and reliable operation of the network.
Patent Information
- Application Number
- CN202510501575.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-08-19
AI Technical Summary
Existing on-chip optical networks lack the ability to quickly self-heal when facing node failures, making it difficult to achieve efficient and reliable adaptive adjustments, affecting network stability and sustainability.
The data monitoring module is used to obtain node health status data, the fault detection module identifies the fault node, the self-healing strategy determination module generates a self-healing strategy, and resets the node link through the self-healing strategy operation module, including taking corresponding self-healing measures under different topological structures, such as building redundant loops, adjusting micro-ring resonator parameters, enabling optical switch matrix, and selecting backup nodes.
It realizes rapid response and adaptive adjustment in the event of node failure, reduces the impact of failure on network performance, ensures seamless transmission of data packets, and improves network reliability and flexibility.
Smart Images

Figure CN120512618A_ABST
Abstract
Description
Technical Field
[0001] One or more embodiments of the present disclosure relate to the field of optical communication network technology, and in particular to a method, system, electronic device, and storage medium for self-healing of an on-chip optical network node failure. Background Art
[0002] With the rapid development of integrated circuit technology, the number of transistors integrated on a chip has exploded, driven by Moore's Law. This has led to a sharp increase in the demand for data transmission within the chip. The traditional on-chip bus architecture is limited by bandwidth bottlenecks and high power consumption, and it is difficult to meet the growing communication requirements. In this context, photonics technology has become a key path to break through the difficulties of internal chip communication with its high bandwidth, low latency and excellent anti-electromagnetic interference capabilities. Optical Network-on-Chip (ONoC) has emerged. It uses optical signals for data transmission, providing a new mechanism for internal chip communication. It can achieve long-distance, high-speed communication within a single package system, significantly reducing latency and power consumption. It has shown great potential in application scenarios such as multi-core processors and high-performance computing, and is gradually becoming the focus of industry research. Its development is of great significance to meeting the challenges of complex multi-core system design in the future.
[0003] However, current topology fault-tolerance technologies for on-chip optical networks generally lack the ability to adaptively adjust to diverse network topologies, performing poorly in real-time fault detection and dynamic response, and failing to meet the stringent requirements for high reliability and durability in complex network environments. Furthermore, these technologies primarily focus on path decision-making, aiming to improve data transmission efficiency, but ignore the frequent node failures in actual network operations. They fail to develop efficient and reliable self-healing methods for node failure scenarios, making it difficult for networks to quickly resume normal operations in the face of failures, severely impacting network stability and service continuity.
[0004] Therefore, it is urgent to develop an on-chip optical network topology self-healing technology that can adapt to different topologies, have rapid self-healing capabilities, and achieve effective dynamic response and adaptive adjustment when nodes fail. Summary of the Invention
[0005] In view of this, an object of one or more embodiments of the present disclosure is to provide a method, system, electronic device and storage medium for self-healing of an on-chip optical network node failure, so as to solve the problems raised in the background technology.
[0006] Based on the above objectives, one or more embodiments of the present disclosure provide a self-healing method for an on-chip optical network node failure, which is applied to a self-healing system for an on-chip optical network node failure. The system includes a data monitoring module, a fault detection module, a self-healing strategy determination module, and a self-healing strategy operation module. The method includes:
[0007] The data monitoring module obtains the health status data of each node in the on-chip optical network;
[0008] A fault detection module determines a faulty node of the on-chip optical network based on the health status data;
[0009] A self-healing strategy determination module determines a self-healing strategy corresponding to the on-chip optical network in response to a faulty node in the optical network being unable to recover;
[0010] The self-healing strategy operation module resets the node links of the on-chip optical network according to the self-healing strategy.
[0011] Optionally, resetting the node links of the on-chip optical network according to the self-healing strategy includes:
[0012] The standby node of the on-chip optical network receives a control instruction generated according to the self-healing strategy;
[0013] The standby node is activated according to the control instruction, and a global routing table of the on-chip optical network is updated, where the global routing table is used to store routing information of all nodes of the on-chip optical network.
[0014] Optionally, in response to the failure node of the optical network being unable to recover, determining a self-healing strategy corresponding to the on-chip optical network includes:
[0015] In response to a failed node of the on-chip optical network being unable to recover, determining a topology of the on-chip optical network;
[0016] A self-healing strategy corresponding to the on-chip optical network is generated according to the topological structure of the on-chip optical network and the node information of the on-chip optical network.
[0017] Optionally, in response to the topology of the on-chip optical network being a ring topology, resetting the node links of the on-chip optical network according to the self-healing strategy includes:
[0018] Adjusting parameters of microring resonators of all optical routers in the on-chip optical network to construct a redundant loop in the opposite direction of the faulty loop, which serves as a physical basis for optical link adjustment;
[0019] Adjust optical links through the optical switch matrix to enable bidirectional path synchronous transmission;
[0020] Obtaining a proximity score of the backup node of the on-chip optical network according to the distance between the on-chip optical network fault node and the backup node of the on-chip optical network, the signal quality of the link between the on-chip optical network nodes, and the historical failure frequency of the backup node of the on-chip optical network;
[0021] A target backup node of the on-chip optical network is determined according to the proximity score, and the target backup node of the on-chip optical network is used to replace the faulty node of the on-chip optical network to construct a transmission link.
[0022] Optionally, in response to the topology of the on-chip optical network being a star topology, resetting the node links of the on-chip optical network according to the self-healing strategy includes:
[0023] Determining a shadow central node corresponding to a faulty node of the on-chip optical network;
[0024] Constructing a transmission link between a peripheral node of the faulty node of the on-chip optical network and the shadow central node;
[0025] The microring resonator of the shadow central node is adjusted to distinguish the wavelength of the shadow central node from the wavelength of the faulty node of the on-chip optical network.
[0026] Optionally, in response to the topology of the on-chip optical network being a tree topology, resetting the node links of the on-chip optical network according to the self-healing strategy includes:
[0027] Inputting the subnode status of the faulty node of the on-chip optical network and the current load of the on-chip optical network into a preset reinforcement learning model to obtain an optimal alternative path corresponding to the faulty node of the on-chip optical network;
[0028] In response to the optical power loss of the optimal alternative path being greater than a preset threshold, an idle sub-node within a nearby range is searched for to fill the position.
[0029] Optionally, in response to the topology of the on-chip optical network being a mesh grid or a two-dimensional torus network topology, resetting the node links of the on-chip optical network according to the self-healing strategy includes:
[0030] Deploy the nodes of the on-chip optical network in a micro-blockchain framework, elect a temporary master node in the vicinity of the faulty node of the on-chip optical network, and establish a temporary transmission link;
[0031] Inputting the node coordinate matrix, historical traffic data, and current link load of the on-chip optical network into a preset spatiotemporal graph convolutional network to obtain a predicted traffic map, wherein the predicted traffic map shows potential congested areas of the on-chip optical network;
[0032] In response to determining that the proportion of faulty nodes is less than a preset threshold, an optimal alternative path is obtained according to the predicted traffic graph, and the improved ant colony algorithm sets an objective function with the goal of minimizing end-to-end delay and load imbalance;
[0033] In response to determining that the proportion of faulty nodes is greater than or equal to a preset threshold, the on-chip optical network is divided into a plurality of sub-networks and the plurality of sub-networks are converted into a tree topology structure.
[0034] Based on the same inventive concept, one or more embodiments of the present disclosure further provide a self-healing system for an on-chip optical network node failure, comprising: a data monitoring module, a fault detection module, a self-healing strategy determination module, and a self-healing strategy operation module;
[0035] The data monitoring module is configured to obtain health status data of each node in the on-chip optical network;
[0036] The fault detection module is configured to determine a faulty node of the on-chip optical network based on the health status data;
[0037] The self-healing strategy determination module is configured to determine a self-healing strategy corresponding to the on-chip optical network in response to a faulty node in the optical network being unable to recover;
[0038] The self-healing strategy operation module is configured to reset the node links of the on-chip optical network according to the self-healing strategy.
[0039] Based on the same inventive concept, one or more embodiments of the present disclosure also provide an electronic device, including a memory, a processor, and a computer program stored in the memory and runnable on the processor, wherein when the processor executes the program, it implements the self-healing method for on-chip optical network node failure as described in any one of the above items.
[0040] Based on the same inventive concept, one or more embodiments of the present disclosure also provide a non-transitory computer-readable storage medium, which stores computer instructions, and the computer instructions are used to enable the computer to execute any of the above-mentioned methods for self-healing of on-chip optical network node failures.
[0041] As can be seen from the above, the self-healing method for on-chip optical network node failure provided by one or more embodiments of the present disclosure is applied to the self-healing system for on-chip optical network node failure, the system including a data monitoring module, a fault detection module, a self-healing strategy determination module and a self-healing strategy operation module, the method including: the data monitoring module obtains the health status data of each node of the on-chip optical network; the fault detection module determines the faulty node of the on-chip optical network based on the health status data; the self-healing strategy determination module determines the self-healing strategy corresponding to the on-chip optical network in response to the faulty node of the optical network being unable to recover; the self-healing strategy operation module resets the node link of the on-chip optical network according to the self-healing strategy.
[0042] One or more embodiments of the present disclosure utilize a self-healing system to monitor the health of network nodes in real time, quickly detect and respond to node failures, and generate self-healing strategies to reset node links, thereby reducing the impact of failures on overall network performance.
[0043] The self-healing system, electronic device and computer-readable storage medium for on-chip optical network node failure provided by the present disclosure are all capable of implementing the steps of the above-mentioned self-healing method for on-chip optical network node failure, and therefore also have the beneficial effects of the above-mentioned self-healing method for on-chip optical network node failure. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] In order to more clearly illustrate one or more embodiments of the present disclosure or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only one or more embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0045] Figure 1 A flowchart of a method for self-healing a node failure in an on-chip optical network according to one or more embodiments of the present disclosure;
[0046] Figure 2 A schematic diagram of the structure of a self-healing system for an on-chip optical network node failure according to one or more embodiments of the present disclosure;
[0047] Figure 3 A schematic diagram of a self-healing solution for node failure in an on-chip optical network with a ring topology according to one or more embodiments of the present disclosure;
[0048] Figure 4 A schematic diagram of a self-healing solution for node failure in a star topology on-chip optical network according to one or more embodiments of the present disclosure;
[0049] Figure 5 A schematic diagram of a self-healing solution for a node failure in an on-chip optical network with a tree topology according to one or more embodiments of the present disclosure;
[0050] Figure 6 A schematic diagram of the hardware structure of an electronic device according to one or more embodiments of the present disclosure. DETAILED DESCRIPTION
[0051] In order to make the objectives, technical solutions and advantages of the present disclosure more clearly understood, the present disclosure is further described in detail below in conjunction with specific embodiments and with reference to the accompanying drawings.
[0052] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in one or more embodiments of the present disclosure should have the usual meanings understood by people with ordinary skills in the field to which the present disclosure belongs. The "first", "second" and similar words used in one or more embodiments of the present disclosure do not indicate any order, quantity or importance, but are only used to distinguish different components. "Include" or "comprise" and similar words mean that the elements or objects appearing before the word include the elements or objects listed after the word and their equivalents, without excluding other elements or objects. "Connect" or "connected" and similar words are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative positional relationships. When the absolute position of the described object changes, the relative positional relationship may also change accordingly.
[0053] The self-healing method for an on-chip optical network node failure according to one or more embodiments of the present disclosure is applied to Figure 2 The self-healing system for on-chip optical network node failure shown includes a data monitoring module, a fault detection module, a self-healing strategy determination module and a self-healing strategy operation module.
[0054] This system is implemented within a communications architecture encompassing the aforementioned system, signaling layer, and optical network layer. The signaling layer, controlled by the aforementioned system, includes electrical router nodes and electrical channels. The optical network layer includes service generation modules, optical router nodes, transmission waveguides, and E / O and O / E converters.
[0055] refer to Figure 1 , the method comprises the following steps:
[0056] Step S101: The data monitoring module obtains health status data of each node in the on-chip optical network;
[0057] Step S102: The fault detection module determines a faulty node of the on-chip optical network based on the health status data;
[0058] Step S103: In response to the failure of the optical network node being unable to recover, the self-healing strategy determination module determines a self-healing strategy corresponding to the on-chip optical network.
[0059] Step S104: the self-healing strategy operation module resets the node links of the on-chip optical network according to the healing strategy.
[0060] Before implementing the method provided by the present disclosure, the system may be initialized. During the system initialization phase, a comprehensive network operation foundation needs to be established.
[0061] Specifically, initialization includes: creating a global routing table to plan routes for data packet transmission and ensure the orderly flow of data within the network; simultaneously completing the initialization of the node health monitoring module to lay a solid foundation for real-time monitoring of node status and timely detection of potential faults. At the same time, the signaling layer is carefully configured to enable efficient data transmission and ensure the smooth interaction of network communication commands.
[0062] During topology design, it is necessary to ensure that the number of nodes is greater than the number of links in order to create a stable and reliable network infrastructure. Mature network topologies such as star, ring, or tree can be selected to meet the network performance requirements in different application scenarios.
[0063] Furthermore, the initial parameters of the off-chip controller can be appropriately set, enabling precise communication with on-chip nodes and achieving internal and external coordinated control. Simultaneously, a coordinate system is constructed to precisely assign unique spatial coordinates to each network node, anchoring the node in the network topology. This provides strong support for efficient subsequent fault location and comprehensive assessment, significantly improving network management and maintenance efficiency.
[0064] In one or more embodiments of the present disclosure, the data monitoring module can monitor the nodes of the on-chip network using an off-chip controller. The monitoring data may include parameters such as temperature, transmission power, node latency, data packet loss rate, and link quality, which can be collected by sensors (such as micro optical sensors, thermocouples, or electronic sensors). Data monitoring can be divided into two modes: full monitoring and incremental monitoring. In addition, real-time analysis of node behavior can be performed to identify possible abnormal fluctuations in order to process data more efficiently.
[0065] In the embodiment of the present disclosure, the data monitoring module may perform monitoring at a preset time according to settings, or perform monitoring after receiving an instruction or signal.
[0066] The above-mentioned fault detection module can identify node faults through a preset fault detection algorithm based on the above-mentioned health status data. The present disclosure does not limit the specific type or content of the above-mentioned fault detection algorithm. When identifying a node fault, the fault type and fault severity of the node can be identified. Among them, the type may include power drop, path interruption or node loss of connection, etc., and the severity of the fault can be evaluated based on whether the faulty node is a critical node (such as a central node or a high-load node) and the network impact that may be caused. In addition, the specific location where the fault occurs and the nature of the fault can also be determined.
[0067] When it is determined that the faulty node is completely failed and cannot be recovered, a self-healing operation is performed on the on-chip optical network.
[0068] It should be noted that, in the technical solution disclosed herein, the type of self-healing solution is determined according to the topology type of the on-chip optical network node.
[0069] In implementing the present disclosure, the applicant discovered that each topology type of on-chip optical network has unique structural characteristics and node distribution. For example, the ring topology is connected in a ring shape, the star topology is centered around a central node, the tree topology supports multi-level distribution with a hierarchical structure, and the mesh and torus topologies are distributed in a mesh-like manner without redundant nodes. Implementing adaptive self-healing strategies based on different topology types can effectively improve self-healing efficiency.
[0070] For a Ring topology, resetting the node links of the on-chip optical network according to the self-healing strategy may include:
[0071] First, the resonant wavelengths of the micro-ring resonators (MRRs) in all optical routers in the on-chip optical network are activated to construct Figure 3 As shown in the clockwise redundant loop. It can be seen that under normal conditions, data packets are transmitted through the counterclockwise link. When the node (0,0) fails, the clockwise link ( Figure 3 first link in ).
[0072] In this step, the Optical Switch Matrix (OSM) can be used to enable bidirectional synchronous transmission, ensuring multipath load balancing and reducing the impact of node failures on loop communication latency. Priority tags can also be used to distinguish between real-time services and disaster recovery data to avoid congestion.
[0073] Afterwards, by judging the proximity scores of other backup nodes in the vicinity of the faulty node, the backup node with the highest proximity score is determined as the target backup node to generate a new link (such as Figure 3 second link shown).
[0074] The proximity score can be obtained based on the distance between the faulty node and the backup node, the signal quality of the link between the nodes of the on-chip optical network, and the historical failure frequency of the backup nodes of the on-chip optical network. In some embodiments of the present disclosure, the mathematical model of the proximity score can be expressed as:
[0075]
[0076] Wherein, α, β and γ are preset weight coefficients. In some embodiments, the values can be 0.5, 0.3 and 0.2. d represents the distance between the faulty node and the backup node. SNR represents the signal quality of the link between the nodes. F history Indicates the historical failure frequency of the standby node.
[0077] For a star topology, the on-chip optical network is pre-configured with a master central node, corresponding shadow central nodes, and peripheral nodes in addition to the master central node. The embodiments of this disclosure primarily propose technical solutions for situations where the master central node in an on-chip optical network with a star topology fails.
[0078] In one or more embodiments of the present disclosure, for a star topology, resetting the node links of the on-chip optical network according to the self-healing strategy may include:
[0079] Build backup links between the shadow central node corresponding to the faulty node and other peripheral nodes in the faulty link.
[0080] To avoid channel conflicts between the shadow central node and the main central node, the wavelength of the microring resonator can be adjusted to seamlessly transmit data packets based on the new wavelength.
[0081] In one or more embodiments of the present disclosure, a dedicated optical channel may be reserved for high-priority services based on service priority.
[0082] by Figure 4 In the illustrated embodiment, the primary central node is (1, 2) and the corresponding shadow central node is (0, 1). When the primary central node fails, the peripheral nodes of the original link, such as (0, 2) and (1, 1), are remapped to the shadow central node to form a backup node.
[0083] For a tree topology, the backup child nodes of the parent node can be pre-stored, and the link status of the backup child nodes (such as bandwidth, delay, and packet loss rate) can be updated in real time and cached locally.
[0084] Resetting the node links of the on-chip optical network according to the self-healing strategy may include:
[0085] The current state and current load of the child node are input into a preset reinforcement learning model to obtain the optimal waiting path. In the embodiment of the present disclosure, the reinforcement learning model can be based on the Proximal Policy Optimization (PPO) algorithm.
[0086] To further enhance the technical effectiveness, the present disclosure also integrates an on-chip optical power monitoring module to monitor the optical signal strength of the reconstructed link in real time. When the optical power loss exceeds a preset threshold, a local rerouting mechanism is triggered, prioritizing redundant idle subnodes within the immediate area to fill the gap and reassessing the optical signal strength of the link.
[0087] like Figure 5 As shown, in one embodiment of the present disclosure, when the parent node (0,1) fails, the child nodes (1,1) and (2,2) are activated. In this embodiment, the optimal alternative is to activate the child node (1,1) and connect it to the child node (0,0). After verifying that the optical power loss is greater than the preset value, the node (3,2) can be selected to fill the gap.
[0088] For Mesh and Torus topologies, you can use predictive traffic engineering and self-healing strategies for elastic topology conversion. Specifically, these strategies include:
[0089] Nodes in Mesh and Torus topologies are embedded in a micro-blockchain module and periodically broadcast their status (such as load, available links, and fault flags). When a node fails, the Raft consensus algorithm can be used to elect a temporary master node near the failed node. After the temporary master node is determined, neighboring nodes submit link resources to the temporary master node.
[0090] To improve the technical effectiveness of the disclosed technical solution, a spatiotemporal graph convolutional network model can be used to predict potential congested areas. The model takes as input a node coordinate matrix, historical traffic data, and real-time link load, and outputs a traffic map that identifies potential congested areas.
[0091] Based on the traffic graph and the improved ant colony algorithm, the optimal path to bypass the fault point can be further obtained. The objective function of the improved ant colony algorithm can be set based on minimizing end-to-end delay and load imbalance. For example, the objective function can be:
[0092]
[0093] Among them, λ represents the balance coefficient, which can be set to 0.6, L represents the total number of links, L j represents the load of the jth link, Indicates the average link load, D i represents the end-to-end delay of the i-th path,
[0094] D max Indicates the maximum allowed delay, which can be set to D max =50μs.
[0095] In order to further improve the technical effect, a topology adjustment strategy based on the fault node density can be implemented in the fuzzy logic controller.
[0096] Specifically, when the node failure density of the on-chip optical network is lower than a preset threshold, the original topology is maintained; when the node failure density is greater than or equal to the preset threshold, the network can be divided into multiple sub-networks, and each sub-network can be converted into a tree topology, and sub-domains can be interconnected through the top parent node optical router.
[0097] That is to say, the input of the fuzzy logic controller is the proportion, location and business type weight of the faulty nodes, and the output is the sub-network topology generated by dynamic adjustment based on the location of the faulty nodes and the business type weight of each node. In computationally intensive scenarios, it prioritizes reducing the number of hops and improving throughput.
[0098] by Figure 5 For example, the coordinates of the failed node are (3, 2). The coordinates of the temporary master node are calculated to be (2, 2). Neighboring nodes (3, 1) and (3, 3) submit available link resources to the temporary master node. Calculation then determines that the potential congestion area is the link load around coordinates (2, 3) with a load greater than 80%. Based on this, the optimal alternative path is ultimately determined to be (2, 2) to (2, 3) to (3, 3).
[0099] In this embodiment, since the fault density is greater than a preset threshold, it is determined through calculation that the network is divided into multiple 2×2 sub-grids (such as coordinates (0,0)-(1,1)).
[0100] After determining the self-healing strategy, the self-healing operation module sends a control signal to the standby node marked during topological self-healing. Upon receiving the signal, the optical router adjusts the resonant wavelength of the microring resonator to open or close the corresponding optical link, ensuring that the standby node completes link takeover in the shortest possible time, achieving seamless handover of data packet communication functions. Furthermore, to avoid optical signal interference or transmission errors caused by parameter adjustments, the system performs multiple status verifications after the adjustment is complete to ensure that the standby node is fully capable of taking over, and then returns an ACK confirmation message to the off-chip controller.
[0101] After receiving the ACK message indicating that the backup node has successfully established a link, the off-chip controller updates the node path information stored in the global routing table to reflect the changes in the topology structure, and runs the shortest path optimization algorithm to recalculate the routes for all nodes in the network to ensure that the data flow can be smoothly transmitted according to the self-healing topology structure. The signaling layer is used to broadcast the routing table information to the relevant nodes. The broadcast adopts a hierarchical design, and prioritizes the synchronization of key nodes (central nodes and high-traffic nodes) to minimize the risk of network interruption.
[0102] After self-healing is completed, the system will verify the current network operation status and evaluate the network recovery by comparing parameters such as the average delay, throughput, bit error rate and link utilization of data packet transmission before and after self-healing. If the ratio of the same coefficients is found to be less than the preset threshold, there is an abnormality in the link. The system will dynamically return to step 4 and re-run the self-healing strategy based on the node feedback. If the ratio is greater than the threshold, there is no abnormality and the system will restart the real-time monitoring module to prepare for changes in the next node failure.
[0103] Therefore, the technical solution disclosed in the present invention adaptively adjusts the self-healing strategy according to the topology structure. Under different topologies, the node self-healing strategy will be targetedly optimized to minimize the data path reconfiguration overhead and link reconstruction time. This strategy effectively reduces the delay caused by the link switching process, while ensuring the efficiency and continuity of the overall data transmission, so as to achieve seamless activation of the replacement node.
[0104] To summarize, in the embodiments of the present disclosure, specifically, in a Ring topology, when an off-chip controller detects a node failure, the self-healing strategy system performs the following operations: adjusts the resonant wavelength of the microring resonator (MRR), constructs a clockwise redundant loop, activates adjacent backup nodes, selects the optimal replacement node based on a comprehensive scoring model of Euclidean distance, link quality, and historical failure rate, uses an optical switch matrix (OSM) to adjust the optical link connection, enables bidirectional path load balancing, and distinguishes real-time services from disaster recovery data through priority labels; in a star topology, when a central node fails, the self-healing strategy system performs the following operations: implements backup switching through a master-shadow dual-center node architecture, the shadow node takes over control within 10μs, adjusts the wavelength of the microring resonator (λ0→λ1) to avoid channel conflicts, and reserves dedicated optical channels for high-priority services; in a tree topology, when a parent node fails, the self-healing strategy system performs the following operations: calls a PPO reinforcement learning model, generates an optimal replacement path based on a child node state snapshot and real-time load, detects the optical signal strength of the reconstructed link in real time through an on-chip integrated optical power monitoring module, and triggers a local rerouting mechanism if the optical power attenuation exceeds a threshold.
[0105] In Mesh and Torus topologies, the self-healing strategy performs the following operations: it deploys a lightweight blockchain framework to periodically broadcast its own status, elects a temporary master node, and uses a spatiotemporal graph convolutional network (ST-GCN) to predict the traffic distribution in the next 10ms when a fault occurs. It then uses an improved ant colony algorithm (ACO) to generate the optimal path around the faulty node, and only updates the local routing table. When the proportion of faulty nodes exceeds a threshold, the network is divided into multiple autonomous subdomains, which are then converted to a tree topology and interconnected through the top parent node optical router. The generated subnetwork topology is dynamically adjusted based on the relevant information of the faulty node through a fuzzy logic controller, and the network structure is optimized according to the fault density and service type to reduce the number of hops.
[0106] The technical solution disclosed in the present invention monitors the health status of network nodes in real time and responds quickly to node failures. It selects a self-healing strategy based on the topology of on-chip optical network nodes, adapts to the self-healing requirements under different topologies, ensures seamless transmission of data packets, and improves the reliability and flexibility of the network.
[0107] The present disclosure is more concise and efficient in determining and handling fault types, reduces the computational burden of the system, and reduces the complexity and overhead of network reconstruction. Especially in high-speed transmission and large-scale network applications, the present disclosure can effectively avoid real-time and scalability issues and enhance the long-term stability and reliability of the network.
[0108] It can be understood that the method can be executed by any device, equipment, platform, or device cluster with computing and processing capabilities.
[0109] The technical carriers involved in the payment described in the embodiments of the present disclosure may include, for example, Near Field Communication (NFC), WIFI, 3G / 4G / 5G, POS card swiping technology, QR code scanning technology, barcode scanning technology, Bluetooth, infrared, Short Message Service (SMS), Multimedia Message Service (MMS), etc.
[0110] The biometric features involved in the biometric identification described in the embodiments of the present disclosure may include, for example, eye prints, voice prints, fingerprints, palm prints, heartbeat, pulse, chromosomes, DNA, bite marks, etc. Eye prints may include irises, sclera, and other biometric features.
[0111] One or more embodiments of the present disclosure
[0112] It should be noted that the methods of one or more embodiments of the present disclosure can be performed by a single device, such as a computer or server. The methods of these embodiments can also be applied in a distributed scenario, performed by multiple devices working together. In such a distributed scenario, one of the multiple devices may only perform one or more steps of the methods of one or more embodiments of the present disclosure, and the multiple devices will interact with each other to complete the described methods.
[0113] It should be noted that the above description is of specific embodiments of the present disclosure. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. In addition, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0114] Based on the same inventive concept, corresponding to any of the above embodiments and methods, the present disclosure also provides a self-healing system for on-chip optical network node failures. Figure 2 As shown, the above system includes: a data monitoring module 11, a fault detection module 12, a self-healing strategy determination module 13 and a self-healing strategy operation module 14;
[0115] The data monitoring module 11 is configured to obtain health status data of each node in the on-chip optical network;
[0116] The fault detection module 12 is configured to determine a faulty node of the on-chip optical network based on the health status data;
[0117] The self-healing strategy determination module 13 is configured to determine a self-healing strategy corresponding to the on-chip optical network in response to a faulty node in the optical network being unable to recover;
[0118] The self-healing strategy operation module 14 is configured to reset the node links of the on-chip optical network according to the self-healing strategy.
[0119] For the convenience of description, the above devices are described as being functionally divided into various modules. Of course, when implementing one or more embodiments of the present disclosure, the functions of each module can be implemented in the same or multiple software and / or hardware.
[0120] The apparatus of the above embodiment is used to implement the corresponding method in the above embodiment and has the beneficial effects of the corresponding method embodiment, which will not be described in detail here.
[0121] Figure 610 is a schematic diagram showing a more specific hardware structure of an electronic device provided in this embodiment. The device may include: a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. The processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040 are communicatively connected to each other within the device via the bus 1050.
[0122] The processor 1010 can be implemented using a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided by the embodiments of the present disclosure.
[0123] The memory 1020 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage devices, dynamic storage devices, etc. The memory 1020 can store an operating system and other application programs. When the technical solutions provided by the embodiments of the present disclosure are implemented through software or firmware, the relevant program codes are stored in the memory 1020 and are called and executed by the processor 1010.
[0124] The input / output interface 1030 is used to connect input / output modules to implement information input and output. The input / output modules can be configured as components within the device (not shown in the figure) or can be externally connected to the device to provide corresponding functions. Input devices may include a keyboard, mouse, touch screen, microphone, various sensors, etc., and output devices may include a display, speaker, vibrator, indicator light, etc.
[0125] The communication interface 1040 is used to connect to a communication module (not shown) to enable communication between the device and other devices. The communication module can communicate via a wired method (such as USB, network cable, etc.) or a wireless method (such as mobile network, WiFi, Bluetooth, etc.).
[0126] The bus 1050 comprises a path for transmitting information between the various components of the device (eg, the processor 1010 , the memory 1020 , the input / output interface 1030 , and the communication interface 1040 ).
[0127] It should be noted that although the above device only shows the processor 1010, the memory 1020, the input / output interface 1030, the communication interface 1040, and the bus 1050, in a specific implementation, the device may also include other components necessary for normal operation. In addition, those skilled in the art will understand that the above device may only include the components necessary to implement the embodiments of the present disclosure, and does not necessarily include all the components shown in the figure.
[0128] The electronic devices of the above embodiments are used to implement the corresponding methods in the above embodiments and have the beneficial effects of the corresponding method embodiments, which will not be described in detail here.
[0129] The computer-readable media of this embodiment include permanent and non-permanent, removable and non-removable media that can be used to store information by any method or technology. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, read-only compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device.
[0130] Those skilled in the art should understand that the discussion of any of the above embodiments is merely illustrative and is not intended to imply that the scope of the present disclosure (including the claims) is limited to these examples. Within the scope of the present disclosure, the technical features of the above embodiments or different embodiments may be combined, the steps may be implemented in any order, and there are many other variations of different aspects of one or more embodiments of the present disclosure as described above, which are not provided in detail for the sake of simplicity.
[0131] In addition, to simplify the description and discussion, and in order not to obscure one or more embodiments of the present disclosure, known power / ground connections to integrated circuit (IC) chips and other components may or may not be shown in the provided figures. In addition, devices may be shown in the form of block diagrams to avoid obscuring one or more embodiments of the present disclosure, and this also takes into account the fact that the details of the implementation of these block diagram devices are highly dependent on the platform on which one or more embodiments of the present disclosure will be implemented (i.e., these details should be fully within the scope of understanding of those skilled in the art). Where specific details (e.g., circuits) are set forth to describe exemplary embodiments of the present disclosure, it will be apparent to those skilled in the art that one or more embodiments of the present disclosure may be implemented without these specific details or with variations in these specific details. Therefore, these descriptions should be considered illustrative rather than restrictive.
[0132] Although the present disclosure has been described in conjunction with specific embodiments thereof, many alternatives, modifications, and variations of these embodiments will be apparent to those skilled in the art based on the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may use the embodiments discussed.
[0133] The one or more embodiments of the present disclosure are intended to encompass all such substitutions, modifications, and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the one or more embodiments of the present disclosure should be included within the scope of protection of the present disclosure.
Claims
1. A self-healing method for an on-chip optical network node failure, characterized in that: A self-healing system for on-chip optical network node failures includes a data monitoring module, a fault detection module, a self-healing strategy determination module, and a self-healing strategy operation module. The method includes: The data monitoring module obtains the health status data of each node in the on-chip optical network; A fault detection module determines a faulty node of the on-chip optical network based on the health status data; A self-healing strategy determination module determines a self-healing strategy corresponding to the on-chip optical network in response to a faulty node in the optical network being unable to recover; The self-healing strategy operation module resets the node links of the on-chip optical network according to the self-healing strategy.
2. The method according to claim 1, characterized in that Resetting the node links of the on-chip optical network according to the self-healing strategy includes: The standby node of the on-chip optical network receives a control instruction generated according to the self-healing strategy; The standby node is activated according to the control instruction, and a global routing table of the on-chip optical network is updated, where the global routing table is used to store routing information of all nodes of the on-chip optical network.
3. The method according to claim 2, characterized in that In response to the failure of the optical network node being unable to recover, determining a self-healing strategy corresponding to the on-chip optical network includes: In response to a failed node of the on-chip optical network being unable to recover, determining a topology of the on-chip optical network; A self-healing strategy corresponding to the on-chip optical network is generated according to the topological structure of the on-chip optical network and the node information of the on-chip optical network.
4. The method according to claim 3, characterized in that In response to the topology of the on-chip optical network being a ring topology, resetting the node links of the on-chip optical network according to the self-healing strategy includes: Adjusting parameters of microring resonators of all optical routers in the on-chip optical network to construct a redundant loop in the opposite direction of the faulty loop, serving as a physical basis for optical link adjustment; Adjust optical links through the optical switch matrix to enable bidirectional path synchronous transmission; Obtaining a proximity score of the backup node of the on-chip optical network according to the distance between the on-chip optical network fault node and the backup node of the on-chip optical network, the signal quality of the link between the on-chip optical network nodes, and the historical failure frequency of the backup node of the on-chip optical network; A target backup node of the on-chip optical network is determined according to the proximity score, and the target backup node of the on-chip optical network is used to replace the faulty node of the on-chip optical network to construct a transmission link.
5. The method according to claim 3, characterized in that In response to the topology of the on-chip optical network being a star topology, resetting the node links of the on-chip optical network according to the self-healing strategy includes: Determining a shadow central node corresponding to a faulty node of the on-chip optical network; Building a transmission link between the peripheral node of the faulty node of the on-chip optical network and the shadow central node; The microring resonator of the shadow central node is adjusted to distinguish the wavelength of the shadow central node from the wavelength of the faulty node of the on-chip optical network.
6. The method according to claim 3, characterized in that In response to the topology of the on-chip optical network being a tree topology, resetting the node links of the on-chip optical network according to the self-healing strategy includes: Inputting the subnode status of the faulty node of the on-chip optical network and the current load of the on-chip optical network into a preset reinforcement learning model to obtain an optimal alternative path corresponding to the faulty node of the on-chip optical network; In response to the optical power loss of the optimal alternative path being greater than a preset threshold, an idle sub-node within a nearby range is searched for to fill the position.
7. The method according to claim 3, characterized in that In response to the topology of the on-chip optical network being a mesh grid or a two-dimensional torus network topology, resetting the node links of the on-chip optical network according to the self-healing strategy includes: Deploy the nodes of the on-chip optical network in a micro-blockchain framework, elect a temporary master node in the vicinity of the faulty node of the on-chip optical network, and establish a temporary transmission link; Inputting the node coordinate matrix, historical traffic data, and current link load of the on-chip optical network into a preset spatiotemporal graph convolutional network to obtain a predicted traffic map, wherein the predicted traffic map shows potential congested areas of the on-chip optical network; In response to determining that the proportion of faulty nodes is less than a preset threshold, an optimal alternative path is obtained according to the predicted traffic graph, and the improved ant colony algorithm sets an objective function with the goal of minimizing end-to-end delay and load imbalance; In response to determining that the proportion of faulty nodes is greater than or equal to a preset threshold, the on-chip optical network is divided into a plurality of sub-networks and the plurality of sub-networks are converted into a tree topology structure.
8. A self-healing system for node failure in an on-chip optical network, characterized in that: include: Data monitoring module, fault detection module, self-healing strategy determination module and self-healing strategy operation module; The data monitoring module is configured to obtain health status data of each node in the on-chip optical network; The fault detection module is configured to determine a faulty node of the on-chip optical network based on the health status data; The self-healing strategy determination module is configured to determine a self-healing strategy corresponding to the on-chip optical network in response to a faulty node in the optical network being unable to recover; The self-healing strategy operation module is configured to reset the node links of the on-chip optical network according to the self-healing strategy.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.
10. A non-transitory computer-readable storage medium, characterized in that The non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 7.
Citation Information
Cited By
Satellite-ground fusion hybrid ad hoc network routing protocol and fault self-healing method and device, equipment and storage medium
CN121037883A
Optimized communication network health degree prediction and collaboration method
CN121664677A