Design method for fault-tolerant adaptive link between super-large-scale network-on-chip routers

By designing multi-drive strength physical interfaces, redundant interconnection lines, adaptive adjustment mechanisms and connection detection and repair technologies among super-large-scale on-chip network routers, the problems of signals being susceptible to interference, connection instability and environmental factors in the prior art are solved, and high reliability and stable interconnection performance are achieved.

CN120075167APending Publication Date: 2025-05-30GUANGDONG INST OF INTELLIGENT SCI & TECH
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510240179.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The prior art is susceptible to electromagnetic interference and noise in the interconnection between ultra-large-scale on-chip network routers, resulting in signal distortion; it is difficult to ensure a stable connection between the new access chip and the original system when the system is expanded; and it is susceptible to external environmental factors, resulting in a degradation of interconnection performance.

Method used

The multi-drive strength physical interface design, redundant interconnection line design, adaptive adjustment mechanism and connection detection and repair technology are adopted to ensure the reliability and stability of signal transmission by dynamically adjusting signal strength, automatically switching redundant lines, real-time monitoring and adjustment of interconnection parameters, periodic detection of link status and automatic repair programs.

Benefits of technology

It improves the reliability of the interconnection of scalable chip systems, reduces the incidence of communication failures caused by various factors, ensures stable connection between the new access chip and the original system, and maintains a good working condition in different environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120075167A_ABST
    Figure CN120075167A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of chip system interconnection, in particular to a method for designing a fault-tolerant self-adaptive link between super-large-scale on-chip network routers, which comprises the following steps of: configuring a plurality of multi-drive strength physical interfaces on a physical layer, and dynamically adjusting signal strength according to environmental requirements so as to balance transmission quality and power consumption; a main interconnection line and at least one redundant interconnection line are arranged for each chip node in a data layer, and when the main interconnection line breaks down or the signal quality is reduced, the main interconnection line is automatically switched to the redundant interconnection line for data transmission; interconnection parameters are dynamically adjusted based on the signal quality, the environmental parameters and the chip load condition monitored in real time, and the interconnection parameters comprise the transmission rate and the signal strength; and the link state is diagnosed on the protocol layer through the periodic detection signal, and an automatic repair program is triggered when abnormity or fault occurs. According to the method, the interconnection stability and the data transmission accuracy of the chip system in various working environments can be improved, and the system can be ensured to be reliably expanded and operated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of chip system interconnection, and particularly to a fault-tolerant and adaptive link design method for ultra-large-scale network-on-chip routers. Background Art

[0002] Ultra-large-scale network-on-chip routers are important components of network-on-chip (NoC), and are used to achieve efficient communication in multi-core chips or system-on-chip (SoC). With the development of integrated circuit technology, the number of processing cores integrated on a chip is increasing continuously. The traditional bus architecture can no longer meet the high-bandwidth and low-latency communication requirements in multi-core systems. Therefore, network-on-chip routers, as a new type of communication architecture, have emerged as the times require.

[0003] In scalable chip systems, with the expansion of the chip system scale and the improvement of complexity, interconnection reliability has become a key challenge. The existing technologies have the following problems: 1. Traditional interconnection methods are vulnerable to electromagnetic interference and noise, resulting in signal distortion, and further affecting the accuracy and reliability of data transmission; 2. It is difficult to ensure a stable connection between newly added chips and the original system during system expansion, and problems such as unstable connection and communication interruption are likely to occur; 3. It is vulnerable to external environmental factors (such as temperature changes, mechanical vibrations, etc.), resulting in a decline in interconnection performance, and further seriously affecting the overall performance and stability of the scalable chip system.

[0004] Therefore, there is an urgent need for a new technical solution to solve the above technical problems. Summary of the Invention

[0005] The purpose of the present invention is to overcome the problems of the above existing technologies, and provides a fault-tolerant and adaptive link design method for ultra-large-scale network-on-chip routers, so as to solve the technical problems that the existing technologies are vulnerable to electromagnetic interference and noise, resulting in signal distortion; it is difficult to ensure a stable connection between newly added chips (nodes) and the original system during system expansion, and the interconnection performance declines due to environmental factors (such as temperature changes, mechanical vibrations, etc.).

[0006] The above purpose is achieved through the following technical solutions: A fault-tolerant and adaptive link design method for ultra-large-scale network-on-chip routers, including: Multi-drive strength physical interface design, configuring multiple multi-drive strength physical interfaces at the physical layer, and dynamically adjusting the signal strength according to environmental requirements to balance transmission quality and power consumption; Redundant interconnection line design, where a main interconnection line and at least one redundant interconnection line are set for each chip node at the data layer. When the main interconnection line fails or the signal quality deteriorates, it automatically switches to the redundant interconnection line serving as a backup for data transmission; Adaptive adjustment mechanism design, which dynamically adjusts interconnection parameters based on real-time monitored signal quality, environmental parameters, and chip load conditions. The interconnection parameters include transmission rate and signal strength; Connection detection and repair technology design, which diagnoses the link state by periodically detecting signals at the protocol layer. When an anomaly or fault occurs, it triggers an automatic repair program, and the automatic repair program includes re-initializing the connection or switching to a redundant interconnection line.

[0007] Furthermore, in the multi-drive strength physical interface design, the multi-drive strength physical interface supports multiple drive modes, including a high drive mode for anti-interference environments and a low drive mode for low-power scenarios.

[0008] Furthermore, in the redundant interconnection line design, the main interconnection line and the redundant interconnection line adopt heterogeneous physical interfaces. Among them, the main interconnection line is in a high-speed and low anti-interference mode, and the redundant interconnection line is in a low-speed and high anti-interference mode.

[0009] Furthermore, the adaptive adjustment mechanism design further includes: when the signal performance deteriorates due to an increase in temperature, reducing the transmission rate and enhancing the signal strength.

[0010] Furthermore, the adaptive adjustment mechanism design also further includes: when electromagnetic interference is detected, switching to an anti-interference mode and increasing the redundant coding ratio to 30% - 50%.

[0011] Furthermore, in the connection detection and repair technology design, the diagnosis of the link state by periodically detecting signals specifically means sending detection signals to the interconnection nodes at preset time intervals and judging the connection state based on the signal delay or packet loss rate.

[0012] Furthermore, the hardware implementation of the method includes: integrating redundant interconnection interfaces and signal processing modules inside the chip.

[0013] Furthermore, the hardware implementation of the method also includes: setting up a dedicated connection detection hardware circuit to monitor and report connection anomalies in real time.

[0014] Furthermore, the software implementation of the method includes: an interconnection parameter adaptive adjustment software algorithm that generates parameter adjustment instructions based on monitoring data.

[0015] Furthermore, in software implementation, the method further includes: executing a connection detection and repair software program for fault diagnosis and repair operations, and recording fault information for system optimization.

[0016] A fault-tolerant adaptive link design method for routers between very large-scale on-chip networks provided by the present invention, through the comprehensive application of redundant interconnection line design, signal enhancement and anti-interference technology, adaptive adjustment mechanism, and connection detection and repair technology, not only improves the reliability of the interconnection of scalable chip systems, but also effectively reduces the incidence of communication failures caused by various factors. When the chip system is expanded, the interconnection method of the present invention can ensure reliable connection between newly accessed chips and the original system, and does not affect the stability and performance of the entire system. Also, through the adaptive adjustment mechanism, the interconnection system can maintain a good working state in different working environments (such as environments with large temperature and humidity changes), further improving the environmental adaptability of the system. Brief Description of the Drawings

[0017] Figure 1 It is a schematic diagram of the interconnection of chips in a fault-tolerant adaptive link design method for routers between very large-scale on-chip networks according to the present invention; Figure 2 It is a schematic diagram of the hierarchical structures within a chip in a fault-tolerant adaptive link design method for routers between very large-scale on-chip networks according to the present invention; Figure 3 It is a schematic diagram of the data transmission path between chips in a fault-tolerant adaptive link design method for routers between very large-scale on-chip networks according to the present invention. Detailed Embodiments

[0018] The present invention will be further described in detail below with reference to the drawings and embodiments. The described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.

[0019] As Figures 1 - 3 shown, the present solution provides a fault-tolerant adaptive link design method for routers between very large-scale on-chip networks, including: Multi-drive strength physical interface design, configuring multiple multi-drive strength physical interfaces at the physical layer, and dynamically adjusting the signal strength according to environmental requirements to balance transmission quality and power consumption; used to adapt to different requirements for signal quality in different environments, and reduce transmission power consumption as much as possible while ensuring signal transmission; Redundant interconnection line design, where a primary interconnection line and at least one redundant interconnection line are set for each chip node at the data layer. When the primary interconnection line fails or the signal quality deteriorates, it automatically switches to the redundant interconnection line serving as a backup for data transmission; by designing redundant interconnection lines at the data layer, the reliability of the interconnection can be greatly improved, reducing the risk of communication interruption caused by line failures. Even if one of the data paths is damaged, it will not affect data transmission; Adaptive adjustment mechanism design, based on real-time monitored signal quality, environmental parameters (such as temperature, humidity, etc.) and chip load conditions, dynamically adjusts interconnection parameters, and the interconnection parameters include transmission rate and signal strength, etc.; for example, when the system detects that the increase in temperature causes a decline in signal transmission performance, it automatically reduces the transmission rate and simultaneously enhances the signal strength to ensure reliable data transmission; Connection detection and repair technology design, periodically detects signals at the protocol layer to diagnose the link status, and triggers an automatic repair program when abnormal or faulty. The automatic repair program includes re-initializing the connection or switching to a redundant interconnection line; the automatic repair program first diagnoses the abnormality or fault to determine the cause of the fault, and then takes corresponding repair measures, such as re-initializing the connection, switching to a redundant interconnection line, etc. For example, at regular time intervals, the system sends detection signals to each interconnected chip node and judges whether the connection is normal according to the returned signal status.

[0020] As Figure 2 shown, the specific structures of each layer within the chip are as follows: At the physical layer, multiple driving modes and corresponding physical interfaces are adopted to adapt to different application scenarios.

[0021] At the data layer, a redundant data path design is adopted, and the damage of one data path will not affect data transmission.

[0022] At the protocol layer, multiple fault-tolerant algorithms are supported to achieve data fault tolerance and error correction.

[0023] In this embodiment of the multi-driving strength physical interface design, the multi-driving strength physical interface supports multiple driving modes, including a high driving mode for anti-interference environments and a low driving mode for low-power consumption scenarios. This embodiment adopts multiple driving modes and corresponding physical interfaces at the physical layer to adapt to different application scenarios.

[0024] In this embodiment of the redundant interconnection line design, the primary interconnection line and the redundant interconnection line adopt heterogeneous physical interfaces. Among them, the primary interconnection line is in a high-speed and low anti-interference mode, and the redundant interconnection line is in a low-speed and high anti-interference mode.

[0025] The adaptive adjustment mechanism design in this embodiment further includes: when the signal performance degrades due to an increase in temperature (where the change in temperature can be achieved through the input of temperature sensor data), reducing the transmission rate and enhancing the signal strength.

[0026] The adaptive adjustment mechanism design also further includes: when electromagnetic interference is detected (where the electromagnetic interference can be achieved through the monitoring of the interference monitoring module), switching to the anti-interference mode and increasing the redundancy coding ratio to 30% - 50%.

[0027] In the connection detection and repair technology design in this embodiment, the specific method of diagnosing the link state by periodically detecting signals is to send detection signals to the interconnected nodes at preset time intervals (such as 10 ms), and judge the connection state according to the signal delay or packet loss rate.

[0028] This embodiment also provides a hardware implementation of a fault-tolerant adaptive link design method for routers in a very large scale on-chip network, including: integrating redundant interconnection interfaces and signal processing modules inside the chip; that is, during the chip design stage, adding redundant interconnection interfaces for redundant interconnection lines to each chip to ensure convenient connection of redundant interconnection lines; at the same time, integrating a signal processing module inside the chip to implement signal enhancement and anti-interference functions.

[0029] This method in the hardware implementation also includes: setting up a dedicated connection detection hardware circuit to monitor and report connection anomalies in real time. In this embodiment, a dedicated connection detection hardware circuit is designed to monitor the interconnection status in real time. This connection detection hardware circuit can quickly and accurately detect connection anomalies and send an alarm signal to the system in a timely manner.

[0030] This embodiment also provides a software implementation of a fault-tolerant adaptive link design method for routers in a very large scale on-chip network, including: an interconnection parameter adaptive adjustment software algorithm for generating parameter adjustment instructions based on monitoring data. This algorithm generates corresponding parameter adjustment instructions through complex calculations and logical judgments according to data such as signal quality and environmental parameters monitored by the hardware, to achieve automatic adjustment of interconnection parameters.

[0031] This method in the software implementation also includes: a connection detection and repair software program for performing fault diagnosis and repair operations, and recording fault information for system optimization. This software program is responsible for receiving the alarm signal sent by the hardware circuit, diagnosing and analyzing the fault, and performing corresponding repair operations. At the same time, this software program also records the fault information for subsequent system maintenance and optimization.

[0032] The above is only to illustrate the embodiments of the present invention and is not intended to limit the present invention. For those skilled in the art, any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for designing fault-tolerant adaptive links between routers in a very large-scale network on chip, characterized by: include: Multiple drive strength physical interface design: multiple multiple drive strength physical interfaces are configured at the physical layer to dynamically adjust signal strength according to environmental requirements to balance transmission quality and power consumption; Redundant interconnection line design: a main interconnection line and at least one redundant interconnection line are set for each chip node at the data layer. When the main interconnection line fails or the signal quality decreases, the redundant interconnection line as a backup is automatically switched to transmit data; Adaptive adjustment mechanism design, based on real-time monitoring of signal quality, environmental parameters and chip load, dynamically adjusts interconnection parameters, including transmission rate and signal strength; The connection detection and repair technology is designed to diagnose the link status through periodic detection signals at the protocol layer, and trigger an automatic repair program when an abnormality or failure occurs. The automatic repair program includes reinitializing the connection or switching redundant interconnection lines.

2. The method for designing fault-tolerant adaptive links between routers in a very large-scale network on chip according to claim 1, characterized in that: In the multi-drive strength physical interface design, the multi-drive strength physical interface supports multiple drive modes, including a high drive mode for an anti-interference environment, and a low drive mode for a low power consumption scenario.

3. The method for designing fault-tolerant adaptive links between routers in a very large-scale network on chip according to claim 1, characterized in that: In the redundant interconnection line design, the main interconnection line and the redundant interconnection line adopt heterogeneous physical interfaces, wherein the main interconnection line is in a high-speed and low-interference mode, and the redundant interconnection line is in a low-speed and high-interference mode.

4. The method for designing fault-tolerant adaptive links between routers in a very large-scale network on chip according to claim 1, characterized in that: The adaptive adjustment mechanism design further includes: when the temperature rises and causes the signal performance to decrease, the transmission rate is reduced and the signal strength is enhanced.

5. A method for designing fault-tolerant adaptive links between routers in a very large-scale network on chip according to claim 1 or 4, characterized in that: The adaptive adjustment mechanism design further includes: when electromagnetic interference is detected, switching to anti-interference mode and increasing the redundant coding ratio to 30%-50%.

6. The method for designing fault-tolerant adaptive links between routers in a very large-scale network on chip according to claim 1, characterized in that: The link status diagnosis by periodic detection signal in the connection detection and repair technology design is specifically to send a detection signal to the interconnected nodes at every preset time interval and determine the connection status according to the signal delay or packet loss rate.

7. The method for designing fault-tolerant adaptive links between routers in a very large-scale network on chip according to claim 1, characterized in that: The method comprises, in hardware implementation, integrating a redundant interconnection interface and a signal processing module inside a chip.

8. The method for designing fault-tolerant adaptive links between routers in a very large-scale network on chip according to claim 7, characterized in that: The method further comprises, in hardware implementation, setting a dedicated connection detection hardware circuit to monitor and report connection anomalies in real time.

9. The method for designing fault-tolerant adaptive links between routers in a very large-scale network on chip according to claim 1, characterized in that: The method comprises, in software implementation, an interconnection parameter adaptive adjustment software algorithm that generates parameter adjustment instructions based on monitoring data.

10. The method for designing fault-tolerant adaptive links between routers in a very large-scale network on chip according to claim 1, characterized in that: The method also includes, in software implementation, a connection detection and repair software program that performs fault diagnosis and repair operations, and records fault information for system optimization.

Citation Information

Cited By

  • Core particle interconnection chip, core particle interconnection method and device

    CN120762975A

  • Clock synchronization method for chip computing network

    CN121508726A

  • A clock synchronization method for a chip computing network

    CN121508726B