Intelligent test method and system for high-speed network communication port
By integrating a collaborative physical layer testing engine within the processing unit and utilizing bidirectional pseudo-random bit sequence and eye diagram feature extraction techniques, the high cost and insufficient accuracy of traditional testing methods are solved. This enables efficient and accurate end-to-end link performance evaluation and adaptive control, improving the operational efficiency and reliability of the processing unit's direct-connect architecture.
Patent Information
- Application Number
- CN202610120249.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-28
- Publication Date
- 2026-05-15
AI Technical Summary
In a high-throughput, low-latency processing unit direct-connect architecture, traditional testing methods rely on external equipment, which is costly and cannot accurately assess end-to-end link performance. They are also difficult to integrate into production self-inspection or online operation and maintenance processes and cannot distinguish whether the fault originates from the transmitter, receiver, or cable physical layer.
By integrating a collaborative physical layer test engine within the processing unit, and utilizing bidirectional pseudo-random bit sequence testing and receiver eye diagram feature extraction techniques, fault location and link health assessment are performed. By combining bit error rate and physical layer characteristic parameters, intelligent assessment and adaptive control of the end-to-end link are achieved.
It enables efficient and low-cost end-to-end link performance evaluation without the need for external devices, accurately diagnoses fault sources and cable degradation, improves operation and maintenance efficiency and reliability, and supports early warning and adaptive control.
Smart Images

Figure CN122053434A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of communication technology, specifically relating to a method and system for intelligent testing of high-speed network communication ports. Background Technology
[0002] In current high-throughput, low-latency applications such as artificial intelligence training, high-performance computing, and large-scale distributed inference, multiple processing units need to work closely together through high-speed interconnection networks. Traditional data centers typically use a tree-like network architecture, and communication between servers needs to be relayed through multiple levels of switches. This approach introduces additional electronic processing steps, resulting in a significant increase in end-to-end latency, making it difficult to meet the stringent requirements of extremely high data throughput and extremely low response between processing units.
[0003] To overcome the aforementioned limitations, a direct-connect architecture for processing units is adopted, which directly connects the serializer / deserializer ports of the processing units via high-speed cables, thereby significantly increasing bandwidth and reducing transmission latency. This direct-connection method also has significant advantages in terms of power consumption and cost, providing a scalable physical connection foundation for large-scale clusters. However, this architecture also brings new testing challenges, as the traditional link status monitoring mechanism that relies on switch feedback becomes ineffective due to the elimination of intermediate switching devices.
[0004] Currently, testing of high-speed direct-connect cables still mainly relies on external dedicated testing instruments. These devices are expensive, complex to operate, and difficult to integrate into production self-inspection or online operation and maintenance processes. Although some processing units have built-in loopback testing functions, they can only achieve local verification and cannot truly reflect end-to-end link performance, nor can they accurately distinguish the specific sources of failure such as the transmitter, receiver, or cable physical layer. This results in a lack of accuracy and diagnosability in link reliability assessment.
[0005] Therefore, there is an urgent need for a test solution that does not require external instruments, can be embedded in the processing unit's workflow, and can achieve intelligent evaluation of end-to-end physical links. Summary of the Invention
[0006] To address the aforementioned technical problems, this invention provides a method and system for intelligent testing of high-speed network communication ports.
[0007] In a first aspect, a method for intelligent testing of high-speed network communication ports is applied to a first processing unit and a second processing unit directly interconnected by a high-speed cable, the method comprising: A bidirectional pseudo-random bit sequence test is performed between the first processing unit and the second processing unit to obtain the bit error rate in the first direction and the second direction, respectively. Fault location is based on bidirectional bit error rate, which is used to distinguish whether the fault originates from the transmitting circuit, the receiving circuit, or the physical medium of the high-speed cable. At the receiving end of the first processing unit and / or the second processing unit, eye diagram sampling is performed on the received test signal to extract physical layer feature parameters including at least eye height, eye width, jitter and signal-to-noise ratio; The physical medium health status of high-speed cables is evaluated based on the comparison results of physical layer characteristic parameters and preset thresholds. Based on the fault location results of the combined bidirectional bit error rate and the health status of the physical medium, a health level assessment is performed on the end-to-end link composed of high-speed cables. Based on the health level assessment results, trigger the corresponding link parameter adaptive adjustment strategy or link recovery strategy.
[0008] Preferably, performing a bidirectional pseudo-random bit sequence test includes: A first test bit stream is generated by a first pseudo-random sequence generator located at the transmitting end of the first processing unit and transmitted to the receiving end of the second processing unit via a high-speed cable. At the receiving end of the second processing unit, the sampling clock is locked and restored using a clock data recovery circuit, which drives the local first reference sequence generator to generate a reference sequence synchronized with the first test bit stream. By comparing the received bit stream with the reference sequence, the bit error rate in the first direction is calculated. A second test bit stream is generated by a second pseudo-random sequence generator located at the transmitting end of the second processing unit, and then transmitted in reverse to the receiving end of the first processing unit via a high-speed cable. At the receiving end of the first processing unit, synchronous clock recovery and sequence comparison are performed to calculate the bit error rate in the second direction.
[0009] Preferably, before starting the bidirectional pseudo-random bit sequence test, the method further includes: A handshake synchronization is performed between the first processing unit and the second processing unit through a dedicated out-of-band management channel to negotiate and confirm the initial seed values used for the first pseudo-random sequence generator and the first reference sequence generator, ensuring that the starting phase of the test sequence is consistent.
[0010] Preferably, fault location based on bidirectional bit error rate specifically includes: If the bit error rate in the first direction is higher than the first warning threshold and the bit error rate in the second direction is lower than the second health threshold, then the fault is determined to be located in the transmitting circuit of the first processing unit or the connection part of the high-speed cable near the first processing unit. If the bit error rate in the second direction is higher than the first warning threshold and the bit error rate in the first direction is lower than the second health threshold, then the fault is determined to be located in the transmitting circuit of the second processing unit or the connection part of the high-speed cable near the second processing unit. If the bit error rate in both the first and second directions is higher than the first warning threshold, it is determined that the fault may be located in the middle section of the high-speed cable or that the connectors at both ends are degraded.
[0011] Preferably, eye diagram sampling is performed on the received test signal to extract physical layer feature parameters, including: Within a unit symbol interval, multi-phase sampling in the horizontal direction is controlled by a delay-locked loop, and multiple decision thresholds in the vertical direction are set by a programmable comparator array. Accumulate sampling points over multiple consecutive signal periods to construct an eye diagram statistical matrix; Based on the eye diagram statistical matrix, calculate and output eye height, eye width, jitter peak-to-peak value, and signal-to-noise ratio.
[0012] Preferably, the health level assessment of the end-to-end link, which combines the fault location results of the bidirectional bit error rate with the physical medium health status, includes: The link health status is divided into multiple levels, and the judgment rules include at least the following: when the bit error rate in both directions is below the excellent threshold and the physical medium health status is normal, it is evaluated as a healthy level; when the bit error rate in either direction is within the warning range or the physical medium health status indicates deterioration, it is evaluated as a warning level; when the bit error rate in either direction is above the fault threshold, it is evaluated as a fault level.
[0013] Preferably, the method also includes trend analysis and early warning steps: Maintain a historical database of link status, recording the bit error rate and physical layer characteristic parameters of each test; Trend analysis is performed on physical layer characteristic parameters. When a specified parameter is detected to show a statistically significant monotonic degradation trend in multiple consecutive tests, an early warning event is proactively triggered even if the current health level has not reached the warning level, and the health level of the link is adjusted to the warning level.
[0014] Preferably, the corresponding adaptive adjustment strategy for link parameters triggered based on the health level assessment results includes: When the health level assessment result is a warning level, the adaptive control process is initiated. The process includes: based on the fault location result, adjusting the pre-emphasis or deemphasis coefficient of the transmitter driver of the suspected problem terminal; and / or, triggering the receiver equalizer to perform adaptive reconfiguration to compensate for channel distortion. After each parameter adjustment, the steps from bidirectional pseudo-random bit sequence testing to health level assessment are re-executed until the link performance recovers to the target health level or the adjustment attempt limit is reached.
[0015] Preferably, based on the health level assessment results, the corresponding link recovery strategy triggered includes: When the health level assessment result is a fault level, the physical layer link retraining process is triggered to recalibrate the clock data recovery circuit and optimize the equalizer coefficients. If a preset redundant link exists, the service data stream will be automatically switched to the redundant link, and an alarm message containing fault location and media degradation details will be generated and reported.
[0016] Secondly, a smart testing system for high-speed network communication ports is provided, applied to a first processing unit and a second processing unit directly interconnected by a high-speed cable, the system comprising: The bidirectional PRBS collaborative generation module is used to synchronously generate a pseudo-random bit sequence test stream at the transmitting ends of the first and second processing units and send it to the other end via a high-speed cable. The bidirectional bit error rate detection module is used at the receiving end of the processing unit to capture the test stream sent by the other end, compare it with the locally synchronously generated reference sequence, and calculate and output the bit error rate in both transmission directions in real time. The eye diagram feature extraction module is integrated into the receiving end of the processing unit. It is used to perform high-precision eye diagram sampling on the input test signal and extract physical layer feature parameters including at least eye height, eye width, jitter and signal-to-noise ratio. The link health assessment module is used to receive bit error rate and physical layer characteristic parameters, locate faults based on bit error rate, assess the medium status based on physical layer characteristic parameters, and combine the results of both to assess the health level of the end-to-end link. The adaptive control module is used to dynamically adjust the SerDes physical layer parameters, trigger link retraining, or perform service channel switching operations based on the health level assessment results.
[0017] In summary, this application includes at least one of the following beneficial technical effects: 1. This invention integrates a collaborative physical layer test engine within the xPU and utilizes a bidirectional pseudo-random bit sequence (PRBS) interaction mechanism and receiver eye diagram feature extraction technology to achieve complete end-to-end performance evaluation of high-speed direct-connect links without relying on expensive external dedicated test equipment. This method not only reduces testing costs and complexity but can also be seamlessly embedded into production self-inspection, online verification, and online operation and maintenance processes, achieving in-situ and automated testing capabilities.
[0018] 2. This invention, through bidirectional bit error rate comparison, can effectively distinguish whether the fault originates from the local transmitting circuit, the remote receiving circuit, or the physical medium of the cable (such as connector degradation or cable bending), thereby achieving a diagnostic upgrade from "link connectivity" to "component-level problem localization." At the same time, by combining real-time monitoring and historical trend analysis of eye diagram characteristic parameters (such as eye height, eye width, jitter, and signal-to-noise ratio), the system can identify the progressive degradation trend of the link in advance and trigger early warnings before performance deteriorates significantly or hard faults occur, providing key data support for preventive maintenance.
[0019] 3. The system of this invention integrates the diagnostic results of bidirectional bit error rate and physical layer characteristic parameters to intelligently grade and evaluate the health status of the link (e.g., excellent, good, warning, fault). Based on the evaluation level, corresponding adaptive control strategies can be automatically triggered: for the warning state, the pre-emphasis coefficient at the transmitter or the equalizer parameters at the receiver are dynamically adjusted to attempt to restore the link performance to a healthy range; for the fault state, link retraining or switching to a redundant path can be automatically triggered, and a detailed diagnostic report is reported, forming a complete intelligent operation and maintenance closed loop, improving the reliability, availability, and operation and maintenance efficiency of large-scale high-speed interconnection networks. Attached Figure Description
[0020] Figure 1 This is a flowchart illustrating an intelligent testing method for high-speed network communication ports according to the present invention. Figure 2 This is a schematic diagram of the process of bidirectional PRBS collaborative testing and joint diagnosis of eye map features in this invention; Figure 3 This is a schematic diagram of the link health assessment and adaptive control process in this invention; Detailed Implementation
[0021] In applications such as artificial intelligence training, high-performance computing, and large-scale distributed inference, multiple xPUs work closely together through high-speed interconnect networks. Traditional data centers generally adopt a Spine-Leaf tree network architecture, in which data flow between servers must be relayed through a top-of-rack switch. This architecture introduces additional electronic processing steps, resulting in increased end-to-end latency, typically reaching the microsecond level. At the same time, the link bandwidth is limited by the switching chip's capabilities, making it difficult to meet the TB-level data throughput requirements between xPUs.
[0022] To overcome the above problems, an xPU direct connection architecture is adopted, which directly connects the SerDes ports of multiple xPUs through passive high-speed cables that conform to IEEE 802.3ck, InfiniBand NDR / HDR or OIF-CEI-112G standards. This achieves ultra-high bandwidth of 800Gbps in one direction and 1.6Tbps in two directions, as well as extremely low latency of 4.5 to 5 nanoseconds per meter. The power consumption of such passive cables is less than 0.1 watts, which is far superior to the 10 to 15 watts of traditional optical modules, significantly reducing the cost of heat dissipation and power infrastructure.
[0023] However, the xPU direct-connect architecture also brings new testing challenges. Traditional networks rely on link status feedback provided by top-of-rack switches for health monitoring, but the direct-connect architecture loses this capability due to the removal of intermediate devices. Currently, the verification of 800Gbps direct-connect cables still relies on expensive external bit error rate testers, which are costly, complex to operate, and cannot be embedded in production self-inspection or online maintenance processes. In addition, even though the xPU's internal SerDes has pseudo-random bit sequence functionality, existing solutions can only perform loopback tests, which cannot truly reflect end-to-end link performance, especially in distinguishing between local transmission problems and remote reception problems, and also cannot detect physical layer degradation such as cable aging and plug-in wear.
[0024] To address the aforementioned technical problems, this invention provides an intelligent testing method for high-speed network communication ports. Its core lies in integrating a collaborative physical layer testing engine within the xPU and utilizing a bidirectional pseudo-random bit sequence interaction mechanism and receiver eye diagram feature extraction technology to achieve joint diagnosis of the transmission performance, reception performance, and physical medium integrity of end-to-end high-speed interconnect links without relying on external testing equipment.
[0025] To further illustrate the technical means and effects adopted by the present invention in order to achieve the intended purpose, the following detailed description is provided in conjunction with the accompanying drawings and preferred embodiments, based on specific implementation methods of the present invention.
[0026] Example 1 A method for intelligent testing of high-speed network communication ports includes the following steps: S1, a pseudo-random bit stream conforming to a preset polynomial sequence is generated at the SerDes transmitter of the first xPU, and transmitted to the SerDes receiver of the second xPU via a passive high-speed cable. The specific implementation includes the following steps: S101: Start the pseudo-random bit sequence generator located at the first xPU SerDes transmitter.
[0027] This generator uses the PRBS31 sequence, and its generator polynomial is: The sequence length is .
[0028] S102: Set the sequence generation rate to be consistent with the nominal rate of the SerDes port, i.e., 800Gbps in one direction.
[0029] The generation process is controlled by the physical layer hard core inside the xPU, and the clock source comes from the master reference clock of the high-speed serial interface, with a frequency of 100MHz.
[0030] The reference clock is multiplied to a frequency that matches the bit rate of each channel by an on-chip phase-locked loop, and then distributed to each parallel channel of the PRBS logic unit via a clock tree to ensure timing synchronization between multiple channels.
[0031] S103: Perform pre-emphasis processing on the generated bitstream.
[0032] The transmitter driver initially uses a default deemphasis factor of 6dB to compensate for high-frequency attenuation of the signal during transmission.
[0033] The deemphasis factor can be dynamically adjusted in 1dB steps within the range of 0-12dB via the SerDes configuration register to adapt to different cable or channel conditions.
[0034] S104: Modulates the processed bitstream into a differential signal and outputs it through the positive and negative channels of a passive high-speed cable.
[0035] The cables used are copper cables conforming to the IEEE 802.3ck standard, with a characteristic impedance of 100Ω, an insertion loss of no more than 20dB at a frequency of 28GHz, and a return loss better than 15dB.
[0036] These cable parameters are used for the initial optimization of the pre-emphasis coefficient during the system design phase. If the cable type is changed, the pre-emphasis preset value can be updated through the management interface.
[0037] S105: During the output process, if clock lockout, drive abnormality, or signal amplitude abnormality is detected, the PRBS generator will automatically pause the output and send a test interrupt signal to the other end through the out-of-band management channel. At the same time, the error event will be recorded at this end for subsequent diagnosis.
[0038] S2, the pseudo-random bit stream is synchronously captured at the SerDes receiver of the second xPU, and bit-by-bit comparison is performed based on the locally generated identical polynomial sequence to calculate the bit error rate in the first direction. The specific implementation includes the following steps: S201: The SerDes receiver of the second xPU continuously monitors the amplitude of the input signal. When the differential signal amplitude is detected to exceed 200mV and last for at least 10 symbol periods, it is determined that a valid pseudo-random bit stream has arrived, and the synchronization acquisition process is then initiated.
[0039] The clock data recovery circuit at the receiver is based on a second-order phase-locked loop with a loop bandwidth set to 1 / 1000 of the symbol rate, which can complete phase locking in about 1μs.
[0040] After locking, the CDR outputs a sampling clock that is strictly aligned with the edge of the bit stream at the transmitting end and generates a lock status flag.
[0041] S202: Use this sampling clock to drive the PRBS31 sequence generator inside the receiver.
[0042] To ensure the sequence starts with the same phase, both ends perform a handshake synchronization via an out-of-band management channel before testing: The first xPU sends a "sequence synchronization request" containing the initial seed value of PRBS31; Upon receiving the second xPU, it initializes the local generator to the seed value and replies "Synchronization ready".
[0043] The out-of-band management channel uses Manchester encoding, has a rate of 100Mbps, and a timeout of 500ms. If synchronization is not completed within the timeout period, the test will be terminated and the error will be recorded.
[0044] S203: Perform a bitwise XOR operation on the captured input bit stream and the locally generated reference sequence in the SerDes physical layer hard core. Bits with an XOR result of "1" are identified as error bits.
[0045] S204: Error bits are accumulated in real time by a dedicated hardware counter with a 32-bit width, automatically reset every second. If the counter overflows within the statistical period, it is considered a serious error, and the bit error rate is recorded as follows: Under normal circumstances, the accumulated value is output in the form of errors per million bits, and the bit error rate in the first direction is obtained by logarithmic transformation. For example, if in the cumulative check... If one error is found in one bit, the bit error rate is... (Or denoted as 1E-9).
[0046] S3, the same pseudo-random bit stream is generated at the SerDes transmitter of the second xPU, and transmitted in reverse to the SerDes receiver of the first xPU via the passive high-speed cable. The specific implementation includes the following steps: S301: At the SerDes transmitter of the second xPU, start its integrated PRBS31 sequence generator.
[0047] The generator uses the same generating polynomial as the first xPU, which is... This is to produce the exact same pseudo-random bitstream.
[0048] S302: Configure the generation rate of this sequence to 800Gbps, which is consistent with the nominal operating rate of the SerDes port. Its clock synchronization mechanism is the same as the principle described in S1, ensuring the accuracy of the sequence bit rate.
[0049] The timing of the reverse transmission is coordinated by the out-of-band management channel: the first xPU immediately notifies the second xPU after initiating its own transmission, and the second xPU initiates reverse transmission within 1μs after receiving the notification, thereby achieving time overlap and phase management for bidirectional testing.
[0050] S303: Apply pre-emphasis processing to the generated bitstream.
[0051] The transmitter driver uses the same deemphasis factor as the first xPU, i.e., 6dB, by default to maintain the consistency of bidirectional test conditions.
[0052] If the difference in bidirectional bit error rate exceeds 3dB in subsequent analysis, the system allows independent adjustment of the pre-emphasis coefficient of the reverse channel, with an adjustment range of 0-12dB and a step of 1dB, to adapt to channel asymmetry.
[0053] S304: Modulate the processed bitstream into a differential signal and send it to the SerDes receiver of the first xPU through the "reverse channel" of the same passive high-speed cable currently in use.
[0054] Here, "reverse channel" refers to the independent differential pair in the cable used for reverse transmission, which shares the cable's physical resources with the forward channel.
[0055] To improve diagnostic accuracy, the system can perform channel calibration during the initialization phase, measuring and recording the performance differences of the bidirectional baselines for subsequent threshold compensation.
[0056] S4, the reverse pseudo-random bit stream is synchronously captured at the SerDes receiver of the first xPU, and bit-by-bit comparison is performed to calculate the bit error rate in the second direction. The specific implementation includes the following steps: S401: When the signal monitoring circuit of the SerDes receiver of the first xPU detects that the amplitude of the input differential signal exceeds 200mV and lasts for at least 10 symbol periods, it determines that the reverse pseudo-random bit stream from the second xPU has arrived stably and then starts the synchronization acquisition process.
[0057] The clock data recovery circuit at the receiving end works first, locking the phase of the inverted input signal and extracting a sampling clock that is strictly aligned with the second xPU transmitter.
[0058] S402: Use this recovered sampling clock to drive the PRBS31 sequence generator inside the first xPU receiver.
[0059] This generator uses the same generating polynomial ( The sequence start phase is guaranteed by a common seed value negotiated with the second xPU during the test preparation phase (i.e., the handshake process described in S202), thus eliminating the need for a synchronization handshake during reverse testing and ensuring that the generated reference sequence is consistent with the sequence start phase of the second xPU transmitter.
[0060] S403: Perform a bit-by-bit XOR operation between the actual captured inverse input bit stream and the locally generated reference sequence in the physical layer hard core. Bits that are inconsistent between the two sequences will be marked as errors in the operation result.
[0061] S404: The number of these error bits is accumulated using a dedicated hardware counter.
[0062] The statistical results are usually output in the form of errors per million bits, and after logarithmic transformation, the bit error rate in the second direction is finally obtained. This bit error rate quantitatively reflects the signal transmission quality of the reverse link from the second xPU to the first xPU.
[0063] S405: Reverse error detection and forward error detection (S2) share the same physical layer hard core comparison and counting resources within the xPU, but are logically separate into two independent instances.
[0064] The system firmware ensures that test data in both directions has independent address spaces when written to the status register or memory, and avoids access conflicts through time-division multiplexing or hardware arbitration mechanisms, thereby supporting concurrent execution of bidirectional tests and independent data reporting.
[0065] S5, based on the bit error rates in the first and second directions, determine the health status of the local transmit link and the remote receive link, respectively. The specific implementation includes the following steps: S501: The link health assessment module receives and compares the bit error rate data from steps S2 and S4 in two directions, namely the bit error rate in the first direction (from the first xPU to the second xPU). Bit error rate in the second direction (from the second xPU to the first xPU) .
[0066] Bit error rate threshold ( and The settings are based on the PAM4 coding scheme and typical receiver sensitivity used in the 800Gbps SerDes system, with a certain engineering margin reserved.
[0067] in, Performance thresholds that are considered to require early warning This is considered a typical level of excellence in the health chain.
[0068] S502: Perform the first type of fault determination.
[0069] If satisfied and If the fault is located in the SerDes transmitter circuit of the first xPU, or in the cable connector directly connected to the first xPU and the adjacent cable segment (i.e., the near end of the link).
[0070] S503: Perform the second type of fault determination.
[0071] If satisfied and If the fault is located in the SerDes transmitter circuit of the second xPU, or in the cable connector directly connected to the second xPU and the adjacent cable segment (i.e., the near end of the other end of the link).
[0072] S504: Perform a third-class fault determination.
[0073] If both bidirectional bit error rates are high, that is, if the conditions are met simultaneously... and If the problem is found to be in the middle section of the cable (such as bending or squeezing that causes a decrease in overall performance), or if both ends of the connector are deteriorated (such as oxidation or wear).
[0074] S505: Based on the above judgment results, generate a preliminary diagnostic label containing "suspected faulty components" (such as: transmitter on side A, connector on side B, cable mid-section, etc.) and "confidence level".
[0075] The preliminary diagnostic label will be used as one of the key inputs, and together with the physical layer feature parameters extracted in subsequent steps S6 and S7, it will be submitted to step S8 for comprehensive health assessment and decision-making.
[0076] S6, at the SerDes receiver of any xPU, perform eye diagram sampling on the captured signal and extract four physical layer feature parameters: eye height, eye width, jitter peak-to-peak value, and signal-to-noise ratio. The specific implementation includes the following steps: S601: When the xPU's SerDes receiver is capturing a pseudo-random bitstream (from step S2 or S4), the eye diagram sampling process is initiated synchronously. This process is executed by the analog front-end hardware inside the SerDes, and its core is a multi-phase sampling mechanism based on a delay-locked loop and a comparator array.
[0077] S602: Configure horizontal sampling parameters.
[0078] For an 800Gbps system (using PAM4 encoding, symbol rate of 200GBaud), the unit symbol interval is 5ps. Within the 5ps unit interval, the horizontal position of the sampling point is precisely controlled by the delay-locked loop. At least 64 uniformly distributed sampling points are set, and the time step accuracy between adjacent sampling points is about 0.078ps (5ps / 64). The hardware design can achieve a step control accuracy of 1ps to meet the requirements.
[0079] S603: Configure vertical decision parameters and provide vertical decision thresholds through a programmable comparator array.
[0080] Considering that the nominal differential swing of the transmitter is about 800mV, the decision level range is set to cover ±400mV (a total span of 800mV).
[0081] Within this range, at least 32 different voltage levels are set, with a voltage step accuracy of 10mV between adjacent levels, to ensure that signal amplitude changes and noise levels can be effectively distinguished.
[0082] S604: Repeat the above sampling process over multiple consecutive signal periods to accumulate sufficient statistical samples.
[0083] Sampling should last at least 1 millisecond, or accumulate for more than 1 millisecond. A number of symbols (corresponding to a symbol rate of 200 GBaud) are used to ensure that the constructed eye diagram statistical matrix is statistically meaningful.
[0084] The information (time, voltage, number of occurrences) corresponding to each sampling point is recorded, and a 64 (time) × 32 (voltage) two-dimensional eye diagram statistical matrix is finally constructed. Each element value in the matrix represents the cumulative number of times the signal crosses the corresponding time-voltage coordinate point.
[0085] In summary, based on the above eye diagram statistical matrix, a predetermined algorithm is used to analyze the data distribution in the matrix and identify characteristic regions such as eye diagram opening areas and intersection points, providing a processed data basis for the specific calculations in subsequent steps S605 to S608.
[0086] S605: Calculate the first feature parameter (eye height).
[0087] First, "open regions" are determined from the eye diagram statistics matrix according to predefined rules: connected regions that appear more than 20% of the global maximum value in the matrix, and that cover at least 80% of the time units in the vertical direction.
[0088] Eye height is defined as the voltage span of the "opening area" in the vertical direction.
[0089] During the calculation, find the index of the highest voltage level corresponding to this region. and minimum voltage level index High standards : .
[0090] S606: Calculate the second feature parameter (eye width).
[0091] Eye width is defined as the time span of the aforementioned "open area" in the horizontal direction.
[0092] Find the index of the rightmost time sampling point corresponding to this region. and leftmost time sampling point index Wide eyes : .
[0093] S607: Calculate the third characteristic parameter (jitter peak-to-peak value).
[0094] This parameter reflects the time jitter range of the signal at the zero crossing point.
[0095] Extract the data closest to the zero decision level (e.g., the row with index 16) from the eye diagram statistics matrix to obtain a set of time-occurrence relationships.
[0096] Find all the local minimum points in this relationship, and their corresponding time indices are the estimated intersection points.
[0097] Collect the time index values of all intersection points, calculate the difference between the maximum and minimum values, and then multiply it by the time step to obtain the peak-to-peak value of the jitter. : .
[0098] S608: Calculate the fourth characteristic parameter (signal-to-noise ratio).
[0099] Select a 3×3 region near the center of the identified "opening region," and calculate the average voltage value (based on row index) of all points within this region, weighted by the frequency of occurrence. This average value will be used as the signal amplitude estimate. .
[0100] Within a 5% interval at each end of the eye diagram time axis, sampling points for all voltage levels are selected, and the standard deviation of the voltage values at these points is calculated as an estimate of the noise voltage standard deviation. .
[0101] The signal-to-noise ratio is calculated using the following formula: .
[0102] S7. Compare the four physical layer feature parameters with a preset threshold set. If any parameter exceeds the corresponding threshold, it is determined that the physical medium has deteriorated. The specific implementation includes the following steps: S701: Set the first feature parameter (the assessment threshold for eye height).
[0103] The eye height assessment threshold primarily ensures that the signal amplitude can still be reliably determined by the receiver after cable attenuation and noise.
[0104] For a system with a nominal swing of 800mV, considering channel insertion loss and receiver noise margin, the vertical clearance of the eye diagram opening area must be no less than 70% of the nominal swing. Accordingly, the eye height threshold is set to 560mV.
[0105] S702: Set the second feature parameter (the evaluation threshold for eye width).
[0106] The eye width assessment threshold is used to ensure that the signal still has sufficient stable sampling time when jitter exists.
[0107] Based on a unit bit period (1.25ps) and to allow for clock jitter and data correlation jitter, the time width of the eye diagram opening region must be no less than 40% of the bit period. Therefore, the eye width threshold is set to 0.5ps.
[0108] S703: Set the third characteristic parameter (evaluation threshold for jitter peak-to-peak value).
[0109] The threshold for evaluating jitter peak-to-peak values is also set based on the unit bit period.
[0110] Excessive jitter will compress the eye width and increase the risk of bit errors. In order to control the impact of total jitter on eye width within an acceptable range, the peak-to-peak jitter value shall not exceed 15% of the unit bit period, and the corresponding specific threshold is 0.1875ps.
[0111] S704: Set the fourth characteristic parameter (evaluation threshold for signal-to-noise ratio).
[0112] Sufficient signal-to-noise ratio is fundamental to ensuring a low bit error rate. According to the PAM4 coding scheme, the target bit error rate (e.g., ...) is... The required theoretical signal-to-noise ratio (SNR) value is determined, and a certain system margin is reserved. The SNR threshold is set to be no less than 15dB.
[0113] S705: Perform parameter comparison and deviation calculation.
[0114] The link health assessment module obtains the real-time eye height extracted in step S6. ), wide eyes ( ), peak-to-peak jitter ( ) and signal-to-noise ratio ( ).
[0115] The module first compares each measured value with the corresponding reference threshold set in steps S701-S704. =560mV, =0.5ps, =0.1875ps, =15dB) are compared one by one. Simultaneously, the "deviation" of each parameter can be calculated; for example, for eye height, the deviation can be defined as: (when hour).
[0116] Other parameters can be calculated similarly.
[0117] These thresholds are loaded during system initialization, serve as a baseline reference, and allow for fine-tuning via the management interface based on the actual cable type, link length, or environmental conditions.
[0118] S706: Comprehensive assessment of the degradation of the execution medium.
[0119] The evaluation module makes a judgment based on the comparison results, and the judgment strategy may vary depending on the strictness level configured in the system: Basic judgment: If the measured value of any one of the four characteristic parameters is clearly inferior to (i.e. lower or higher than) its corresponding reference threshold, a "medium anomaly" flag is generated.
[0120] Grading determination (optional): To further refine the diagnosis, grading logic can be introduced. For example: Slight degradation: Only one parameter slightly exceeds the limit (e.g., deviation within 10%), and the bidirectional bit error rate (from S5) remains excellent (e.g., below). ).
[0121] Significant deterioration: One parameter is severely out of control (deviation > 10%) or multiple parameters are out of control simultaneously.
[0122] Hierarchical determination can provide richer status information and serve different maintenance strategies.
[0123] S707: Conclusion on the health status of the generated medium.
[0124] The judgment results are encapsulated and output. The conclusion includes at least an overall status indicator (such as "normal", "abnormal" or "slightly deteriorated" or "significantly deteriorated"), and details which parameter (or parameters) exceeded the standard and the comparison between its measured value and the threshold.
[0125] In addition, the calculated parameter deviation can be output together. This conclusion, as the core diagnostic information for assessing link health from the perspective of physical media, will be provided to the subsequent step S8 for integration with the bit error rate diagnostic results to generate the final link health assessment report.
[0126] S8, combining bidirectional bit error rate and physical layer characteristic parameters, generates a link health assessment report and triggers corresponding alarms or self-healing strategies. The specific implementation includes the following steps: S801: Set comprehensive evaluation levels and judgment rules.
[0127] The link health assessment module receives the bidirectional bit error rate diagnostic conclusion from step S5 (including...). , (and preliminary fault location label) and physical medium health status conclusion from step S7 (including measured values of each parameter, deviation and overall status label).
[0128] The link health assessment module classifies link health status into four levels based on the following predefined rule set: (1) Level 1 (Excellent / Healthy): Must meet the following conditions simultaneously: Both bidirectional bit error rates are below the excellent threshold (i.e.) and ); The physical medium state conclusion is "normal", meaning that the measured values of all four eye diagram feature parameters (eye height, eye width, jitter peak-to-peak value, and signal-to-noise ratio) are better than (eye height, eye width, and signal-to-noise ratio are higher than, and jitter peak-to-peak value is lower than) their corresponding thresholds set in S7.
[0129] (2) Level 2 (Good / Sub-healthy): Must meet the following conditions simultaneously: Both bidirectional bit error rates are below the warning threshold (i.e.) and ); The physical medium condition does not meet the "Level 1" standard, but does not trigger the "Level 3" condition. Typical cases include at least one eye diagram parameter measured value being in the "near threshold" state. Here, "near threshold" can be defined as the measured value being between 95% and 100% of the corresponding threshold (for eye height, eye width, and signal-to-noise ratio) or between 100% and 105% (for jitter peak-to-peak value). This tolerance range can be adjusted through the management interface.
[0130] (3) Level 3 (Early Warning): Triggered when any of the following conditions are met: The bit error rate in either direction or both directions is within the warning range (i.e.) ); The physical medium condition conclusion indicates the presence of "deterioration" (i.e., at least one eye diagram parameter clearly exceeds the threshold, or is judged as "slight deterioration" / "significant deterioration" according to the S706 rule); Early warnings triggered by the trend analysis module (see S802).
[0131] (4) Level 4 (Fault): Bit error rate in any direction or both directions is higher than or equal to .
[0132] S802: Execution trend analysis and early warning.
[0133] The evaluation module maintains a historical database of link states, continuously recording the timestamp, bidirectional bit error rate, measured values of four physical layer parameters, and calculated deviation for each test.
[0134] The database supports data analysis by configurable time windows (e.g., the most recent 24 hours), and the trend analysis algorithm examines the numerical sequence of each key physical layer parameter (such as eye height and signal-to-noise ratio) within a recent window.
[0135] When the system detects that a parameter’s deviation (or measured value) shows a statistically significant monotonically decreasing trend (for eye height, eye width, signal-to-noise ratio) or monotonically increasing trend (for jitter peak-to-peak value) in N consecutive tests (e.g., N=3), even if the current comprehensive evaluation level is still Level 2 (Good) according to the S801 rule, the trend analysis module will proactively generate an “early warning” event, which will directly lead to the comprehensive evaluation level being upgraded to Level 3 (Warning).
[0136] Both parameter N and the threshold for determining trend significance can be configured according to the operation and maintenance strategy.
[0137] S803: Generate and output a link health assessment report.
[0138] Based on the level determination results of step S801 and the trend analysis conclusions of S802, the link health assessment module generates a structured assessment report. The core content of the report should include: Final comprehensive evaluation level (Level 1 / Level 2 / Level 3 / Level 4); Performance details: Specific values for bidirectional bit error rate and comparison with the threshold; Physical layer diagnostic details: measured values, reference thresholds, deviations of four eye diagram characteristic parameters, and the conclusion on the medium state given in step S7; Fault location reference: Preliminary fault location label from step S5 (e.g., "Suspected A-side transmitter"); Trend Analysis Summary (if applicable): Indicate which parameter shows a deterioration trend and its severity; Timestamp and Test ID: Used for tracking and tracing.
[0139] This report will serve as the basis for decisions to trigger subsequent adaptive control (S804) or recovery / switchover strategies (S805), and can also be reported to the cluster management system or operation and maintenance platform.
[0140] S804: Perform adaptive parameter optimization for Level 3 (early warning) status.
[0141] When the link is assessed as Level 3 (early warning), the adaptive control module is activated, aiming to restore the link to Level 2 or Level 1 by optimizing physical layer parameters. The optimization process is guided by the diagnostic information from steps S5 and S7: (1) Targeted parameter adjustment: If the initial location of the problem in S5 indicates that it may be a problem at the transmitter or near end, then prioritize optimizing the transmission parameters of the suspected problem end. For example, if the first xPU (A side) transmitter is suspected, then adjust the deemphasis coefficient of the SerDes transmitter driver on the A side. The adjustment direction is usually to increase the compensation amount, that is, based on the default value (e.g., 6dB), increase it in 1dB increments towards the maximum settable value (e.g., 12dB). After each adjustment, re-execute the test procedure from S1 to S8.
[0142] If the S7 indicates that the physical medium parameters (such as eye height and signal-to-noise ratio) are mainly degraded, then it may be possible to simultaneously or sequentially try to optimize the transmitter pre-emphasis / deemphasis and the receiver equalizer.
[0143] (2) Optimization objectives and termination conditions: Verification tests are conducted after each parameter adjustment to optimize the bit error rate corresponding to the direction (e.g., adjusting the transmitter on side A, then observe...). Whether to restore to the warning threshold The following are the main success criteria; at the same time, observe whether the eye diagram parameters improve.
[0144] The optimization process continues until the target is reached, the limit of the parameter adjustment range is reached, or the preset maximum number of attempts is exceeded.
[0145] (3) Receiver equalizer reconfiguration: As an auxiliary or parallel means, an adaptive equalization algorithm can be used to optimize the receiver equalizer.
[0146] With the PRBS test stream continuously transmitted, the equalizer tap weights are updated using an algorithm such as the minimum mean square error, leveraging the known characteristics of the PRBS sequence. The formula can be expressed as: .
[0147] in, and These are the updated and current tap weight vectors, respectively. Step size factor The error signal generated by comparing the received signal, after equalization and decision, with the local PRBS reference sequence. The input signal vector.
[0148] This process is designed to automatically compensate for channel distortion.
[0149] S805: Execute the recovery and handling strategy for Level 4 (fault) status.
[0150] When a link is determined to be at level four (fault), it indicates severe performance degradation, and the adaptive control module will perform more in-depth recovery operations or initiate a handling process: (1) Physical layer link retraining: Trigger the SerDes physical layer control state machine to execute the complete retraining process.
[0151] This includes recalibrating the clock data recovery circuit, re-optimizing the equalizer coefficients, and re-establishing a stable symbol lock, in an attempt to restore communication capabilities at the hardware level.
[0152] This process does not involve a change in communication rate.
[0153] (2) Service channel switching and maintenance alarms: If the current system has a preset redundant link (e.g., a hot standby SerDes channel or cable), the affected service data stream will be automatically switched seamlessly or interrupted to the backup link, and the original faulty link will be marked as offline.
[0154] Regardless of whether redundant channels exist, the system will generate the highest priority fault alarm, along with a complete link health assessment report (including fault location suggestions for S5 and media degradation details for S7), and report it to the cluster operation and maintenance system through the management interface, clearly indicating that manual intervention is required for maintenance.
[0155] The original faulty link will be isolated by the system before it is repaired and passes the full S1-S8 test verification to prevent it from being reused.
[0156] Example 2 The system proposed in this invention, to implement the aforementioned intelligent testing method for high-speed network communication ports, specifically includes a bidirectional PRBS collaborative generation module, a bidirectional bit error detection module, an eye diagram feature extraction module, a link health assessment module, and an adaptive control module. The structure and function of each module are as follows: The bidirectional PRBS collaborative generation module is built into two directly connected xPUs. Its core function is to synchronously generate identical pseudo-random bit sequences at their respective SerDes transmitters, providing an excitation source for bidirectional testing. Specifically: Implementation details: Each xPU contains a PRBS31 sequence generator hardware, whose generator polynomial is... The sequence length is The sequence generation rate is strictly synchronized with the nominal rate of the SerDes port (e.g., 800Gbps).
[0157] Synchronization mechanism: Initial sequence synchronization between two xPUs is accomplished through a dedicated out-of-band management channel. This channel uses Manchester encoding format and has a transmission rate of 100Mbps. The handshake protocol is controlled by the management processor within the xPU, which exchanges and confirms the initial seed value of the PRBS sequence to ensure that the starting phase of the test sequences at both ends is consistent.
[0158] Signal processing: This module is integrated into the SerDes physical layer hard core. The generated bit stream is pre-emphasized / de-emphasized by the transmitter driver (default coefficient is 6dB) and then modulated into a differential signal output.
[0159] A bidirectional bit error detection module is installed at the SerDes receiver of each of the two xPUs, responsible for real-time detection and quantification of signal quality in both transmission directions. Specifically: Specific implementation: Each receiver integrates this module, the core of which is a high-speed comparison circuit. When it receives the PRBS test stream sent by the other end, the clock data recovery circuit first locks and recovers the sampling clock.
[0160] The comparison logic uses the recovered clock to drive the local PRBS31 reference sequence generator (synchronized with the transmitter) to perform a bit-by-bit hardware XOR operation between the received bit stream and the local reference sequence. Bits that result in "1" are identified as errors.
[0161] Error statistics: Error bits are accumulated in real time by a dedicated 32-bit hardware counter. The statistical results are output in the form of errors per million bits, and are calculated and converted into the bit error rate for the corresponding direction (e.g., ...). The entire comparison and statistical delay does not exceed 100 nanoseconds to meet the real-time requirements at 800Gbps.
[0162] The eye diagram feature extraction module, integrated into the analog front-end of the xPU's SerDes receiver, is used for high-precision sampling of the input signal and extraction of physical layer features. Specifically: Specific implementation: Its core is a multi-phase sampling system based on a delay-locked loop and a programmable comparator array.
[0163] Sampling mechanism: Within a unit symbol / bit period, the delay-locked loop provides high-precision time control in the horizontal direction, with no fewer than 64 sampling points and a time step accuracy of 1 picosecond (ps); at the same time, the programmable comparator array provides decision thresholds in the vertical direction, with no fewer than 32 voltage levels and a voltage step accuracy of 10 millivolts (mV).
[0164] Data processing: By continuously sampling the test signal (e.g., accumulating more than...) (Symbols) to construct a 64×32 two-dimensional eye diagram statistical matrix. After analog-to-digital conversion, the sampled data is processed by an embedded digital signal processor or dedicated hardware logic, which calculates and outputs four characteristic parameters from the matrix according to a predetermined algorithm: eye height, eye width, jitter peak-to-peak value, and signal-to-noise ratio.
[0165] The link health assessment module, running on the xPU's management processor (or dedicated coprocessor), acts as the system's "brain," responsible for comprehensive analysis and decision-making. Specifically: Data aggregation: This module receives bit error rate data from the bidirectional bit error detection module and four physical layer parameters from the eye diagram feature extraction module through the internal data bus or register interface.
[0166] Evaluation logic: The module has a built-in evaluation rule engine (as described in S5, S7, and S8 of Example 1), which evaluates the rules based on preset thresholds (such as the bit error rate threshold). The system uses a combination of parameters such as eye height threshold (560mV, etc.) and judgment logic to comprehensively diagnose and classify the link status (excellent, good, warning, fault).
[0167] Historical and Trend Analysis: This module maintains a historical database of link status, records data from each test, and supports trend analysis based on time windows. When a continuous deterioration trend of key physical parameters is detected, an early warning can be triggered.
[0168] Report generation: The final output is a structured link health assessment report, which includes status level, detailed parameters, diagnostic conclusions and trend information.
[0169] The adaptive control module, acting as the system's "actuator," dynamically adjusts the physical layer parameters based on the evaluation results. Specifically: Control Interface: This module directly accesses and configures the physical layer control registers of SerDes via a memory-mapped register interface or a dedicated configuration bus.
[0170] Regulation strategies: When the assessment is in an "early warning" state, the module performs parameter optimization, such as adjusting the de-emphasis coefficient of the transmitter driver in 1dB steps (the range is usually 0-12dB), or triggering the receiver equalizer to perform adaptive reconfiguration of tap weights using algorithms such as minimum mean square error. When assessed as a "fault" state, the module triggers more thorough recovery operations, such as initiating a SerDes physical layer link retraining process (renegotiating parameters, calibrating clocks and equalizers), or performing automatic switching of service flows when redundant channels are available.
[0171] Response speed: Through direct control at the register level, this module can achieve millisecond-level parameter adjustment and response.
[0172] In summary, this embodiment utilizes the collaborative work of the five modules to form a complete embedded intelligent testing system. Its workflow begins with the bidirectional PRBS collaborative generation module generating test stimuli at both ends; the bidirectional bit error detection module and eye diagram feature extraction module work in parallel, collecting raw link performance data from both digital bit error rate and analog waveform dimensions; this data is aggregated by the link health assessment module for fusion intelligent diagnosis; finally, the diagnostic conclusions drive the adaptive control module to execute a closed-loop operation from parameter fine-tuning to link recovery. Furthermore, this system is fully embedded within the xPU, requiring no external instruments, achieving fully automated intelligent operation and maintenance capabilities for high-speed direct-connect links, from monitoring and diagnosis to self-optimization, thereby improving the reliability and availability of large AI clusters.
[0173] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention. Therefore, the embodiments should be regarded as exemplary and non-limiting in all respects.
[0174] Furthermore, it should be understood that although this specification describes embodiments, not every embodiment includes only one independent technical solution. This narrative style is merely for clarity. Those skilled in the art should consider the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.
Claims
1. A method for intelligent testing of high-speed network communication ports, applied to a first processing unit and a second processing unit directly interconnected by a high-speed cable, characterized in that, The method includes: A bidirectional pseudo-random bit sequence test is performed between the first processing unit and the second processing unit to obtain the bit error rate in the first direction and the second direction, respectively. Fault location is based on bidirectional bit error rate, which is used to distinguish whether the fault originates from the transmitting circuit, the receiving circuit, or the physical medium of the high-speed cable. At the receiving end of the first processing unit and / or the second processing unit, eye diagram sampling is performed on the received test signal to extract physical layer feature parameters including at least eye height, eye width, jitter and signal-to-noise ratio; The physical medium health status of high-speed cables is evaluated based on the comparison results of physical layer characteristic parameters and preset thresholds. Based on the fault location results of the combined bidirectional bit error rate and the health status of the physical medium, a health level assessment is performed on the end-to-end link composed of high-speed cables. Based on the health level assessment results, trigger the corresponding link parameter adaptive adjustment strategy or link recovery strategy.
2. The intelligent testing method for high-speed network communication ports according to claim 1, characterized in that, Performing a bidirectional pseudo-random bit sequence test includes: A first test bit stream is generated by a first pseudo-random sequence generator located at the transmitting end of the first processing unit and transmitted to the receiving end of the second processing unit via a high-speed cable. At the receiving end of the second processing unit, the sampling clock is locked and restored using a clock data recovery circuit, which drives the local first reference sequence generator to generate a reference sequence synchronized with the first test bit stream. By comparing the received bit stream with the reference sequence, the bit error rate in the first direction is calculated. A second test bit stream is generated by a second pseudo-random sequence generator located at the transmitting end of the second processing unit, and then transmitted in reverse to the receiving end of the first processing unit via a high-speed cable. At the receiving end of the first processing unit, synchronous clock recovery and sequence comparison are performed to calculate the bit error rate in the second direction.
3. The intelligent testing method for high-speed network communication ports according to claim 2, characterized in that, Before initiating the bidirectional pseudo-random bit sequence test, the method further includes: A handshake synchronization is performed between the first processing unit and the second processing unit through a dedicated out-of-band management channel to negotiate and confirm the initial seed values used for the first pseudo-random sequence generator and the first reference sequence generator, ensuring that the starting phase of the test sequence is consistent.
4. The intelligent testing method for high-speed network communication ports according to claim 1, characterized in that, Fault location based on bidirectional bit error rate specifically includes: If the bit error rate in the first direction is higher than the first warning threshold and the bit error rate in the second direction is lower than the second health threshold, then the fault is determined to be located in the transmitting circuit of the first processing unit or the connection part of the high-speed cable near the first processing unit. If the bit error rate in the second direction is higher than the first warning threshold and the bit error rate in the first direction is lower than the second health threshold, then the fault is determined to be located in the transmitting circuit of the second processing unit or the connection part of the high-speed cable near the second processing unit. If the bit error rate in both the first and second directions is higher than the first warning threshold, it is determined that the fault may be located in the middle section of the high-speed cable or that the connectors at both ends are degraded.
5. The intelligent testing method for high-speed network communication ports according to claim 1, characterized in that, Eye diagram sampling is performed on the received test signal to extract physical layer feature parameters, including: Within a unit symbol interval, multi-phase sampling in the horizontal direction is controlled by a delay-locked loop, and multiple decision thresholds in the vertical direction are set by a programmable comparator array. Accumulate sampling points over multiple consecutive signal periods to construct an eye diagram statistical matrix; Based on the eye diagram statistical matrix, calculate and output eye height, eye width, jitter peak-to-peak value, and signal-to-noise ratio.
6. The intelligent testing method for high-speed network communication ports according to claim 1, characterized in that, Based on the fault location results of the combined bidirectional bit error rate and the health status of the physical medium, the health level assessment of the end-to-end link includes: The link health status is divided into multiple levels, and the judgment rules include at least the following: when the bit error rate in both directions is below the excellent threshold and the physical medium health status is normal, it is evaluated as a healthy level; when the bit error rate in either direction is within the warning range or the physical medium health status indicates deterioration, it is evaluated as a warning level; when the bit error rate in either direction is above the fault threshold, it is evaluated as a fault level.
7. The intelligent testing method for high-speed network communication ports according to claim 6, characterized in that, The method also includes trend analysis and early warning steps: Maintain a historical database of link status, recording the bit error rate and physical layer characteristic parameters of each test; Trend analysis is performed on physical layer characteristic parameters. When a specified parameter is detected to show a statistically significant monotonic degradation trend in multiple consecutive tests, an early warning event is proactively triggered even if the current health level has not reached the warning level, and the health level of the link is adjusted to the warning level.
8. The intelligent testing method for high-speed network communication ports according to claim 6, characterized in that, Based on the health level assessment results, the corresponding adaptive adjustment strategies for link parameters include: When the health level assessment result is a warning level, the adaptive control process is initiated. The process includes: based on the fault location result, adjusting the pre-emphasis or deemphasis coefficient of the transmitter driver of the suspected problem terminal; and / or, triggering the receiver equalizer to perform adaptive reconfiguration to compensate for channel distortion. After each parameter adjustment, the steps from bidirectional pseudo-random bit sequence testing to health level assessment are re-executed until the link performance recovers to the target health level or the adjustment attempt limit is reached.
9. The intelligent testing method for high-speed network communication ports according to claim 6, characterized in that, Based on the health level assessment results, the corresponding link recovery strategies triggered include: When the health level assessment result is a fault level, the physical layer link retraining process is triggered to recalibrate the clock data recovery circuit and optimize the equalizer coefficients. If a preset redundant link exists, the service data stream will be automatically switched to the redundant link, and an alarm message containing fault location and media degradation details will be generated and reported.
10. An intelligent testing system for high-speed network communication ports, applied to a first processing unit and a second processing unit directly interconnected by a high-speed cable, characterized in that... The system includes: The bidirectional PRBS collaborative generation module is used to synchronously generate a pseudo-random bit sequence test stream at the transmitting ends of the first and second processing units and send it to the other end via a high-speed cable. The bidirectional bit error rate detection module is used at the receiving end of the processing unit to capture the test stream sent by the other end, compare it with the locally synchronously generated reference sequence, and calculate and output the bit error rate in both transmission directions in real time. The eye diagram feature extraction module is integrated into the receiving end of the processing unit. It is used to perform high-precision eye diagram sampling on the input test signal and extract physical layer feature parameters including at least eye height, eye width, jitter and signal-to-noise ratio. The link health assessment module is used to receive bit error rate and physical layer characteristic parameters, locate faults based on bit error rate, assess the medium status based on physical layer characteristic parameters, and combine the results of both to assess the health level of the end-to-end link. The adaptive control module is used to dynamically adjust the SerDes physical layer parameters, trigger link retraining, or perform service channel switching operations based on the health level assessment results.