Three-dimensional topology self-sensing management system based on PCIe equipment

The three-dimensional topology self-aware management system based on PCIe devices solves the problems of device identity confusion, resource allocation bottlenecks and weak topology management capabilities of traditional PCIe devices in large-scale clusters. It realizes the self-description and accurate positioning of devices, and improves the computing power utilization and stability of the cluster.

CN122070537APending Publication Date: 2026-05-19HAILI COMPUTING (BEIJING) TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HAILI COMPUTING (BEIJING) TECHNOLOGY CO LTD
Filing Date
2025-12-05
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Traditional PCIe bus devices suffer from problems such as device identity confusion, resource configuration bottlenecks, weak topology management capabilities, and low computing power utilization due to management dependence on the main CPU in large-scale, high-performance computing scenarios, making it difficult to adapt to the flexible resource scheduling of large-scale clusters.

Method used

A 3D topology self-aware management system based on PCIe devices is adopted. Through global device group coding, enhanced resource declaration and dynamic topology awareness mechanism, combined with FPGA acceleration and heterogeneous computing hardware, it realizes device self-description, environmental awareness and precise positioning, supports dynamic reconfiguration, and builds a multi-dimensional coding management system.

Benefits of technology

It improves equipment management efficiency, enhances the overall computing power utilization of the cluster, reduces fault location time, strengthens the flexibility of equipment authentication and resource allocation, ensures the stability and compatibility of the cluster, and adapts to the complex application needs of large-scale computing power clusters.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122070537A_ABST
    Figure CN122070537A_ABST
Patent Text Reader

Abstract

The invention discloses a three-dimensional topology self-sensing management system based on PCIe equipment, the management system is applied to a PCIe board card, the management system is characterized in that the PCIe board card comprises a virtual CPU management end and a PCIe board end, the PCIe board end is standard PCIe equipment which comprises the three-dimensional topology self-sensing management system and accords with a preset standard specification, and the virtual CPU management end is connected with the PCIe board end. The virtual CPU management end is installed on a standard PCIe bus, and the virtual CPU management end reads and writes the PCIe board end through a computer bus. Through the three-dimensional topology self-sensing management system, equipment can be converted into an intelligent entity with self-description, environment sensing and precise positioning capabilities from a passive resource consumer, so that rapid topology positioning is realized, dynamic reconfiguration is supported, the resource demand bottleneck of high-performance equipment is solved, and the service life of the equipment is prolonged. And flexible resource scheduling of a large-scale cluster is adapted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention generally relates to the field of computer bus technology. More specifically, this invention relates to a three-dimensional topology self-sensing management system based on PCIe devices. Background Technology

[0002] In traditional computer systems, PCIe bus devices rely on BDF addresses (bus number, device number, function number) for addressing and combine vendor ID and device ID to identify the type. This static and passive enumeration mechanism has many key problems and is difficult to adapt to the needs of large-scale, high-performance computing scenarios.

[0003] First, there is a lack of globally unique identifiers. The BDF address is dynamically allocated by the software and may change with each system startup. Furthermore, the vendor ID and device ID can only distinguish the device type and cannot identify multiple specific device instances of the same type, leading to device identity confusion and difficulty in tracing faults. This is especially true in large-scale clusters with more than 200 devices, where device management efficiency is extremely low.

[0004] Secondly, there is a bottleneck in resource allocation. The traditional 32-bit base address register (BAR) only supports 4GB of address space, which is difficult to meet the large memory mapping requirements of high-performance devices such as GPUs and FPGAs. In addition, the validity of resources needs to be determined by readback detection, resulting in low configuration efficiency and inability to dynamically adapt to changes in device resource requirements.

[0005] Furthermore, the system has weak topology management capabilities, lacks a precise device location positioning mechanism, and cannot build a global topology view. When a device fails, it needs to rely on manual troubleshooting or long-term enumeration to locate the device, which usually takes more than 30 minutes and seriously affects the system's operation and maintenance efficiency. At the same time, the system has difficulty in perceiving the difference between the physical number of devices and the enumerated number, and cannot provide early warnings of hardware failures or configuration errors.

[0006] Finally, management relies on the main CPU. Tasks such as device enumeration, status monitoring, and parameter configuration all require main CPU resources, which prevents the main CPU from focusing on computing power. This results in low overall cluster computing power utilization. Furthermore, traditional management solutions have poor compatibility and require modification of the operating system kernel during deployment, increasing the difficulty of deploying large-scale clusters and the risk to stability.

[0007] In view of this, there is an urgent need to provide a three-dimensional topology self-aware management system based on PCIe devices in order to quickly and accurately adapt to the flexible resource scheduling of large-scale clusters. Summary of the Invention

[0008] In order to at least solve one or more of the technical problems mentioned above, the present invention proposes a three-dimensional topology self-sensing management system based on PCIe devices in several aspects.

[0009] In a first aspect, the present invention provides a three-dimensional topology self-sensing management system based on a PCIe device. The management system is applied to a PCIe board, characterized in that the PCIe board includes a virtual CPU management terminal and a PCIe board terminal, wherein the PCIe board terminal is a standard PCIe device conforming to a preset standard specification and containing the three-dimensional topology self-sensing management system, and the virtual CPU management terminal is installed on a standard PCIe bus, which reads and writes to the PCIe board terminal through a computer bus.

[0010] In some embodiments, the preset standard specification is a global device group coding, and the global device group coding rules include: a device declaration area, a standard PCIe configuration space header area compatibility layer, a standard PCIe capability and status register group, an encrypted heritable device DNA identifier register, a device resource and topology management register group, a device historical status and reliability management register group, and a device communication affinity and behavior measurement register group, wherein the device resource and topology management register group includes a three-dimensional topology self-sensing management system.

[0011] In some embodiments, the three-dimensional topology self-aware management system includes: an enhanced resource declaration mechanism and / or dynamic topology awareness and verification.

[0012] In some embodiments, the device resource and topology management register group works collaboratively through a three-dimensional topology self-sensing management system, including: intelligent resource allocation, advanced diagnostics and maintenance, and / or collaborative computing paradigms.

[0013] Through the aforementioned 3D topology self-aware management system based on PCIe devices, this invention integrates two innovative mechanisms—"enhanced resource declaration" and "topology state awareness"—to transform devices from passive resource consumers into intelligent entities with self-description, environmental awareness, and precise positioning capabilities. Furthermore, it meets large resource demands through scalable resource declarations, achieves rapid topology positioning through 3D coordinates and snapshot verification, and supports dynamic reconfiguration, thereby solving the resource demand bottleneck of high-performance devices and enabling flexible resource scheduling for large-scale clusters. Attached Figure Description

[0014] The above and other objects, features, and advantages of exemplary embodiments of the present invention will become readily apparent upon reading the following detailed description with reference to the accompanying drawings. In the drawings, several embodiments of the invention are illustrated by way of example and not limitation, and like or corresponding reference numerals denote like or corresponding parts, wherein:

[0015] Figure 1 An exemplary structural block diagram of a PCIe board device according to an embodiment of the present invention is shown;

[0016] Figure 2 A diagram illustrating the structure of global device grouping coding rules according to an embodiment of the present invention is shown.

[0017] Figure 3 A structural diagram of a three-dimensional topology self-sensing management system based on a PCIe device according to an embodiment of the present invention is shown. Detailed Implementation

[0018] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0019] Figure 1 An exemplary structural block diagram of a PCIe board device according to an embodiment of the present invention is shown. Figure 1 As shown, the PCIe board device 100 runs on the PCIe board, which includes a virtual CPU management terminal 101 and a PCIe board terminal 102. The PCIe board terminal is a standard PCIe device that conforms to a preset standard specification and includes the three-dimensional topology self-sensing management system 1021. The virtual CPU management terminal is installed on the standard PCIe bus and reads and writes to the PCIe board terminal through the computer bus.

[0020] The following description uses a large-scale computing cluster scenario as an example. This embodiment can manage over 200 PCIe devices (including GPU accelerator cards, NVMe SSDs, etc.) simultaneously, achieving efficient device control using the encoding management system provided in this embodiment. The large-scale PCIe device encoding management system is integrated onto a standard PCIe board, which may contain two core modules: a virtual CPU management terminal and a PCIe board terminal.

[0021] The PCIe board can be a standard PCIe device conforming to the EX-DCC global device grouping coding specification. Specifically, it can be a 16-way PCIe 4.0 expansion interface module, with each interface capable of connecting peripherals such as GPUs and NVMe hard drives. All peripherals expand the traditional PCIe configuration space and have built-in encrypted genetic device DNA identifier registers, device resource and topology management register groups, etc., to ensure compliance with preset standard specifications. The virtual CPU management end can be an independent board unit conforming to the PCIe 4.0 specification, installed in a standard PCIe slot (bus width x16) on the cluster host. Its hardware configuration includes an ARM Cortex-A53 embedded CPU (1.8GHz), a 1TB onboard NVMe SSD (for storing device information and operation logs), and 2GB DDR4 cache. It can run an embedded Linux operating system and dedicated management software.

[0022] After system startup, the virtual CPU management terminal automatically enumerates all PCIe board peripherals via the computer's PCIe bus, reads the DNA identifier, topology information, and other data of each peripheral, and writes them to the onboard SSD. Simultaneously, it monitors the peripheral's operating status in real time, writing log information such as temperature and bandwidth usage to the onboard cache every minute, and then periodically synchronizing it to the SSD for persistent storage. When peripheral parameters need to be configured, the virtual CPU management terminal writes instructions to the device resource and topology management register groups on the PCIe board via the bus to complete the parameter adjustment, without requiring intervention from the main CPU. Furthermore, this PCIe board is logically recognized by the main CPU as a standard PCIe storage controller, allowing it to adapt to mainstream Windows and Linux systems without modifying the operating system kernel, ensuring cluster compatibility.

[0023] The "virtual CPU management terminal + standard PCIe board terminal" architecture of the above-described embodiment solves the problems of dynamic BDF address changes, lack of globally unique identifiers, and heavy management burden on the main CPU in traditional PCIe device management. On the one hand, the encrypted DNA identifier under the EX-DCC specification assigns a globally unique identity to each device, avoiding confusion between instances of the same type. On the other hand, the virtual CPU management terminal independently undertakes device enumeration, status monitoring, and configuration management tasks, allowing the main CPU to focus on computing power. According to the inventors' experimental calculations, the overall computing power utilization of the cluster can be improved by more than 15%. At the same time, the system is compatible with the traditional PCIe specification, can support the stable operation of more than 200 devices, meets the complex application requirements of large-scale computing power clusters, and deployment does not require modification of the system kernel, ensuring the stability and security of the cluster operation.

[0024] Furthermore, in one embodiment, in a large-scale computing cluster scenario (managing 200+ GPU / NVMe devices), the PCIe board is an onboard memory, and the virtual CPU management terminal is a board compliant with the PCIe specification. Complex device management is achieved through a "multi-level storage architecture + heterogeneous computing hardware + intelligent scheduling algorithm".

[0025] Specifically, the onboard memory on the PCIe board can adopt a three-tier architecture of "DDR5 cache + SLC NAND high-speed cache + QLC NAND main storage": the total capacity can be configured as 2GB DDR5 (timing 3200MT / s, used for real-time data temporary storage), 1TB SLC NAND (100,000 erase / write cycles, for storing hot data), and 8TB QLC NAND (for storing cold data), and redundancy is built through RAID 5 array technology to avoid single point of failure.

[0026] In one implementation scenario, the memory is divided into three functional partitions: ① Device DNA identification area (10GB SLC, read-only, storing the AES-256 encrypted DNA code of each peripheral); ② Real-time log area (200GB DDR5+SLC hybrid, read / write speed ≥2GB / s, storing data such as device temperature, bandwidth usage, and error count within 1 hour); ③ Historical data area (8TB QLC, using the ZFS file system, storing historical logs within 30 days, supporting compression and snapshots).

[0027] In one implementation scenario, the virtual CPU management terminal is configured as a PCIe 5.0 compliant board (16 bus width, 64GB / s bandwidth), employing a heterogeneous computing architecture: the core consists of two ARM Cortex-A76 embedded CPUs (2.4GHz clock speed, responsible for logic control), paired with one Xilinx Artix-7 FPGA (responsible for data acceleration processing), and integrated with a 16-channel PCIe switch (supporting device enumeration and bandwidth allocation). The board runs a customized embedded Linux system, pre-installed with a multi-level data scheduling algorithm and a device health assessment model. The scheduling algorithm dynamically migrates hot data (such as the device status accessed frequently within 1 minute) to SLC / NAND and archives cold data (such as logs from 24 hours ago) to QLC by real-time monitoring of the IOPS (Input / Output Operations Per Second) and bandwidth utilization of each storage partition. The health model is based on the device temperature collected by the memory (weight 40%), PCIe bus error count (weight 30%), and bandwidth fluctuation (weight 30%). It outputs a health score of 0-100 by weighted summation, and triggers an alarm when the score is below 60.

[0028] The specific operating steps of the system are as follows: 1. Upon startup, the FPGA initializes the PCIe Switch using pre-programmed firmware, automatically scans the onboard memory of the 16 PCIe interfaces, obtains the device DNA identifier, and verifies it via AES-256 decryption; 2. The embedded CPU loads a scheduling algorithm to write the real-time collected device logs (temperature, bandwidth, etc.) to the DDR5 cache using PCIe DMA technology (avoiding CPU interrupt overhead), and synchronizes the cached data to the SLC NAND every 5 minutes; 3. At 2:00 AM every day, the algorithm automatically compresses logs exceeding 24 hours in the SLC (using compression algorithm LZ4) and archives them to the QLC, generating a data snapshot; 4. The health model reads the device data in the memory every 10 minutes, calculates the health score, and issues a dual warning via onboard LED indicators and a system pop-up window if the score is below the threshold.

[0029] The combination of the multi-level storage architecture and heterogeneous computing hardware described in the above embodiments is significantly different from the traditional single storage + single CPU PCIe management scheme: On the one hand, the three-level storage architecture can reduce data read and write latency to below 50μs, which is 40% higher than the traditional solution, and RAID 5 and snapshot technology ensure data reliability and avoid log loss; on the other hand, FPGA accelerates device enumeration and data transmission, and the embedded CPU combined with the weighted health model significantly improves the accuracy of device fault early warning and reduces ineffective operation and maintenance; at the same time, the partition management and scheduling algorithm of the onboard memory supports efficient storage of log data from 200+ devices, reduces the average daily storage occupancy of a single device, further reduces the burden on the main CPU and cluster storage, and ensures the stable operation of large-scale computing clusters.

[0030] further, Figure 2 A structural diagram of a global device grouping coding rule 200 according to an embodiment of the present invention is shown.

[0031] As shown in rule 200, in one embodiment, the preset standard specification is EX-DCC (Expanse Domain Clustering Code) global device clustering coding. Specifically, global device clustering coding rule 200 may include: a device declaration area 201, a standard PCIe configuration space header area compatibility layer 202, a standard PCIe capability and status register group 203, an encrypted heritable device DNA identifier register 204, a device resource and topology management register group 205, a device historical status and reliability management register group 206, and a device communication affinity and behavior measurement register group 207, wherein the device resource and topology management register group includes a three-dimensional topology self-aware management system 2051.

[0032] In large-scale computing cluster scenarios (managing 200+ GPU / NVMe devices), the "hardware encryption + FPGA acceleration + reinforcement learning algorithm" technical solution provided in this embodiment of the invention can construct a multi-dimensional encoding management system, achieving a fundamental difference from traditional PCIe management solutions. Specific details are as follows:

[0033] In one implementation scenario, EX-DCC encoding can adopt a 512-bit fixed-length structure, divided into a "compatibility layer (64-bit) + extended function layer (448-bit)". The compatibility layer retains the traditional PCIe configuration space header (first 64 bytes) to ensure that the main CPU can recognize it. The extended function layer corresponds to seven register groups, each of which is allocated an independent address segment (address step size of 4 bytes). It can be directly read and written by the virtual CPU management terminal through the PCIe bus, and the main CPU has no access rights (implemented through FPGA address filtering).

[0034] The specific technical details and implementation steps of the seven major register groups are as follows:

[0035] 1. Device Declaration Area (32-bit Register): The EX-DCC encoded "identity foundation layer," which fundamentally addresses the lack of ownership traceability and key association in traditional PCIe devices. This declaration area can provide declaration information issued by the product owner, company name, information on authorized users, custom text information, and a key storage area for key management when providing software development tools in the future.

[0036] For example, field definitions can use 16-bit device type codes (0x0001 = GPU, 0x0002 = NVMe), 8-bit cluster affiliation codes (0x01 = computing cluster A, 0x02 = computing cluster B), and 8-bit device status codes (0x00 = offline, 0x01 = online, 0x02 = warning).

[0037] In terms of technical implementation, a "one-time programming + dynamic encrypted update" mechanism can be adopted, combined with a TPM security chip to achieve full lifecycle identity management. When the device leaves the factory, the device type code and ownership code are programmed by laser (which cannot be modified). The status code is updated by the virtual CPU management terminal every 50ms after reading the device hardware signals (such as power enable and link activation). The update command is transmitted encrypted using the SM4 national cryptographic algorithm.

[0038] 2. Standard PCIe Configuration Space Header Compatibility Layer (64-bit Register). This design adopts a strategy that emphasizes both compatibility and expansion, directly compatible with and referencing the PCI Express standard configuration space header layout defined by companies such as Intel, using this as the hardware foundation and logical starting point for the entire innovative register system.

[0039] The core purpose and advantages of the above design are:

[0040] 1) By strictly adhering to the standard specifications defined by the PCI-SIG organization, it ensures that hardware devices equipped with this innovative function can be correctly recognized as "standard" PCIe devices by existing standard motherboards, operating systems and BIOSes, and complete the most basic enumeration and initialization process.

[0041] 2) Clear Functional Positioning: This part of the registers (such as Vendor ID, Device ID, Class Code, Status, Command, etc.) defines the most basic identification, basic functions, and status of the device, thus clarifying the fundamental attribute of the device as a "general-purpose I / O device" and positioning its innovation scope as "enhancing memory consistency and topology management capabilities on top of the standard PCIe architecture," rather than creating a completely new and incompatible bus type. For example, the traditional PCIe 3.0 specification header fields are completely retained, such as Vendor ID (0x19E5, corresponding to custom vendors), Device ID (0x2001, corresponding to EX-DCC compatible devices), and the Command register (bit 0 controls device enable).

[0042] 3) Providing a foundation for expansion: The standard PCIe configuration space offers expansion capabilities (such as the Capabilities List mechanism). This design cleverly utilizes this mechanism, mounting subsequently defined innovative register sets (registers starting at offset 08H) as a set of custom, vendor-specific capabilities within the standard configuration space. This allows all innovative features to be organically integrated into the existing PCIe architecture ecosystem, and software can locate and access these new features through the standard PCIe capability discovery mechanism. For example, an 8-bit "EX-DCC flag" (0xAA indicating support for this specification) can be added to the end of the header area. During virtual CPU management initialization, this flag is scanned to quickly filter compliant devices.

[0043] The above-described embodiments utilize the industry-standard PCIe standard as a "universal language" and "passport," thereby ensuring device identifiability and interoperability, and laying a solid hardware and protocol foundation for the implementation and access of all subsequent innovative functions, maximizing the use of existing ecosystem resources.

[0044] 3. Standard PCIe capabilities and status register set.

[0045] This register group is the "basic function support layer" of EX-DCC encoding. Its core is "integrated standard capabilities + enhanced status linkage". It fully integrates the key link, device and power status registers in the PCIe standard, and accelerates the acquisition and interrupt linkage through FPGA, so as to provide real-time and accurate low-level status data.

[0046] Specifically, it includes:

[0047] 1) Fully integrated standard registers: Includes three main categories of core registers (128-bit):

[0048] Link Status / Control Register (LNK_STS / LNK_CTL): Records link width (0x08 = 8 channels), speed (0x04 = PCIe 4.0), and error count, used to determine the physical connection status of the device.

[0049] Device Status / Control Register (DEV_STS / DEV_CTL): Records internal device errors (such as ECC errors, checksum errors), and supports error mask configuration.

[0050] Power Management Register (PWR_MGMT): Records current power consumption (in W) and power consumption threshold, and supports switching between high performance and power saving modes (0x01 / 0x02).

[0051] 2) FPGA-accelerated state acquisition: The FPGA's built-in "state machine + DMA" module can, for example, acquire all register data every 10ms and write it directly to the virtual CPU's onboard cache via PCIe DMA, without occupying embedded CPU resources. Differential signal sampling technology is used during data acquisition to reduce data errors caused by electromagnetic interference (error rate ≤ 0.1%).

[0052] 3) Dynamic interrupt linkage mechanism: Different thresholds are set for different registers (e.g., GPU power consumption threshold 0x46 = 70W, NVMe power consumption threshold 0x1E = 30W; link error count threshold 0x05 = 5 times). When the data exceeds the threshold, the FPGA triggers a level interrupt, and the virtual CPU can respond within 1ms (e.g., when the power consumption exceeds the limit, the GPU frequency is reduced to 2.0GHz, and when the link error exceeds the limit, the link is restarted), thereby improving the efficiency of traditional polling response (≥10ms).

[0053] The above embodiments of the present invention achieve real-time acquisition and rapid response of status data through FPGA-accelerated acquisition and dynamic interruption, while providing underlying status input (such as link rate affecting communication delay assessment) for subsequent topology management and affinity calculation. This solves the problem that existing technologies rely heavily on CPU polling to read status, resulting in slow response and high resource consumption.

[0054] 4. Encryptable genetic device DNA identifier register

[0055] This design register is a "secure root of trust" encoded by EX-DCC. It breaks through the limitation of traditional Vendor / Device IDs that can only identify "type". Through "double encryption + genetic design", it gives each device a globally unique, verifiable hardware identity with "lineage" relationship, solving the problems of device counterfeiting and difficulty in tracing.

[0056] Specifically, it includes:

[0057] 1) 128-bit DNA structure design: DNA consists of four parts, ensuring uniqueness and heritability. Specifically, this may include:

[0058] -24-bit vendor root key (0x123456789ABC, stored in TPM 2.0 chip, unreadable).

[0059] -48-bit unique serial number for each device (laser-burned at the factory, globally unique, e.g., 0x100000010001~0x100000010200 corresponds to 200 devices).

[0060] -32-bit family identifier (shared by devices of the same batch / model, such as GPU family identifier 0x00010000, reflecting "lineage").

[0061] -24-bit CRC-32 checksum (calculated from the first 104 bits to ensure data integrity).

[0062] 2) Dual Encryption and Authentication: Employs a dual mechanism of "RSA-2048 asymmetric encryption + AES-256-GCM symmetric encryption".

[0063] - When the device leaves the factory, the DNA is signed using the manufacturer's RSA private key.

[0064] When a device connects to the cluster, the virtual CPU sends a "DNA verification request" → the device encrypts the DNA using an AES key → the virtual CPU decrypts the DNA using the RSA public key in the TPM and verifies the CRC → connection is only allowed after successful verification. Experiments have verified that a single verification takes ≤200μs, and the spoofing detection rate is 100%.

[0065] 3) Genetic Application: The virtual CPU identifies devices in the same batch (such as GPU family identifier 0x00010000) through the "family identifier hash matching algorithm", automatically assigns them the same initial communication affinity weight (0xFFFF), and enables "cluster collaborative optimization" (such as when multi-GPU parallel computing, computing groups are quickly formed based on family identifiers, reducing data synchronization latency by 30%).

[0066] The technical solution in this embodiment ensures identity credibility through double-encrypted DNA and identifies the "lineage" of devices through family identifiers, thereby distinguishing devices of the same model, providing accurate traceability, and providing technical support for accurate operation and maintenance (fault warning for the same batch) and collaborative computing.

[0067] 5. Device Resource and Topology Management Register Set

[0068] This register group is the "resource and location hub" of EX-DCC encoding. It breaks through the limitations of the traditional PCIe static BAR (base address register) and solves the problems of high resource demand and difficult location of complex topologies for high-performance devices through "enhanced resource declaration + dynamic topology awareness", providing support for resource scheduling of large-scale clusters.

[0069] The embodiments of this technical solution include:

[0070] 1) Enhanced resource declaration mechanism: Breaking through the traditional 32-bit BAR 4GB addressing limit, it adopts a "base size + scaling factor" design, including:

[0071] SELF_BARn_BASE_ADDR (32-bit): The physical base address allocated by the storage system (e.g., 0x80000000);

[0072] SELF_BARn_SIZE (16-bit): Base space size (e.g., 0x1000 = 4KB);

[0073] SELF_BAR_SIZE_SCALE (8 bits): scaling factor (e.g., 0x0A = 1024), actual address space = 4KB × 1024 = 4GB, supporting the large memory mapping requirements of devices such as GPUs;

[0074] SELF_BAR_VALID (8-bit bitmap): 1 bit corresponds to 1 BAR. bit0=1 indicates that BAR0 is enabled. The software can determine the BAR status without reading back to detect, which can improve the configuration efficiency by 50%.

[0075] 2) Dynamic Topology Awareness and Verification: Constructing a three-dimensional coordinate system of "domain-hierarchy-ordination", including:

[0076] CUR_TIER_IN_DOMAIN (8 bits): Device domain (0x02 = second domain).

[0077] SELF_NUM_IN_TIER (8 bits): Domain level number (0x05 = fifth level).

[0078] The virtual CPU updates the coordinates hourly using a "topology scan algorithm" (traversing the downstream ports of the PCIe switch) and compares them with the topology snapshot in the onboard ZFS file system. Simultaneously, by analyzing the difference between CUR_TIER_DEV_RECOG_CNT (number of physical devices) and CUR_TIER_SYS_RECOG_CNT (system enumeration count) (e.g., difference = 1), unenumerated devices can be located within 50ms, thereby improving fault location efficiency.

[0079] 3) Dynamic resource reconfiguration: Supports real-time adjustment of BAR in FPGA partial reprogramming scenarios: When the memory space needs to be expanded after FPGA reprogramming, the virtual CPU updates BAR_SIZE and SCALE through "SM4 encrypted instructions" without restarting the device. The configuration time is ≤100ms, which meets the needs of dynamic computing power expansion.

[0080] The above-described embodiments meet large resource requirements through scalable resource declarations, achieve rapid topology positioning through three-dimensional coordinates and snapshot verification, support dynamic reconfiguration, and adapt to flexible resource scheduling for large-scale clusters.

[0081] 6. Equipment historical status and reliability management register group

[0082] This register group is the "health management hub" of EX-DCC encoding. The solution is designed with the concept of "recording history + benchmark comparison + fault recovery". By integrating error records, performance benchmarks and historical configuration backups, it solves the problems of traditional technologies that can only view the current status, are difficult to predict faults and have slow recovery, and realizes the reliability management of the entire life cycle of the device.

[0083] The technical solution of this embodiment includes:

[0084] 1) Error logging and trend analysis: A "circular buffer + weighted counting" design is used.

[0085] SELF_ERR_CREDIT (8 bits): Error counter, different error types are assigned different weights (ECC error weight 0.6, communication timeout weight 0.4), the weighted threshold is set to 0x0A (10 points), and an alarm is triggered when the threshold is exceeded.

[0086] HIST_ERR_TYPE (64-bit): 16 bits / each stores 4 recent error types (0x01 = ECC error, 0x02 = communication timeout). New errors overwrite the oldest records and support root cause analysis (e.g., 3 consecutive ECC errors are determined to be hardware failure).

[0087] The virtual CPU runs a "weighted moving average (WMA) algorithm" to calculate a reliability score (0-100) based on the most recent 128 error records. A score ≤60 is marked as "high risk", and the early warning accuracy rate is ≥95%.

[0088] 2) Performance benchmark and loss quantification:

[0089] LINK_LATENCY (32-bit): Ideal communication latency for storage devices (e.g., 200ns for GPU, 150ns for NVMe), updated monthly via a "benchmark calibration algorithm" (averaging three tests under no-load conditions) to ensure benchmark accuracy.

[0090] During runtime, the actual latency is compared with the benchmark to quantify the performance loss (e.g., actual latency of 300ns = 50% loss), providing a basis for performance optimization (e.g., cleaning up link interference when the loss exceeds 30%).

[0091] 3) Backup and restore of historical configurations:

[0092] Back up topology configurations such as LAST_SELF_TIER (level) and LAST_MASTER_TIER (upper level) and store them in onboard non-volatile memory (NVM).

[0093] When the current configuration error count is ≥3 (such as hierarchical anomaly), the backup configuration is automatically loaded, and the recovery time is ≤200ms, which is 99% more efficient than traditional restart recovery (≥5min).

[0094] This technical solution achieves predictive maintenance through error trend analysis, quantifies performance loss through benchmark comparison, and enables rapid recovery through backup configuration, significantly improving equipment reliability and maintainability.

[0095] 7. Device Communication Affinity and Behavior Measurement Register Set

[0096] This group serves as the "intelligent collaboration hub" of the EX-DCC encoding. Its core breakthrough lies in overcoming the limitations of traditional static communication management. Through "quantified behavior + machine learning optimization," it achieves dynamic perception and optimization of device communication relationships, solving the problems of high communication latency and resource waste in large-scale clusters. The specific implementation method is as follows:

[0097] 1) Quantitative measurement of communication behavior: A "multi-dimensional indicator + sliding window" design is adopted.

[0098] SELF_VISIT_WEIGHT / OTHER_VISIT_WEIGHT (16 bits): Quantizes the priority of the device "initiating access" and "being accessed" respectively (e.g., GPU access to NVMe weight 0xFFFF, access to network card weight 0x8000).

[0099] SELF_VISIT_QUTIY / OTHER_VISIT_QUTIY (32-bit): 1-minute sliding window (50% overlap, reducing statistical error) statistics on traffic (in KB), such as GPU-NVMe traffic of 100,000 KB.

[0100] VISIT_WEIGHT_SCALE (8 bits): Weight scaling factor (e.g., 0x02 = 2 times), adapting to different business priority requirements (e.g., increasing GPU weights in AI training scenarios).

[0101] 2) GNN+Q-Learning Intelligent Optimization:

[0102] Every 5 minutes, the virtual CPU constructs a communication affinity graph (node ​​= device, edge weight = affinity score) using a graph neural network (GNN) based on weights, traffic, and topological distance (calculated in three-dimensional coordinates).

[0103] Run the Q-Learning algorithm to optimize affinity: With the goal of "lowest cluster communication latency", dynamically adjust the weight (e.g., if the communication volume is ≥100MB / min for 5 consecutive minutes, the weight increases by 10%), and adjust the step size ≤5% to avoid fluctuations;

[0104] By allocating high-affinity devices (such as GPUs and NVMe) to the same PCIe switch based on the graph, communication latency can be reduced from 500ns to 300ns, and the overall cluster transmission efficiency can be improved by 40%.

[0105] 3) Traffic load balancing linkage: Affinity data is synchronized to the traffic management module of the PCIe Switch, and more bandwidth is allocated among high affinity devices (such as 40%), thereby avoiding link congestion and effectively improving bandwidth utilization.

[0106] Compared to existing technologies that statically configure communication weights without intelligent optimization, the embodiments of this invention quantify communication behavior in multiple dimensions, use GNN+Q-Learning to achieve dynamic affinity optimization, link traffic management, significantly reduce communication latency, and improve cluster collaboration efficiency.

[0107] In summary, this solution achieves a comprehensive breakthrough in traditional PCIe management technology through the collaborative design of the seven register groups of EX-DCC. Experimental results show that: in terms of compatibility, the underlying layer fully complies with the PCIe standard, with a 100% device enumeration success rate, and can be deployed without modifying the system kernel. In terms of security, the encrypted DNA combined with the TPM chip provides 100% protection against device spoofing, and fault tracing accuracy reaches the single-device level. In terms of management efficiency, topology location time is reduced from 30 minutes to 50ms, fault warning accuracy is ≥95%, and device recovery time is reduced from 5 minutes to 200ms. In terms of performance, communication latency is reduced by 40%, GPU computing power utilization is increased by 20%, and bandwidth utilization is increased by 25%. Overall, it can support the stable operation of 200+ GPU / NVMe devices, fully meeting the secure, efficient, and intelligent management requirements of large-scale computing clusters.

[0108] Furthermore, in one embodiment, the PCIe card is logically recognized by the main CPU as a standard PCIe peripheral. The PCIe card (the card housing the virtual CPU management terminal) enables the main CPU to recognize it as a "normal peripheral" by precisely matching the PCIe standard device attributes.

[0109] 1. Device type definition: In the header area of ​​the standard PCIe configuration space, set the "Class Code" to 0x0108 (standard storage controller class). This is a device type natively supported by the main CPU (such as Intel Xeon series) and can be recognized without additional drivers.

[0110] 2. Core Identifier Configuration: The Vendor ID uses 0x19E5 (custom compliant vendor code), and the Device ID is set to 0x2002 (standard storage controller ID), matching the device list in the main CPU driver library to ensure that it is classified as "identifiable storage peripheral" during enumeration.

[0111] 3. Basic Capability Declaration: In the standard PCIe capability register, only the "necessary capabilities of the storage controller" (such as SATA protocol compatibility and NVMe protocol support) are exposed, while core functions such as FPGA acceleration and EX-DCC management are hidden to avoid the main CPU triggering alarms for unknown devices.

[0112] Furthermore, to prevent the main CPU from interfering with the board's encoding management functions, an isolation mechanism for identification and management is implemented through "hardware address filtering + function permission division":

[0113] 1. Address Space Partitioning: The board's address space is divided into a "Main CPU Visible Area" (containing only the standard memory controller configuration space, address 0x0000-0x0FFF) and a "Management Dedicated Area" (EX-DCC register, policy algorithm memory, address 0x1000-0xFFFF). The FPGA has a built-in address filter, and the main CPU automatically returns an "invalid address" when accessing the "Management Dedicated Area".

[0114] 2. Driver interaction limitation: The main CPU can only read and write to the board's 4TB NVMe SSD (such as reading management log backups) through the native storage driver, and cannot access the processes and registers of the encoding management system, ensuring that the management function runs independently.

[0115] 3. Enumeration process adaptation: When the main CPU starts up, the board simulates the enumeration response of the standard storage controller (such as sending the configuration space read / write ACK signal). The enumeration time is ≤10ms, which is consistent with traditional storage peripherals and does not affect the startup efficiency of the main CPU.

[0116] Through the standard peripheral logical profile and functional isolation of the above embodiments of the present invention, the following can be achieved: First, the main CPU does not need to install any custom drivers, and the native recognition rate is 100%, avoiding the stability risks caused by modifying the operating system kernel; Second, the board's encoding management function is completely hidden, and the main CPU can only perform basic storage interaction, which not only ensures management independence, but also integrates into the cluster as an "ordinary peripheral", adapting to various mainstream server hardware and significantly reducing the deployment difficulty of large-scale clusters.

[0117] Furthermore, in one embodiment, the PCIe board encodes standard PCIe peripherals according to the preset standard specification and writes the encoding into the onboard memory. This achieves standardization, encryption, and persistence of peripheral encoding through a "code generation + hierarchical storage writing mechanism based on the EX-DCC specification," solving the problems of traditional PCIe devices lacking unified encoding and easily lost identifiers. Specific details are as follows:

[0118] I. Encoding Generation Process Based on EX-DCC Specification

[0119] The PCIe board (the board where the virtual CPU management terminal is located) uses EX-DCC as the default standard and generates a unique peripheral code in three steps:

[0120] 1. Information Acquisition: The FPGA uses PCIe DMA technology to read the core information of the seven registers of the peripheral EX-DCC, including a 64-bit encrypted DNA identifier (device unique serial number), 16-bit topological coordinates (domain + hierarchical number), and 8-bit device type code (0x01 = GPU / 0x02 = NVMe). The acquisition time is ≤50μs / unit, ensuring data real-time performance.

[0121] 2. Encoding and Encryption: The embedded CPU combines the collected information into a 128-bit code according to the EX-DCC rule (field allocation: DNA identifier 64 bits + topology coordinates 16 bits + device type 8 bits + CRC32 check 32 bits + reserved 8 bits), and then encrypts it through the onboard hardware encryption engine (supporting SM4 algorithm) to generate an "encrypted encoded string" to prevent the code from being tampered with.

[0122] 3. Encoding, Storage, and Verification: The encrypted encoding string is written to the onboard 128MB NVM (non-volatile memory) via the SPI interface. The storage address is hash-mapped according to the device topology coordinates (e.g., devices in domain 0x02 and tier 0x03 are stored at address 0x203000), and power-loss retention is supported. Simultaneously, an encoding digest (SHA-256) is generated and stored in the NVMe log layer for subsequent integrity verification.

[0123] II. Encoding and writing to onboard memory can employ a "three-level storage collaboration" mechanism to adapt to different access requirements:

[0124] 1. Cache storage: The encrypted code that passes the verification is first written to 8GB DDR5 cache (address range 0x90010000-0x9001FFFF) to respond to the real-time query of the management system (access latency ≤100ns) and meet the requirements of high-frequency encoding calls.

[0125] 2. Hot data persistence: Every 5 minutes, the newly added / updated code in DDR5 is synchronized to 1TB SLC NAND (100,000 erase / write cycles) through the "incremental synchronization algorithm". Redundancy is built using RAID 5 array technology to prevent code loss.

[0126] 3. Cold data archiving: Every day at 2:00 AM, the encoding of SLC NAND exceeding 7 days is compressed (LZ4 algorithm, compression ratio 2:1) and archived to 8TB QLC NAND, retaining 30 days of archived records and supporting historical encoding traceability.

[0127] Through the EX-DCC standardized encoding and hierarchical storage writing of the above-described embodiments, the following technical effects are achieved: First, the 128-bit encrypted encoding ensures that each peripheral device identifier is unique and tamper-proof, solving the problem of dynamic address changes in traditional BDF; second, the three-level storage collaboration reduces the encoding access latency to within 100ns, and the archived retention is 30 days, meeting the needs of real-time management and traceability; third, the entire process of generation and writing is automated, requiring no manual intervention, with a 100% encoding generation success rate, adapting to large-scale cluster management of 200+ devices, improving the efficiency and security of encoding management.

[0128] Figure 3A structural diagram of a three-dimensional topology self-sensing management system 300 based on a PCIe device according to an embodiment of the present invention is shown.

[0129] like Figure 3 As shown, the three-dimensional topology self-sensing management system 300 includes an enhanced resource declaration mechanism 301 and / or dynamic topology sensing and verification 302.

[0130] The device resource and topology management register group, where the 3D topology self-aware management system resides, is the "resource and location hub" coded by EX-DCC, improving from basic compatibility to advanced resource management and intelligent topology awareness. By constructing an integrated hardware abstraction layer, devices can not only declare their resource requirements to the system but also accurately perceive their logical coordinates and local environmental states in complex interconnected topologies, thus providing a fine-grained resource control and collaborative computing foundation for heterogeneous computing platforms. It breaks through the limitations of traditional PCIe static BARs (base address registers) by using "enhanced resource declaration + dynamic topology awareness," surpassing the static resource configuration mode in the PCIe standard. By introducing a dynamic, scalable, and context-sensitive enhanced management scheme, it can solve the problems of high resource requirements and difficult location in complex topologies for high-performance devices, providing support for resource scheduling in large-scale clusters.

[0131] In one embodiment, the enhanced resource declaration mechanism achieves an innovative breakthrough over the 4GB address space limitation of the traditional 32-bit base address register (BAR) by employing a collaborative design architecture of "base size + scaling factor". This architecture mainly consists of three core parts: the base address register set (SELF_BARn_BASE_ADDR), the scalable space size register set (SELF_BARn_SIZE and SELF_BAR_SIZE_SCALE), and the resource validity bitmap (SELF_BAR_VALID). These components together construct a highly flexible and scalable address mapping mechanism, endowing high-performance computing devices with more powerful resource configuration capabilities.

[0132] Furthermore, in one implementation scenario, a cache-optimized resource prefetching strategy can be introduced on top of the existing enhanced resource declaration mechanism. By conducting in-depth analysis of historical resource access patterns and combining this with the current system's operating status, resources that may be accessed in the future can be predicted, and these resources can be pre-loaded into the cache.

[0133] The above-mentioned cache-optimized resource prefetching strategies include:

[0134] 1) Resource access predictor: Using machine learning algorithms, such as recurrent neural networks (RNNs), it analyzes and models historical resource access records to predict the resource addresses and types that may be accessed in the near future.

[0135] 2) Cache Manager: Based on the output of the resource access predictor, the predicted resources are prefetched from main memory to the cache. Simultaneously, intelligent cache replacement algorithms, such as an improved version of the Least Recently Used (LRU) algorithm, are employed to ensure efficient utilization of cache space.

[0136] The above mechanisms can effectively reduce resource access latency and improve system response speed and overall performance. The performance improvement is particularly significant for applications with regular resource access patterns.

[0137] Furthermore, in one implementation scenario, a multi-dimensional resource allocation and scheduling mechanism can be adopted on top of the existing enhanced resource declaration mechanism. By comprehensively considering multiple dimensions of resources, such as bandwidth, computing power, and storage capacity, based on the existing address mapping mechanism, resource allocation and scheduling are performed. Resource allocation and scheduling operations are dynamically implemented according to the specific needs of the application and the real-time status of system resources.

[0138] The aforementioned multi-dimensional resource allocation and scheduling mechanisms include:

[0139] 1) Resource multidimensional descriptor: Define a descriptor for each resource that contains information in multiple dimensions, covering various attributes of the resource, such as bandwidth, computing power, storage capacity, access latency, etc.

[0140] 2) Resource Scheduler: Based on the application's resource requirements and multi-dimensional resource descriptors, it uses multi-objective optimization algorithms, such as genetic algorithms, to allocate and schedule resources. Simultaneously, it monitors system resource usage in real time and dynamically adjusts the resource allocation scheme.

[0141] The aforementioned multi-dimensional resource allocation and scheduling mechanism improves resource utilization and avoids resource waste and bottlenecks. It better meets the diverse resource needs of different types of applications, enhancing the overall performance and flexibility of the system.

[0142] Furthermore, in one implementation scenario, an adaptive resource isolation and sharing mechanism can be adopted on top of the existing enhanced resource declaration mechanism. This mechanism enables adaptive resource isolation and sharing in multi-user or multi-task environments. The degree of resource isolation and sharing strategy are dynamically adjusted based on the application's security level, performance requirements, and resource usage. Specifically, this includes:

[0143] 1) Resource Isolation Controller: Set different resource isolation levels according to the security level and performance requirements of the application, such as hardware-level isolation, virtualization-level isolation, etc.

[0144] 2) Resource Sharing Coordinator: Monitors system resource usage in real time. When an application's resource requirements are low, it dynamically shares its idle resources with other applications that need them. Simultaneously, it ensures security and stability during the resource sharing process.

[0145] The adaptive resource isolation and sharing mechanism described above improves resource sharing efficiency and reduces system operating costs while ensuring system security and stability. It is particularly suitable for multi-user, multi-tasking application environments such as cloud computing and data centers.

[0146] First, the Base Address Registers (BARs) are crucial hardware configuration units in PCIe devices. They are responsible for receiving the physical memory base address, which is uniformly allocated and set by the system firmware (such as BIOS or UEFI) or the operating system. This mechanism establishes a precise and unique mapping starting point in the system's global address space for each device's functional area (such as device control registers, data buffers, or status areas), enabling the host CPU to communicate directly with the device through memory access instructions.

[0147] In one embodiment scenario, the implementation of device functional area mapping includes:

[0148] (I) Detailed classification and function of functional areas

[0149] PCIe devices include functional areas such as device control registers, data buffers, and status areas. Device control registers are used to control various operating modes, working states, and parameter settings of the device. For example, in a storage device, the control register can set parameters such as read / write mode and data transfer rate. The data buffer is used for temporary storage of data transferred between the device and the system; it can balance differences in data transfer rates and improve data transfer efficiency. The status area reflects the current operating status of the device, such as whether it is busy or whether an error has occurred. The host CPU can understand the device's operating status by reading the information in the status area.

[0150] (II) Specific Implementation and Guarantee Mechanism of Mapping

[0151] By working in conjunction with the address decoding logic of the PCIe bus, Memory Access Arrays (BARs) precisely map the internal functional areas of a device to system memory addresses. This process involves complex address calculation and matching algorithms. First, the BARs calculate the specific address of each functional area in the system's global address space based on their stored base address information and the offset of the device's functional area. Then, the PCIe bus's address decoding logic decodes the memory access instructions issued by the host CPU and matches the address in the instruction with the address mapped by the BARs.

[0152] To ensure the accuracy and uniqueness of the mapping, the system employs multiple safeguards. For example, during device initialization, the mapping relationships of BARs are verified and calibrated to ensure that the address mappings for each functional area are correct. Simultaneously, the system regularly checks and maintains the mapping relationships to promptly identify and correct any potential mapping errors.

[0153] Furthermore, the efficient implementation of communication between the host CPU and the device also includes:

[0154] (I) Processing flow of memory access instructions

[0155] When the host CPU needs to interact with a PCIe device, it sends read / write instructions to the corresponding memory address, just like accessing regular memory. These instructions first undergo address translation by the CPU's Memory Management Unit (MMU), converting the virtual address into a physical address. Then, the instructions are transmitted to the PCIe bus via the system bus, where the PCIe bus routes the instructions to the appropriate PCIe device based on the address decoding logic.

[0156] After receiving an instruction, the BARs will accurately transmit the instruction to the corresponding functional area of ​​the device according to the pre-established mapping relationship. For example, if the CPU sends a write instruction, the BARs will write the data in the instruction to the device's data buffer; if it is a read instruction, the BARs will read the data from the device's functional area and return the data to the host CPU.

[0157] (II) Error Detection and Correction Mechanism

[0158] To ensure reliable communication, BARs possess robust error detection and correction mechanisms. During data transmission, BARs verify the transmitted data, employing algorithms such as parity checking and cyclic redundancy check (CRC) to detect errors. If an error is detected, the BARs send an error signal to the host CPU and attempt error correction.

[0159] For some correctable errors, BARs will automatically retransmit data or perform error correction; for serious errors, the system will take corresponding measures, such as interrupting data transmission or notifying the operating system to handle the error.

[0160] (III) Implementation and Optimization of Burst Transmission Mode

[0161] To improve communication efficiency, BARs support burst transfer mode. In burst transfer mode, the host CPU can transfer multiple data blocks consecutively in a single operation, reducing communication overhead and latency. When the CPU needs to perform continuous data transfer, it sends a burst transfer instruction to the BARs, specifying the number of data blocks to be transferred and the transfer direction.

[0162] BARs automatically manage the data transmission process according to instructions. During transmission, BARs employ pipeline technology, simultaneously reading, transmitting, and processing data, thus improving the parallelism of data transmission. Furthermore, the system optimizes bandwidth for burst transmissions, rationally allocating bandwidth resources based on device performance and system load to ensure efficient burst transmission.

[0163] It should be noted that this design is not only a key prerequisite for achieving efficient and reliable Direct Memory Access (DMA) transfers, but also forms an important foundation for low-latency, high-bandwidth data interaction between the host and PCIe devices. By strictly adhering to and maintaining complete compatibility with the PCIe standard configuration space, the design of this base address register set ensures that devices from different vendors can be correctly identified, resource-enumerated, and initialized in various hardware platforms and operating system environments. Therefore, the compatibility and stability of this design provide a solid guarantee for reliable startup, long-term stable operation, and efficient device management of the computer system.

[0164] To enable those skilled in the art to more clearly understand the technical solution of the present invention, the following detailed description of the base address register group is provided through some embodiments:

[0165] 1) Example of BAR configuration for a graphics card: A certain NVIDIA RTX 4090 graphics card contains 6 base address registers (BAR0-BAR5), of which BAR0 (32-bit prefetchable) maps 16MB of video memory for register access, and BAR2 (64-bit prefetchable) maps 24GB of GDDR6 video memory for data throughput. During system startup, the UEFI firmware allocates physical addresses 0xA00000000000-0xA00058000000 to BAR2 through an enumeration process, enabling the CPU to directly access the graphics card's video memory via memory instructions, achieving 800GB / s DMA data transfer.

[0166] 2) Storage Controller BAR Conflict Resolution: When two NVMe SSDs are inserted into the server simultaneously, if the initial firmware allocation causes BAR address overlap, the operating system's PCIe driver will trigger a remapping mechanism. This mechanism will adjust the BAR1 of the second SSD from 0x20000000 to 0x30000000 and update the address resources of the device path \_SB.PCI0.RP21.SSD2 through the ACPI table. This ensures that both devices can be addressed normally by the CPU, avoiding I / O access conflicts.

[0167] 3) Virtualization Environment BAR Isolation Scheme: In a KVM virtualization scenario, QEMU uses IOMMU technology to map physical BAR addresses 0x50000000-0x60000000 to the virtual machine's virtual address 0xC0000000, while configuring the EPT memory table to achieve virtual-physical address translation. When the virtual machine performs an MMIO write operation 0xC0000010, the hardware MMU automatically translates it into an access to the physical BAR space, ensuring device passthrough performance while achieving address space isolation between different virtual machines.

[0168] Secondly, in the expandable address space size register group, the SELF_BARn_SIZE register is mainly used to set the number of bytes of the base address space size, which is usually configured in the form of an integer power of 2. The SELF_BAR_SIZE_SCALE register, on the other hand, acts as a multiplier factor, and together with the base space size, it enables dynamic adjustment of the actual accessible address window size through multiplication.

[0169] In modern computer systems, high-performance peripherals such as graphics processing units (GPUs), field-programmable gate arrays (FPGAs), and smart network interface cards (NICs) are increasingly widely used. These peripherals often require mapping extremely large amounts of onboard memory or complex register spaces. Traditional 32-bit base address registers (BARs) can only support a fixed addressing range of up to 4GB, which is insufficient to meet the needs of these peripherals. The composite address space configuration mechanism provided in this invention effectively overcomes the limitation of traditional 32-bit base address registers (BARs) supporting only a fixed addressing range of up to 4GB. It provides a large address space allocation scheme that is both highly flexible and backward compatible for high-performance peripherals such as GPUs, FPGAs, and smart NICs—especially for applications that require mapping massive amounts of onboard memory or complex register spaces. Compared to alternative implementations such as directly using 64-bit BARs, which may bring hardware compatibility challenges and software adaptation complexity, this mechanism significantly expands the physical address coverage capability while also having the advantages of clear structure and simple configuration, thereby greatly reducing the engineering complexity that may occur during the overall system design and integration process.

[0170] Specifically, the SELF_BARn_SIZE register is a fundamental component of the expandable address space size register set, and its core function is to define the size of the base address space. This register is configured in powers of 2 because in computer systems, memory allocation and management are typically performed in units of powers of 2, ensuring address space alignment and improving memory access efficiency. For example, it can be configured to 1KB (2^250 ... 10 bytes), 4MB (2 22 Common aligned capacity units such as bytes.

[0171] Specifically, when the system allocates address space, the hardware reads the value of the SELF_BARn_SIZE register and uses it as a basis to determine the initial address space range. This initial range provides a baseline for subsequent address space expansion.

[0172] Furthermore, the SELF_BAR_SIZE_SCALE register plays a crucial role as a multiplier factor in the entire technical solution. It dynamically adjusts the size of the actual accessible address window by multiplying it with the base address space size defined by the SELF_BARn_SIZE register.

[0173] Assuming the SELF_BARn_SIZE register is configured to 4MB (2 22 (bytes), while the value of the SELF_BAR_SIZE_SCALE register is 8 (2 bytes). 3 If so, the actual accessible address window size will become 4MB × 8 = 32MB (2 25 (bytes). This multiplication method makes the expansion of the address space more flexible, allowing for precise configuration according to the specific needs of different peripherals.

[0174] Therefore, when dealing with high-performance peripherals that require mapping very large onboard memory or complex register spaces, the traditional 32-bit base address register (BAR) is limited to a fixed addressing range of only 4GB due to its bit limitation. This limitation is particularly pronounced.

[0175] The Scalable Address Space Register (BAR) technology, through the collaborative operation of the two core registers mentioned above, successfully overcomes the limitation of traditional 32-bit base address registers (BARs), which can only support a fixed addressing range of up to 4GB due to their bit limitations, when dealing with high-performance peripherals that require mapping very large onboard memory or complex register spaces. It allows the system to flexibly adjust the address space size according to the actual needs of the peripherals, providing greater address space support for high-performance peripherals. This flexibility enables the system to better adapt to different types and sizes of peripherals, improving the overall system performance and compatibility.

[0176] Furthermore, to further enhance the address space's scalability, a multi-level expansion mechanism can be introduced. In addition to the existing SELF_BAR_SIZE and SELF_BAR_SIZE_SCALE registers, an extra set of extended registers can be added. For example, a two-level extended register set can be set up, including the SECOND_BAR_SIZE and SECOND_BAR_SIZE_SCALE registers.

[0177] When primary address space expansion (using the SELF_BAR_SIZE and SELF_BAR_SIZE_SCALE registers) is insufficient to meet the address space requirements of peripherals, the system can enable secondary address space expansion. Specifically, primary expansion first determines an intermediate address space range, and then, based on this intermediate range, secondary expansion is performed using the SECOND_BAR_SIZE and SECOND_BAR_SIZE_SCALE registers. This allows for larger-scale address space expansion, meeting the needs of high-performance peripherals in extreme cases.

[0178] Furthermore, a dynamic adaptive expansion mechanism can be introduced, enabling the system to automatically adjust the address space size based on the real-time needs of peripherals. The system can assess the address space requirements of peripherals in real time by monitoring their memory access patterns and data traffic.

[0179] For example, when the system detects a sudden increase in the memory access frequency of a peripheral device, and the existing address space may not be able to meet its subsequent data processing needs, the system will automatically adjust the value of the SELF_BAR_SIZE_SCALE register, or enable a multi-level expansion mechanism to dynamically increase the accessible address space of the peripheral device. When the memory access needs of the peripheral device decrease, the system can correspondingly shrink the address space to improve the utilization of system resources.

[0180] Furthermore, to better manage the expanded address space, a segmented address space management approach can be adopted. This involves dividing the entire expanded address space into multiple segments, each with different attributes and uses.

[0181] For example, the address space can be divided into data segments, code segments, and register segments. Different segments can employ different access permissions and management strategies, improving system security and stability. Furthermore, during address space allocation, different segments can be assigned to different functional modules based on the specific needs of peripherals, enabling more granular resource management.

[0182] In summary, the embodiments provided by this invention offer significant advantages over alternatives such as directly using a 64-bit BAR. While directly using a 64-bit BAR can provide a larger addressing range, it introduces a series of hardware compatibility issues. Many existing computer systems and devices are still based on 32-bit architectures, and using a 64-bit BAR may cause these devices to malfunction. Furthermore, software adaptation complexity increases significantly, requiring substantial modifications and optimizations to existing operating systems, drivers, and other components.

[0183] Therefore, the scalable address space register group technology provided by the embodiments of the present invention significantly improves physical address coverage while possessing the advantages of clear structure and simple configuration. It can expand the address space through simple register configuration without changing the existing hardware architecture and software foundation, effectively reducing the complexity of system design and integration, and improving the overall performance and compatibility of the system.

[0184] To facilitate understanding by those skilled in the art, some specific implementation scenarios of scalable space-size register groups are given below:

[0185] 1) GPU Onboard Memory Expansion Scenario: A certain GPU model needs to map 32GB of onboard GDDR6 memory. By configuring SELF_BAR0_SIZE = 0x20 (representing 2^32 bytes, i.e., 4GB base space) and SELF_BAR_SIZE_SCALE = 8 (multiplier factor), an actual address space of 4GB × 8 = 32GB is achieved. During system initialization, the firmware automatically detects and configures these two registers, enabling the CPU to directly access the complete 32GB of video memory. This improves addressing capability by 8 times compared to the traditional 32-bit BAR scheme, meeting the large-scale data throughput requirements in AI training.

[0186] 2) FPGA Dynamic Configuration Case: An adaptive computing accelerator card (FPGA) supports dynamic hardware logic reconfiguration, requiring its configuration space to be expanded from the initial 8MB to 128MB. By writing SELF_BAR1_SIZE = 0x10 (2^16 = 64KB base space) and SELF_BAR_SIZE_SCALE = 2048 (multiplier factor), 64KB × 2048 = 128MB is calculated. After the FPGA completes the logic reconfiguration, the firmware updates these two register values ​​via the I2C interface, completing the address space expansion without restarting the system, achieving a 95% improvement in configuration efficiency.

[0187] 3) Smart NIC Multi-Buffer Management: A certain 25G smart NIC needs to map four independent buffers simultaneously (8GB each). By configuring four sets of SELF_BARn_SIZE and SELF_BAR_SIZE_SCALE registers: BAR0 (0x20,2) = 8GB, BAR1 (0x20,2) = 8GB, BAR2 (0x20,2) = 8GB, and BAR3 (0x20,2) = 8GB, the total address space reaches 32GB. The operating system dynamically adjusts the size of each buffer according to network traffic. When the receive buffer is full, it automatically adjusts the SCALE value of BAR0 from 2 to 4, quickly expanding the buffer to 16GB, avoiding packet loss and increasing network throughput by 200%.

[0188] Finally, the resource availability bitmap provides an intuitive and accurate mechanism for indicating resource availability by using a one-bit-to-one-base-address-register (BAR) mapping. Each bit in the bitmap directly reflects the current configuration status of the corresponding BAR: a bit of 1 indicates that the corresponding BAR space has been successfully enabled; a bit of 0 indicates that the BAR is disabled. The software can quickly obtain the entire bitmap information with a single simple read operation, thus efficiently determining whether a specific BAR is available. This method completely avoids the write-readback verification operation required in traditional resource configuration processes, significantly simplifying the configuration steps for PCIe device resources during initialization and significantly improving the overall efficiency of system initialization.

[0189] Furthermore, this design offers long-term architectural advantages, providing underlying hardware support for future dynamic reconfiguration of runtime resources. For example, in scenarios requiring partial FPGA reprogramming or dynamic adjustment of device functions, the system can reconstruct and remap BAR resources without a complete reset, further expanding the system's potential in terms of flexibility and scalability.

[0190] Furthermore, in one embodiment scenario, the resource validity bitmap can adopt a multi-level hierarchical bitmap architecture design. For example, a three-level indexed bitmap structure can be used to achieve fine-grained management of resource status. Specifically, the global control bit plane is 32 bits, which implements enable control at the base address register (BAR) group level, supporting batch configuration of 8 groups, each containing 4 BARs. The functional domain status bit plane is 128 bits, which is a sub-bitmap divided according to PCIe functional domains, with each functional domain covering 16 BAR status bits. The extended attribute bit plane is 256 bits, used to record extended attributes such as BAR access permissions (4 bits), address space type (2 bits), and mapping status (2 bits). Fast association lookup of the three-level bitmap is achieved through a bit plane cross-indexing algorithm (XOR hash mapping), keeping the lookup latency within 3 clock cycles.

[0191] Furthermore, in one embodiment scenario, the resource validity bitmap employs a dynamic resource status monitoring mechanism. This mechanism includes a real-time status sampler, a state transition cache, and conflict detection logic. The real-time status sampler integrates a 32-channel parallel comparator to asynchronously monitor the BAR configuration space at a frequency of 200MHz. The state transition cache uses a FIFO structure to record the 16 most recent BAR state transition events, which include a timestamp (16 bits), a BAR number (8 bits), and a status code (4 bits). The conflict detection logic implements cache consistency checks based on the MESI protocol, preventing multiple master devices from simultaneously operating on the same BAR resource. Simultaneously, a state confidence assessment mechanism is introduced, confirming the validity of the state through three consecutive samples, controlling the false positive rate to below 0.001%.

[0192] Furthermore, in one embodiment scenario, the resource validity bitmap also includes an advanced power management extension. This extension includes a power state map, a dynamic power optimizer, and a wake-up priority encoder. The power state map establishes the association between BAR states and PCIe power management states (D0-D3hot). The dynamic power optimizer automatically adjusts the power state of the corresponding PCIe port based on the access frequency of the BAR (statistics are collected every 10ms). The wake-up priority encoder sets wake-up priorities of 0-7 for different BARs, supporting fast recovery based on service priorities. By implementing an adaptive power adjustment algorithm, the power consumption of the PCIe subsystem can be reduced by approximately 18% while ensuring performance.

[0193] Furthermore, in one embodiment scenario, the resource validity bitmap may also include a security enhancement module. This module consists of an access control matrix, an abnormal behavior detector, and a security log register. The access control matrix implements a 32×32 permission control table based on the RBAC model, supporting fine-grained authorization for 8 user groups. The abnormal behavior detector has a built-in LSTM neural network inference engine (quantized to 8-bit weights) capable of monitoring abnormal BAR access patterns in real time. The security log register records 64 security event logs, including operation type (4 bits), user ID (8 bits), timestamp (20 bits), and result code (4 bits). It also supports log signature functionality using the SM2 national cryptographic algorithm to ensure that audit information cannot be tampered with.

[0194] Furthermore, in one embodiment scenario, the resource validity bitmap may also include high-speed interfaces and protocol extensions. Specifically, it may include an AXI4-Stream interface, a PCIe MSI-X extension, and a JTAG debug interface. The AXI4-Stream interface implements a 128-bit data width stream interface, supporting DMA batch transfers in BAR state. The PCIe MSI-X extension provides 32 MSI-X interrupt vectors, which can be configured as interrupt sources for events such as BAR state changes and collision detection. The JTAG debug interface supports forced bitmap state setting and single-step tracing in bit debug mode. An integrated SerDes interface hard core achieves a maximum state information transmission rate of 8Gbps to meet the centralized monitoring needs of large systems.

[0195] The above-described embodiments, in terms of FPGA dynamic reconfiguration support, enable seamless switching of BAR resources during partial reconfiguration, keeping reconfiguration interruption time within 1.2ms. Regarding virtualization environment optimization, it supports independent management of VF BAR states using SR-IOV technology, with each PF scalable to 64 VFs. For edge computing adaptation, it provides network sharing functionality for BAR resources and supports remote BAR mapping based on RDMA. In NFV environments, it can reduce virtual machine switching time by 40%; in FPGA acceleration scenarios, it can improve resource utilization by approximately 25%.

[0196] Furthermore, in one implementation scenario, the resource validity bitmap can also include diagnostic and self-healing mechanisms. These mechanisms include built-in BIT testing, error injection registers, and redundant bitmap design. Built-in BIT testing implements end-to-end self-testing, covering test items such as bitmap integrity, state machine logic, and interface timing. The error injection register supports the active injection of 16 predefined error modes for system robustness testing. The redundant bitmap design employs a TMR (Triple Modular Redundancy) structure for critical control bits, improving single-event fault tolerance by 100 times. Through a dynamic error recovery algorithm, 99.99% bitmap functionality availability can be achieved.

[0197] In one embodiment, the following specific configuration parameters can be used:

[0198] -SELF_BARn_BASE_ADDR (32-bit register): Used to store the actual physical base address dynamically allocated by the system. For example, this address may be set to 0x80000000 as the starting position of the physical memory region that this BAR can access.

[0199] -SELF_BARn_SIZE (16-bit register): Used to define the size of the base address space mapped by this BAR. For example, when the value of this register is 0x1000, it means that the base space size corresponding to this BAR is 4KB.

[0200] -SELF_BAR_SIZE_SCALE (8-bit register): Specifies the scaling factor for the address space, which is expressed as a power of 2, representing the expansion factor of the address range. For example, if the register is set to 0x0A, it means the scaling factor is 2 to the power of 10 (1024 times), making the actual accessible address space size of the BAR 4KB multiplied by 1024, or 4GB. This design effectively meets the urgent need for large-scale contiguous physical address mapping in devices such as GPUs.

[0201] -SELF_BAR_VALID (8-bit bitmap register): Each bit in this register maps to the enable / disabled state of a BAR, where bit 0 corresponds to BAR0, bit 1 to BAR1, and so on. For example, if bit 0 is 1, it means BAR0 is enabled. Software can quickly determine the availability of each BAR by directly reading this bitmap without performing redundant readback probing operations. Practical applications show that this mechanism significantly optimizes the resource allocation process, improving its efficiency by approximately 50%.

[0202] The following is a specific implementation of the resource validity bitmap and related registers:

[0203] 1. Multi-device initialization scenario: When the server starts, the virtual CPU management terminal reads the SELF_BAR_VALID register (0x05) of all PCIe devices to obtain bitmap 0b00000011, indicating that BAR0 and BAR1 are enabled. Based on this, the system quickly identifies the GPU (BAR0 maps to 32GB of video memory) and the NVMe controller (BAR1 maps to 4GB of register space), skipping the detection of disabled BARs (2-7), reducing initialization time by 40%.

[0204] 2. Dynamic Resource Adjustment Case: When a smart network interface card (NIC) needs to enable a new encryption acceleration engine, its firmware automatically sets bit 3 of SELF_BAR_VALID to 1 (the bitmap becomes 0b00001011) and notifies the system via an MSI interrupt. After reading the bitmap change, the virtual CPU management terminal immediately allocates 0x90000000 as the BAR3 base address, completing the functional expansion without a reboot, with a service interruption time ≤50ms.

[0205] 3. Fault Recovery Mechanism: If BAR2 of the FPGA is disabled due to an ECC error (bit 2 of SELF_BAR_VALID changes from 1 to 0), the system triggers a hot reset process: First, the configuration parameters of BAR0 / 1 are backed up, then the BAR2 address mapping is restored by writing SELF_BARn_BASE_ADDR and SIZE registers through the I2C interface, and finally bit 2 is set to enable (bitmap restored to 0b00000111). The entire recovery process takes ≤200ms, ensuring the continuity of industrial control scenarios.

[0206] Furthermore, in one embodiment, dynamic topology awareness and verification leverages self-localization and environmental awareness capabilities to achieve device collaboration and system-level management. By constructing a three-dimensional coordinate system of "domain-hierarchy-ordination," devices are logically assigned unique identifiers and clear hierarchical relationships within the topology. This coordinate system acts like a precise map, determining a unique position for each device within the system, facilitating unified management and scheduling. Specifically, it may include: a topology location register group (CUR_TIER_IN_DOMAIN, SELF_NUM_IN_TIER) and an environmental status register group (CUR_TIER_DEV_RECOG_CNT, CUR_TIER_SYS_RECOG_CNT, CUR_TIER_MAX_DEV).

[0207] First, the CUR_TIER_IN_DOMAIN register in the topology location register group serves as the core configuration unit, its main function being to accurately identify the domain number to which a device belongs. In complex multi-rooted or partitioned system architectures, this register provides crucial support for the rapid location of device logical subtrees, thus ensuring the efficiency of system resource management and scheduling in multi-domain environments. Simultaneously, the SELF_NUM_IN_TIER register, by assigning a unique sequence number to each device at its respective level, effectively achieves accurate differentiation and identification between devices at the same level, significantly reducing the probability of conflicts during device identification. The globally unique coordinate system (composed of domain number, hierarchy information, and device ordinal number) constructed through the collaborative work of these register groups lays the foundation for accurate allocation of system resources, efficient implementation of fault isolation mechanisms, and deep optimization of Non-Unified Memory Access (NUMA) architectures, thereby improving the reliability and performance of large-scale computing systems.

[0208] The detailed configuration instructions are as follows:

[0209] -CUR_TIER_IN_DOMAIN (8 bits): Used to identify the domain number to which the device belongs. For example, a value of 0x02 indicates that the device is located in the second domain.

[0210] -SELF_NUM_IN_TIER (8 bits): Used to indicate the sequence number of the device in the current hierarchy within its domain. For example, 0x05 means that the device is the fifth position in the hierarchy.

[0211] The following is a specific embodiment of the topology location register group:

[0212] 1. Multi-server domain isolation deployment: In a 4U rackmount server, dual-domain isolation (domain 0: compute node / domain 1: storage node) is achieved by configuring the CUR_TIER_IN_DOMAIN register. When a GPU card (SELF_NUM_IN_TIER = 0x03) is inserted into the PCIe slot of domain 0, its CUR_TIER_IN_DOMAIN value is automatically configured to 0x00, and the system allocates the video memory resources (32GB) to the compute domain memory pool accordingly. Meanwhile, the CUR_TIER_IN_DOMAIN of the NVMe RAID card (SELF_NUM_IN_TIER = 0x05) is set to 0x01, and its storage bandwidth (32GB / s) is directed to the storage domain. This achieves 100% resource isolation between domains, avoiding increased IO latency caused by compute tasks preempting storage bandwidth (actually reducing latency fluctuation by 40%).

[0213] 2. Blade Server Tiered Resource Scheduling: A blade server center contains 16 compute blades (Tier 1) and 4 switching blades (Tier 0). Each blade is configured with a unique ordinal number (0x00-0x0F) via the SELF_NUM_IN_TIER register. When blade 3 (ordinal number 0x03) experiences a memory failure, the virtual CPU management terminal reads its CUR_TIER_IN_DOMAIN = 0x01 (compute tier) and immediately triggers resource reconfiguration at the same tier: the tasks of the failed blade are migrated to the redundant blade with ordinal number 0x07. The migration process directly locates the target device using topological coordinates, reducing the time taken by 65% ​​compared to the traditional PCIe enumeration method (from 2.3s to 0.8s).

[0214] 3. NUMA Architecture Performance Optimization: In a dual-CPU server, the GPU card (domain 0, tier 2, ordinal 0x02) is identified by the system as a high-affinity device close to CPU0 based on the coordinate combination of CUR_TIER_IN_DOMAIN and SELF_NUM_IN_TIER (0x00_02_02). The virtual CPU management terminal then prioritizes allocating the weight parameters (8GB) of the AI ​​training task to the local memory of CPU0, reducing GPU access latency from 120ns to 75ns and improving model training iteration speed by 37%. When accessing CPU1 memory across domains, the system automatically triggers the QoS mechanism, dynamically adjusting PCIe bandwidth allocation based on coordinate information to ensure the priority of critical tasks.

[0215] Secondly, in one embodiment, the CUR_TIER_DEV_RECOG_CNT register in the environment status register group aims to accurately count and report the total number of physical devices actually identified in the current level from a hardware perspective, truly reflecting the "physical truth" of the system's underlying structure. CUR_TIER_SYS_RECOG_CNT, on the other hand, records the number of devices in the same level identified by the system software during the enumeration process. Furthermore, CUR_TIER_MAX_DEV is used to explicitly define the maximum number of devices that the current level can support on the physical hardware, i.e., the so-called "physical capacity" upper limit. Through this register group, the system can provide each device with a comprehensive view of its local operating environment, allowing it to clearly understand the status of its level. By comparing the number of physical devices with the number enumerated by the system in real time, the system can achieve rapid and consistent diagnostics and effective health assessment. The maximum device value further clarifies the theoretical limit of the environment in terms of device expansion, providing an important basis for resource planning and dynamic adjustment.

[0216] During system operation, the virtual CPU triggers a "topology scan algorithm" once per hour. This algorithm dynamically updates device coordinate information by traversing all downstream ports of the PCIeSwitch and compares it with pre-stored topology snapshots in the onboard ZFS file system for verification. Furthermore, by calculating the numerical difference between CUR_TIER_DEV_RECOG_CNT (number of physical devices) and CUR_TIER_SYS_RECOG_CNT (system enumeration count) in real time (e.g., when the difference is 1), the system can accurately locate devices that have not been correctly enumerated by the system within 50 milliseconds, thus significantly improving the efficiency of fault diagnosis and location.

[0217] The following is a specific implementation of the environment status register group:

[0218] 1. Blade Server Tier Device Consistency Verification: A blade server center's 3rd domain tier 2 (coordinates 0x03_02) is configured with 12 GPU compute nodes, and its CUR_TIER_MAX_DEV is set to 0x0C (12 nodes). During system operation:

[0219] 1) Physical identification: The FPGA scans the backplane via the I2C bus, and CUR_TIER_DEV_RECOG_CNT returns 0x0B in real time (11 units).

[0220] 2) System enumeration: OS PCIe driver enumeration result CUR_TIER_SYS_RECOG_CNT = 0x0A (10 units).

[0221] 3) Difference Diagnosis: The virtual CPU detects a physical-system quantity difference of 1 (0x0B-0x0A=0x01) and immediately triggers the localization mechanism: ① Reads the SELF_NUM_IN_TIER register of all devices in layer 2 and finds that the device with ordinal number 0x07 is unresponsive; ② Locates blade 7 using topology coordinates (0x03_02_07) and detects poor contact of its PCIe gold fingers; ③ Resets the blade power supply via SMBus. After 3 seconds, the device comes back online, and both CUR_TIER_DEV_RECOG_CNT and CUR_TIER_SYS_RECOG_CNT are updated to 0x0B. The fault recovery time is reduced by 98% compared to traditional manual troubleshooting.

[0222] 2. Domain Capacity Planning for Multi-Root Servers: A cloud computing platform conducted domain expansion tests on newly deployed dual-root servers.

[0223] 1) Physical capacity verification: Set CUR_TIER_MAX_DEV in the second field (0x02) to 0x10 (16 units). When 17 GPUs are actually connected, bit 15 of the configuration register of the 17th device (SELF_NUM_IN_TIER=0x10) is set (capacity over-limit flag).

[0224] 2) System response: The virtual CPU management terminal refuses to enumerate the over-limit device and records "Domain 0x02 Level 1 Capacity Over-Limit (17>16)" in the log.

[0225] 3) Optimization solution: Based on the hardware limit of CUR_TIER_MAX_DEV, the system automatically allocates the over-limited device to the free slot in the third domain (0x03). The migration is completed by modifying the value of its CUR_TIER_IN_DOMAIN register (updated from 0x02 to 0x03), ensuring that the resource utilization rate reaches 94%.

[0226] 3. Dynamic Expansion of Edge Computing Nodes: The initial configuration of a Layer 1 industrial edge computing gateway is CUR_TIER_MAX_DEV = 0x04 (4 devices). As the production line expands, it needs to connect 6 AI vision sensors.

[0227] 1) Hardware configuration adjustment: Reprogram CUR_TIER_MAX_DEV to 0x08 (8 units) via JTAG interface.

[0228] 2) Hot-swap verification: After the two newly connected sensors (ordinal numbers 0x05-0x06) come online, CUR_TIER_DEV_RECOG_CNT increments from 0x04 to 0x06, and the system enumeration count is updated synchronously, with the difference always being 0.

[0229] 3) Performance monitoring: Continuous monitoring for 72 hours showed that the average response latency of the expanded device was stable at 35ms (<50ms threshold), proving that the adjustment of CUR_TIER_MAX_DEV can effectively support dynamic expansion requirements.

[0230] Furthermore, in one embodiment, the system incorporates an adaptive learning algorithm in dynamic topology sensing and verification. This algorithm dynamically adjusts the topology sensing and verification strategy based on historical operating data of the devices and real-time topology changes. For example, when the system detects frequent changes in devices in a certain area, the algorithm automatically increases the sensing frequency of that area, improving the response speed to topology changes. Simultaneously, through the analysis and learning of a large amount of historical data, the algorithm can predict potential device failures or anomalies and take corresponding preventative and handling measures in advance.

[0231] Furthermore, in one embodiment, to address the storage and processing needs of large-scale device topology data, the system employs distributed storage and processing technology. Topology information and device status data are distributed and stored across multiple nodes, avoiding single points of failure and data bottlenecks.

[0232] Furthermore, in one embodiment, to enhance the device's environmental awareness, the system employs multi-sensor fusion technology. In addition to traditional positioning sensors, it integrates various types of sensors, such as temperature sensors, humidity sensors, and light sensors. By fusing data from different sensors, the system can acquire more comprehensive and accurate environmental information. For example, during topology sensing, combining data from the temperature sensor can determine whether the device is malfunctioning due to overheating, allowing for timely implementation of heat dissipation measures or adjustment of the device's operating status.

[0233] Furthermore, in one embodiment, the system establishes a robust fault diagnosis and fault tolerance mechanism, capable of real-time monitoring of device operating status and topology changes. When a device malfunction or topology anomaly is detected, the system quickly initiates a fault diagnosis program, analyzing device status data and topology information to locate the fault point and determine the cause. Simultaneously, the system possesses fault tolerance capabilities, automatically adjusting the topology and device coordination strategies to ensure normal system operation even when some devices malfunction. For example, when a device malfunctions, the system automatically distributes its tasks to other functioning devices, ensuring uninterrupted system operations.

[0234] The above embodiments dynamically adjust the topology sensing and verification strategy based on processed and historical data using adaptive learning and optimization algorithms. Utilizing the system's "domain-hierarchy-ordinal" three-dimensional coordinate system and register set, the topology structure is sensed and verified in real time. Simultaneously, combined with fault diagnosis and fault tolerance mechanisms, equipment faults and topology anomalies are detected and addressed promptly. This enables the dynamic topology sensing and verification system to operate efficiently and stably in complex and ever-changing environments, providing strong support for the reliable operation of the system.

[0235] Furthermore, in one embodiment, the 3D topology self-aware management system achieves deep collaboration and close cooperation through an enhanced resource declaration mechanism and / or a collaborative working mechanism for dynamic topology awareness and verification, specifically including the following three aspects:

[0236] Firstly, the intelligent resource allocation mechanism. Based on the domain number information of the device, the system can dynamically reallocate and adjust resources in real time, allocating the device's BAR space to adjacent or nearby memory nodes in the topology. This effectively reduces access latency, improves data transmission efficiency, and significantly optimizes the overall system's access performance and resource utilization.

[0237] In one implementation scenario, the AI ​​training cluster BAR spatial topology optimization of a certain intelligent computing center's second domain (0x02) includes 8 GPUs (ordinal numbers 0x01-0x08) and 4 NVMes (ordinal numbers 0x09-0x0C), with its memory nodes distributed on both sides of the motherboard's north and south bridges (topology coordinates 0xA000 / 0xB000). During system operation: 1) Domain resource mapping: The virtual CPU identifies the domain affiliation of all devices through the CUR_TIER_IN_DOMAIN register, and generates a "device-memory distance matrix" by combining PCIe topology scanning, showing that GPU0x01-0x04 is only 2 hop away from the Northbridge memory (0xA000), and GPU0x05-0x08 is 3 hop away from the Southbridge memory (0xB000); 2) Dynamic BAR allocation: According to the distance matrix, the BAR space of GPU0x01-0x04 (a total of 8GB) is mapped to the Northbridge memory node, and the NVMe device BAR (2GB) is allocated to the Southbridge memory; 3) Performance benefits: In the model training task, the latency of GPU accessing the local BAR space is reduced from 120ns to 75ns, the cross-node data transmission bandwidth is increased by 38%, the multi-card synchronization efficiency is increased by 22%, and the training epoch time is shortened by 15 minutes compared with the traditional fixed allocation scheme.

[0238] Secondly, advanced diagnostics and maintenance capabilities. When devices experience abnormal situations such as DMA errors, the system can quickly combine their domain number and device ordinal information to achieve precise fault location and tracing within seconds, greatly improving system maintainability. Furthermore, inconsistencies between the system monitoring parameters CUR_TIER_DEV_RECOG_CNT and CUR_TIER_SYS_RECOG_CNT can serve as important indicators for identifying early hardware failures or potential configuration errors, helping maintenance personnel intervene proactively and avoid more serious system problems.

[0239] The following is a specific example scenario:

[0240] A. DMA error location analysis within seconds

[0241] For example, in a financial server cluster, a DMA transfer error occurs on the GPU device with domain number 0x04 (ordinal number 0x09):

[0242] 1) Error Trigger: The device driver captures a PCIe DMA timeout interrupt (error code 0x1003) and immediately reads the device's CUR_TIER_IN_DOMAIN (returns 0x04) and SELF_NUM_IN_TIER (returns 0x09) registers;

[0243] 2) Topology location: The virtual CPU, combined with the domain-level coordinates (0x04_02_09), queries the physical slot mapping table of this location through the backplane management bus to locate the blade server's 4U 9th slot;

[0244] 3) Fault diagnosis: The system automatically reads the temperature sensor (85℃) and link status register (CRC error count = 23) of the slot and determines that the link instability is caused by overheating of the PCIe gold finger;

[0245] 4) Repair measures: Restart the power supply of the slot by sending a command through IPMI. The device will come back online after 30 seconds, the DMA error will be eliminated, and no manual intervention is required throughout the process. This reduces the time taken by 97% compared to the traditional diagnostic process.

[0246] B. Warning scenario for inconsistent device quantity

[0247] For example, a cloud computing data center has 16 SSDs configured in tier 3 of domain 2 (CUR_TIER_MAX_DEV = 0x10):

[0248] 1) Difference detection: Hourly topology scans found CUR_TIER_DEV_RECOG_CNT = 0x0F (15 units), CUR_TIER_SYS_RECOG_CNT = 0x0E (14 units), with a difference of 1;

[0249] 2) Root cause analysis: The SELF_NUM_IN_TIER of the missing device is 0x0D. Its PCIe configuration space Device Status register bit5 (Link Down) is set, and the PHY chip temperature reaches 92℃ (exceeding the threshold of 85℃).

[0250] 3) Preventive maintenance: The maintenance system generates a work order indicating that "there is an overheating risk in device 0x0D at level 3 of domain 0x02". The engineer replaces the heat sink within 2 hours, preventing the device from going down (MTBF is expected to be extended by 1200 hours).

[0251] Thirdly, the collaborative computing paradigm provides crucial support for the system. The switch (often referred to as the "brain card"), acting as the control core, can intelligently schedule tasks and distribute data with precision based on the ordinal information of downstream devices, the current valid state of each device's BAR (base address register) area, and its actual configuration size. By dynamically assessing the computational load and rationally allocating it to specific computing units, the system can effectively achieve efficient parallel processing and deep resource collaboration among multiple computing nodes. This mechanism not only significantly improves the response speed of the entire heterogeneous computing cluster in complex task environments but also optimizes the overall utilization of system resources, thereby enhancing the execution efficiency and throughput of large-scale computing tasks.

[0252] In a heterogeneous computing scheduling scenario for an AI inference cluster, a certain intelligent computing center's heterogeneous computing cluster includes 1 Brain Card (switch), 8 GPUs (ordinal numbers 0x01-0x08), 4 FPGAs (ordinal numbers 0x09-0x0C), and 12 NVMe storage devices (ordinal numbers 0x0D-0x18). The collaborative scheduling process for running multimodal AI inference tasks is as follows:

[0253] 1) Resource Status Mapping: The brain card scans downstream devices every 200ms, generating a dynamic resource matrix by reading the 0x10-0x17 fields (resource utilization) and 0x20-0x23 fields (temperature) of each device's BAR area. For example, it detects that the GPU 0x03's computing unit utilization is 65% and the remaining BAR space is 8GB, the FPGA 0x0A's DSP resource utilization is 32%, and the NVMe 0x12's IO bandwidth is 1.2GB / s (threshold 2GB / s);

[0254] 2) Task Feature Analysis: The newly integrated video inference task (1080P@30fps) consists of three stages: image decoding (requires high IO bandwidth) → feature extraction (requires FPGA acceleration) → model inference (requires GPU computing power). After the Brain Card analyzes the task metadata, it matches the device function tags (GPU = "Inference", FPGA = "Feature Extraction", NVMe = "Storage").

[0255] 3) Cross-device collaborative scheduling: Based on the dynamic resource matrix and function tags, the brain card performs three-level allocation: ① Storage layer: NVMe0x12 (1.2GB / s) with the lowest IO bandwidth is selected as the video stream storage node, and the address is written to the DMA address mapping table through the BAR space 0x40-0x47 fields; ② Acceleration layer: The feature extraction subtask is assigned to FPGA0x0A, and configuration instructions (including task priority 0x03) are sent through the PCIe message bus; ③ Computation layer: Among the remaining GPU resources, GPU0x03 with the lowest utilization rate is scheduled first, and the inference model parameter address is written through its BAR area 0x80-0x87 fields.

[0256] 4) Real-time load balancing: During task execution, the brain card detects that the GPU 0x03 utilization rate rises to 92% and immediately triggers load migration: 20% of the inference subtasks are diverted to GPU 0x05 (utilization rate = 45%). Seamless switching is achieved by modifying the 0xC0-0xC7 fields (task queue pointer) in the BAR area. The migration takes 1.8ms and there is no task interruption.

[0257] 5) Performance benefits: Through coordinated scheduling of ordinal information and BAR status, the end-to-end latency of tasks is reduced from 85ms to 42ms, the average utilization rate of equipment resources is increased from 62% to 89%, and the number of concurrent inference paths per cluster is increased from 24 to 41, resulting in a 73% improvement in overall energy efficiency compared to the traditional static scheduling scheme.

[0258] Through the above collaborative work, real-time adjustment of BAR can be supported in FPGA partial reprogramming scenarios: when the memory space needs to be expanded after FPGA reprogramming, the virtual CPU updates BAR_SIZE and SCALE through "SM4 encrypted instructions" without restarting the device, and the configuration time is ≤100ms, which meets the dynamic computing power expansion requirements.

[0259] The above-described embodiments, through the integration of two innovative mechanisms—"enhanced resource declaration" and "topology state awareness"—transform devices from passive resource consumers into intelligent entities with self-description, environmental awareness, and precise positioning capabilities. It meets large resource demands through scalable resource declarations, achieves rapid topology positioning through 3D coordinates and snapshot verification, supports dynamic reconfiguration, and adapts to flexible resource scheduling in large-scale clusters. This not only solves the resource requirement bottleneck of high-performance devices but also provides crucial hardware infrastructure for building large-scale, manageable, and collaborative next-generation computing systems.

[0260] While numerous embodiments of the invention have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Many modifications, alterations, and alternatives will occur to those skilled in the art without departing from the spirit and essence of the invention. It should be understood that various alternatives to the embodiments of the invention described herein may be employed in the practice of the invention. The appended claims are intended to define the scope of protection of the invention and therefore cover equivalents or alternatives within the scope of these claims.

Claims

1. A three-dimensional topology self-sensing management system based on PCIe devices, wherein the management system is applied to a PCIe board, characterized in that, The PCIe board includes a virtual CPU management terminal and a PCIe board terminal. The PCIe board terminal is a standard PCIe device that conforms to a preset standard specification and includes the three-dimensional topology self-sensing management system. The virtual CPU management terminal is installed on a standard PCIe bus and reads and writes to the PCIe board terminal through the computer bus.

2. The system according to claim 1, characterized in that, The preset standard specification is a global device group coding, and the global device group coding rules include: The device declaration area, the standard PCIe configuration space header area compatibility layer, the standard PCIe capability and status register group, the encrypted heritable device DNA identifier register, the device resource and topology management register group, the device historical status and reliability management register group, and the device communication affinity and behavior measurement register group, wherein the device resource and topology management register group includes a three-dimensional topology self-sensing management system.

3. The system according to claim 1, characterized in that, The three-dimensional topology self-sensing management system includes: an enhanced resource declaration mechanism and / or dynamic topology sensing and verification.

4. The system according to claim 3, characterized in that, The enhanced resource declaration mechanism includes a base address register set, an expandable space size register set, and / or a resource validity bitmap.

5. The system according to claim 4, characterized in that, The resource validity bitmap is mapped in such a way that one bit in the bitmap corresponds to one base address register.

6. The system according to claim 3, characterized in that, The dynamic topology awareness and verification includes: a topology location register group and / or an environment status register group.

7. The system according to claim 6, characterized in that, The topology location register group constructs a globally unique coordinate system to accurately identify the domain number to which the device belongs; wherein, the globally unique coordinate system is composed of the domain number, hierarchical information and device ordinal number.

8. The system according to claim 2, characterized in that, The device resource and topology management register group works collaboratively through a three-dimensional topology self-sensing management system, including: intelligent resource allocation, advanced diagnosis and maintenance, and / or collaborative computing paradigms.

9. The system according to claim 8, characterized in that, The collaborative computing paradigm includes generating a dynamic resource matrix for each downstream device of the system through a switch. The dynamic resource matrix includes: the ordinal information of the device, the current valid state of the base address register area of ​​each device, and its actual configuration size.

10. The system according to claim 9, characterized in that, The switch extracts features based on the inference task to generate device function tags, and performs dynamic evaluation of the computing load based on the dynamic resource matrix and device function tags, and allocates the load to specific computing units according to a preset scheme to complete parallel processing and resource collaboration between computing nodes.