Large-scale PCIe device coding management system
By setting up a virtual CPU management terminal and a PCIe board terminal on the PCIe board, the group coding of all devices in the domain is realized, which solves the problems of non-unique device identification and heavy main CPU burden in traditional PCIe device management, and improves the device management efficiency and stability of large-scale computing clusters.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HAILI COMPUTING (BEIJING) TECHNOLOGY CO LTD
- Filing Date
- 2025-12-05
- Publication Date
- 2026-04-24
AI Technical Summary
In traditional PCIe device management, the dynamic changes in BDF addresses result in a lack of globally unique identifiers, making it impossible to accurately identify the device identity in large-scale, high-performance computing scenarios, and placing a heavy management burden on the main CPU.
A large-scale PCIe device coding management system is adopted. By setting up a virtual CPU management terminal and a PCIe board terminal on the PCIe board, the group coding of all devices in the domain is realized, each device is assigned a unique identification code, and device management is carried out through embedded CPU and onboard memory, reducing the burden on the main CPU.
It enables rapid and accurate identification of device identities, improves cluster computing power utilization by more than 15%, supports the stable operation of more than 200 devices, is compatible with traditional PCIe specifications, and ensures the stability and security of cluster operation.
Smart Images

Figure CN121925645A_ABST
Abstract
Description
Technical Field
[0001] This invention generally relates to the field of computer bus technology. More specifically, this invention relates to a large-scale PCIe device encoding management system. Background Technology
[0002] In traditional computer systems, PCIe bus devices are addressed using a BDF address consisting of a Bus number, a Device number, and a Function number, and their device type is identified by a Vendor ID and a Device ID. This management method is a passive, static enumeration mechanism that lacks globally unique identifiers: the BDF address is assigned by software each time the system starts and may change dynamically. The Vendor / Device ID can only identify the device type and cannot distinguish between multiple specific device instances of the same type.
[0003] In view of this, there is an urgent need to provide a large-scale PCIe device coding management system in order to quickly and accurately identify the device identity in large-scale, high-performance computing scenarios. Summary of the Invention
[0004] In order to at least solve one or more of the technical problems mentioned above, the present invention proposes a large-scale PCIe device coding management system in several aspects.
[0005] In a first aspect, the present invention provides a large-scale PCIe device encoding management system, the encoding management system running on a PCIe board, the PCIe board including: a virtual CPU management terminal and a PCIe board terminal; wherein, the PCIe board terminal is a standard PCIe device conforming to a preset standard specification, the virtual CPU management terminal is installed on a standard PCIe bus, and it reads and writes to the PCIe board terminal through a computer bus.
[0006] In some embodiments, the PCIe board is an onboard memory, and the virtual CPU management terminal is a board that conforms to the PCIe specification.
[0007] In some embodiments, the preset standard specification is a global device group coding.
[0008] In some embodiments, the global device grouping coding rules include: a device declaration area, a standard PCIe configuration space header area compatibility layer, a standard PCIe capability and status register group, an encrypted heritable device DNA identifier register, a device resource and topology management register group, a device historical status and reliability management register group, and a device communication affinity and behavior measurement register group.
[0009] In some embodiments, the virtual CPU management terminal has a built-in embedded CPU and onboard extended memory. The embedded CPU manages devices according to a preset strategy by reading or writing PCIe board-side device information and operation logs on the bus into the onboard extended register.
[0010] In some embodiments, the device management according to the preset strategy includes the virtual CPU management terminal analyzing and processing the collected data through the preset strategy to monitor device status, provide fault warnings, and manage configuration.
[0011] In some embodiments, the PCIe board independently runs the encoding management system, which manages and operates the standard PCIe peripherals of the bus according to encoding rules.
[0012] In some embodiments, the PCIe board is logically recognized by the main CPU as a standard PCIe peripheral.
[0013] In some embodiments, the PCIe board encodes standard PCIe peripherals according to the preset standard specifications and writes the encoding into the onboard memory.
[0014] In some embodiments, the encoding management system includes an encoding-behavior bidirectional mapping memory mechanism to coordinate and organize the allocation of device functions.
[0015] Through the large-scale PCIe device coding management system provided above, this embodiment of the invention expands the traditional PCIe configuration space by setting up a virtual CPU management terminal and a PCIe board terminal on the PCIe board. Furthermore, in some embodiments, by integrating an embedded CPU and onboard memory into the virtual CPU management terminal, all PCIe board terminal devices on the bus can be discovered, their extended registers can be read / written, device information and operation logs can be collected, and device management can be performed according to preset strategies, thereby relieving the management burden on the main CPU. Even further, in some embodiments, by configuring global device group coding on the PCIe board terminal, a unique identity code can be dynamically assigned to each specific device instance, thereby achieving rapid and accurate identification of device identities in large-scale, high-performance computing scenarios. Attached Figure Description
[0016] The above and other objects, features, and advantages of exemplary embodiments of the present invention will become readily apparent upon reading the following detailed description with reference to the accompanying drawings. In the drawings, several embodiments of the invention are illustrated by way of example and not limitation, and like or corresponding reference numerals denote like or corresponding parts, wherein:
[0017] Figure 1An exemplary structural block diagram of a large-scale PCIe device coding management system according to an embodiment of the present invention is shown;
[0018] Figure 2 A diagram illustrating the structure of global device grouping coding rules according to an embodiment of the present invention is shown. Detailed Implementation
[0019] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0020] Figure 1 An exemplary structural block diagram of a large-scale PCIe device coding management system 100 according to an embodiment of the present invention is shown. Figure 1 As shown, the device coding management system 100 runs on a PCIe board, which includes a virtual CPU management terminal 101 and a PCIe board terminal 102. The PCIe board terminal 102 is a standard PCIe device that conforms to a preset standard specification. The virtual CPU management terminal 101 is installed on a standard PCIe bus and reads and writes to the PCIe board terminal through the computer bus.
[0021] The following description uses a large-scale computing cluster scenario as an example. This embodiment can manage over 200 PCIe devices (including GPU accelerator cards, NVMe SSDs, etc.) simultaneously, achieving efficient device control using the encoding management system provided in this embodiment. The large-scale PCIe device encoding management system is integrated onto a standard PCIe board, which may contain two core modules: a virtual CPU management terminal and a PCIe board terminal.
[0022] The PCIe board can be a standard PCIe device that conforms to the EX-DCC global device grouping coding specification. Specifically, it can be a 16-way PCIe 4.0 expansion interface module. Each interface can connect to peripherals such as GPUs and NVMe hard drives. All peripherals expand the traditional PCIe configuration space and have built-in encrypted genetic device DNA identifier registers, device resource and topology management register groups, etc., to ensure compliance with preset standard specifications.
[0023] The virtual CPU management terminal can be an independent board unit compliant with the PCIe 4.0 specification, installed in the standard PCIe slot (bus width x16) of the cluster host. Its hardware configuration includes an ARM Cortex-A53 embedded CPU (1.8GHz), a 1TB onboard NVMe SSD (for storing device information and operation logs), and 2GB DDR4 cache. It can also run an embedded Linux operating system and dedicated management software.
[0024] After system startup, the virtual CPU management terminal automatically enumerates all PCIe board peripherals via the computer's PCIe bus, reads the DNA identifier, topology information, and other data of each peripheral, and writes them to the onboard SSD. Simultaneously, it monitors the peripheral's operating status in real time, writing log information such as temperature and bandwidth usage to the onboard cache every minute, and then periodically synchronizing it to the SSD for persistent storage. When peripheral parameters need to be configured, the virtual CPU management terminal writes instructions to the device resource and topology management register groups on the PCIe board via the bus to complete the parameter adjustment, without requiring intervention from the main CPU. Furthermore, this PCIe board is logically recognized by the main CPU as a standard PCIe storage controller, allowing it to adapt to mainstream Windows and Linux systems without modifying the operating system kernel, ensuring cluster compatibility.
[0025] The "virtual CPU management terminal + standard PCIe board terminal" architecture of the above-described embodiment solves the problems of dynamic BDF address changes, lack of globally unique identifiers, and heavy management burden on the main CPU in traditional PCIe device management. On the one hand, the encrypted DNA identifier under the EX-DCC specification assigns a globally unique identity to each device, avoiding confusion between instances of the same type. On the other hand, the virtual CPU management terminal independently undertakes device enumeration, status monitoring, and configuration management tasks, allowing the main CPU to focus on computing power. According to the inventors' experimental calculations, the overall computing power utilization of the cluster can be improved by more than 15%. At the same time, the system is compatible with the traditional PCIe specification, can support the stable operation of more than 200 devices, meets the complex application requirements of large-scale computing power clusters, and deployment does not require modification of the system kernel, ensuring the stability and security of the cluster operation.
[0026] Furthermore, in one embodiment, in a large-scale computing cluster scenario (managing 200+ GPU / NVMe devices), the PCIe board is an onboard memory, and the virtual CPU management terminal is a board compliant with the PCIe specification. Complex device management is achieved through a "multi-level storage architecture + heterogeneous computing hardware + intelligent scheduling algorithm".
[0027] Specifically, the onboard memory design used on the PCIe board employs a three-tier storage architecture, including a DDR5 cache, an SLC NAND cache, and a QLC NAND main storage device. In terms of capacity, the system supports 2GB of DDR5 cache with a timing speed of 3200MT / s, suitable for fast temporary storage of real-time data. To further improve the performance of the DDR5 cache, Dynamic Voltage and Frequency Scaling (DVFS) technology can be introduced to dynamically adjust the voltage and frequency of the DDR5 based on the real-time data processing load, reducing power consumption while maintaining performance. The 1TB SLC NAND cache has a write / erase cycle life of up to 100,000 cycles and is used to store frequently accessed hot data. To extend the lifespan of the SLC NAND, a wear leveling algorithm is used to evenly distribute write operations across all storage cells, preventing excessive wear on local cells. An 8TB high-capacity QLC NAND serves as the main storage, responsible for storing less frequently accessed cold data. To improve the data read speed of the QLC NAND, Multi-Level Cell (MLC) mapping technology can be used to simulate MLC cells for reading, improving read efficiency. To improve data reliability, the entire storage system employs RAID 5 array technology for data redundancy, effectively avoiding the risk of data loss due to single points of failure. Furthermore, erasure coding technology can be introduced to further enhance data fault tolerance, ensuring data integrity even when multiple storage devices fail simultaneously.
[0028] In one implementation scenario, the memory is divided into three functional partitions: ① Device DNA Identification Area: Occupying 10GB SLC, set to read-only, storing the AES-256 encrypted DNA code of each peripheral. To further improve the security of the device DNA code, quantum key distribution technology is used to update and manage the encryption key. Quantum key distribution technology is based on the principles of quantum mechanics, ensuring the absolute security of the key. ② Real-time Log Area: Composed of 200GB DDR5 + SLC hybrid memory, with a read / write speed ≥2GB / s, it can record the device's operating status data for one hour, including key indicators such as temperature, bandwidth usage, and error count. To improve the processing efficiency of real-time logs, streaming computing technology is used. Streaming computing can process and analyze the log data generated in real time, promptly detect abnormal device behavior, and take corresponding measures. ③ Historical Data Area: Built using 8TB QLC and employing the ZFS file system, it can store historical logs for 30 days and supports data compression and snapshot functions, facilitating subsequent data analysis and system auditing. To improve the query efficiency of historical data, a distributed file system and indexing technology are used. Distributed file systems can distribute historical data across multiple nodes for storage, improving data read and write performance. Indexing techniques can speed up data retrieval and reduce query time.
[0029] Furthermore, to monitor the operating status of the hardware devices in real time, numerous sensors are deployed on the PCIe board and on the GPU and NVMe devices to monitor parameters such as temperature, humidity, voltage, and current. These sensors transmit data to the monitoring system in real time, which analyzes the data using machine learning algorithms to predict potential device failures and issue timely warnings. Simultaneously, to prevent interference and damage from the external environment, electromagnetic shielding technology and waterproof, dustproof, and shockproof designs are employed to improve the stability and reliability of the hardware devices.
[0030] Furthermore, regarding intelligent scheduling algorithms, the original algorithms primarily handle task allocation and resource management. To further improve scheduling efficiency, deep reinforcement learning algorithms can be employed. This algorithm continuously interacts with the environment to learn the optimal scheduling strategy, dynamically adjusting task allocation and resource usage based on factors such as task priority, device load, and data storage location, thereby improving overall system performance. Simultaneously, to support concurrent processing by multiple users and tasks, a distributed scheduling architecture is adopted, distributing scheduling tasks across multiple nodes for processing, thus enhancing scheduling parallelism and efficiency.
[0031] Furthermore, in terms of data management, in addition to using the ZFS file system to store historical logs, a data lake architecture can also be adopted. A data lake can store various types of structured and unstructured data, including device operation logs, performance data, and user behavior data. Through a data lake, data can be managed and analyzed in a unified manner, unlocking its potential value. Regarding data security, in addition to AES-256 encryption of the device's DNA code, homomorphic encryption technology is used to encrypt sensitive data stored in onboard memory. Homomorphic encryption allows computation on encrypted data without decryption, ensuring data security during processing. Simultaneously, to prevent data leakage, access control lists (ACLs) and authentication technologies are used to strictly control data access.
[0032] In one implementation scenario, the virtual CPU management terminal is configured as a PCIe 5.0 compliant board (bus width x16, bandwidth 64GB / s), employing a heterogeneous computing architecture: the core consists of two ARM Cortex-A76 embedded CPUs (2.4GHz, responsible for logic control), paired with one Xilinx Artix-7 FPGA (responsible for data acceleration processing), and integrated with a 16-channel PCIe switch (supporting device enumeration and bandwidth allocation). The board runs a customized embedded Linux system, pre-installed with a multi-level data scheduling algorithm and a device health assessment model. The scheduling algorithm dynamically migrates hot data (such as the device status accessed frequently within 1 minute) to SLC / NAND and archives cold data (such as logs from 24 hours ago) to QLC by real-time monitoring of the IOPS (input / output operations per second) and bandwidth utilization of each storage partition. The health model is based on the device temperature collected by the memory (weight 40%), PCIe bus error count (weight 30%), and bandwidth fluctuation (weight 30%). It outputs a health score of 0-100 by weighted summation, and triggers an alarm when the score is below 60.
[0033] The specific operating steps of the system are as follows: 1. Upon startup, the FPGA initializes the PCIe Switch using pre-programmed firmware, automatically scans the onboard memory of the 16 PCIe interfaces, obtains the device DNA identifier, and verifies it via AES-256 decryption; 2. The embedded CPU loads a scheduling algorithm to write the real-time collected device logs (temperature, bandwidth, etc.) to the DDR5 cache using PCIe DMA technology (avoiding CPU interrupt overhead), and synchronizes the cached data to the SLC NAND every 5 minutes; 3. At 2:00 AM every day, the algorithm automatically compresses logs exceeding 24 hours in the SLC (using compression algorithm LZ4) and archives them to the QLC, generating a data snapshot; 4. The health model reads the device data in the memory every 10 minutes, calculates the health score, and issues a dual warning via onboard LED indicators and a system pop-up window if the score is below the threshold.
[0034] The combination of the multi-level storage architecture and heterogeneous computing hardware described in the above embodiments is significantly different from the traditional PCIe management scheme of single storage + single CPU: On the one hand, the three-level storage architecture can reduce data read and write latency to below 50μs, which is 40% higher than the traditional solution, and RAID 5 and snapshot technology ensure data reliability and avoid log loss; on the other hand, FPGA accelerates device enumeration and data transmission, and the embedded CPU combined with the weighted health model significantly improves the accuracy of device fault early warning and reduces ineffective operation and maintenance; at the same time, the partition management and scheduling algorithm of the onboard memory supports efficient storage of log data from 200+ devices, reduces the average daily storage occupancy of a single device, further reduces the burden on the main CPU and cluster storage, and ensures the stable operation of large-scale computing clusters.
[0035] further, Figure 2 A structural diagram of a global device grouping coding rule 200 according to an embodiment of the present invention is shown.
[0036] As shown in rule 200, in one embodiment, the preset standard specification is EX-DCC (Expanse Domain Clustering Code) global device clustering coding. Specifically, global device clustering coding rule 200 may include: a device declaration area 201, a standard PCIe configuration space header area compatibility layer 202, a standard PCIe capability and status register group 203, an encrypted heritable device DNA identifier register 204, a device resource and topology management register group 205, a device historical status and reliability management register group 206, and a device communication affinity and behavior measurement register group 207.
[0037] In large-scale computing cluster scenarios (managing 200+ GPU / NVMe devices), the "hardware encryption + FPGA acceleration + reinforcement learning algorithm" technical solution provided in this embodiment of the invention can construct a multi-dimensional encoding management system, achieving a fundamental difference from traditional PCIe management solutions. Specific details are as follows:
[0038] In one implementation scenario, EX-DCC encoding can adopt a 512-bit fixed-length structure, divided into a "compatibility layer (64-bit) + extended function layer (448-bit)". The compatibility layer retains the traditional PCIe configuration space header (first 64 bytes) to ensure that the main CPU can recognize it. The extended function layer corresponds to seven register groups, each of which is allocated an independent address segment (address step size of 4 bytes). It can be directly read and written by the virtual CPU management terminal through the PCIe bus, and the main CPU has no access rights (implemented through FPGA address filtering).
[0039] The specific technical details and implementation steps of the seven major register groups are as follows:
[0040] 1. Device Declaration Area (32-bit Register): The EX-DCC encoded "identity foundation layer," which fundamentally addresses the lack of ownership traceability and key association in traditional PCIe devices. This declaration area can provide declaration information issued by the product owner, company name, information on authorized users, custom text information, and a key storage area for key management when providing software development tools in the future.
[0041] For example, field definitions can use 16-bit device type codes (0x0001 = GPU, 0x0002 = NVMe), 8-bit cluster affiliation codes (0x01 = computing cluster A, 0x02 = computing cluster B), and 8-bit device status codes (0x00 = offline, 0x01 = online, 0x02 = warning).
[0042] In terms of technical implementation, a "one-time programming + dynamic encrypted update" mechanism can be adopted, combined with a TPM security chip to achieve full lifecycle identity management. When the device leaves the factory, the device type code and ownership code are programmed by laser (which cannot be modified). The status code is updated by the virtual CPU management terminal every 50ms after reading the device hardware signals (such as power enable and link activation). The update command is transmitted encrypted using the SM4 national cryptographic algorithm.
[0043] Furthermore, in one embodiment, the device declaration area can be extended to a 128-bit hierarchical structure. Through "multi-dimensional field splitting, multi-level encryption protection, full lifecycle management, supply chain-level traceability, fine-grained permission control, and dedicated interaction protocols", a device declaration system adapted to large-scale, high-security, and high-reliability scenarios can be constructed.
[0044] Specifically, the device declaration area can be divided into 5 functional domains + 1 reserved domain. Each domain field can be encrypted and validated independently, and cross-domain association validation is supported. The specific field definitions are shown in the table below:
[0045]
[0046]
[0047]
[0048] To address the security requirements of different functional domains, a multi-layered encryption protection system (cross-algorithm + hardware anchoring) can be used. Specifically, a triple protection approach of "layered encryption + hardware anchoring + integrity verification" can be adopted to prevent field tampering and identity impersonation.
[0049] For the encryption of immutable fields in the basic identity field and the property rights traceability field, the following methods can be adopted: Basic Identity Field: RSA-2048 asymmetric encryption is used. The manufacturer signs fields such as device main type and hardware version using a private key. The virtual CPU management terminal verifies the validity of the signature using a pre-set manufacturer public key to prevent device type forgery. Property Rights Traceability Field: ECC-256 elliptic curve encryption is used. Each node in the production process (chip factory, board factory, testing factory) signs the production batch and serial number using a dedicated ECC key. The signature information is stored in the onboard OTP memory (one-time programmable, physically immutable) and associated with the manufacturer node on the blockchain (such as a consortium blockchain), supporting on-chain traceability verification.
[0050] For the encryption of variable fields in dynamic status fields / access control fields, the following methods can be used: Dynamic Status Fields: SM4-GCM authentication encryption is used. When the device status is updated (e.g., the health level drops from "Good" to "Medium"), the embedded CPU generates a random nonce value, which, combined with the session key generated by TPM 2.0, encrypts the status field. Simultaneously, an authentication tag is generated. The virtual CPU verifies the tag's integrity upon receiving it to prevent status tampering. Access Control Fields: AES-128-CCM encryption is used. Permission changes require an "application-approval-signature" process. For example, modifying configuration permissions requires the administrator to generate a permission change command through TPM 2.0, attaching the administrator's ECC signature. The device verifies the signature before allowing the update of the permission field.
[0051] For global integrity verification, each subfield can have an independent CRC check (CRC-16 / CRC-8 / CRC-4) built-in, and tampering of a single field can be detected in real time. The SHA-384 hash value of the entire 128-bit declaration area is stored in the PCR (Platform Configuration Register) 17 of TPM 2.0. After each declaration area is updated, the virtual CPU reads the PCR
[17] value and compares it with the calculated hash value. If they are inconsistent, a hardware-level alarm is triggered (the device PCIe link is cut off). The encryption keys are all stored in the TPM 2.0 security chip and adopt the "key hierarchical derivation" mechanism (root key → domain key → field key). The root key is melted down in the chip hardware and cannot be exported, thus preventing key leakage.
[0052] 2. Standard PCIe Configuration Space Header Compatibility Layer (64-bit Registers). This design adopts a strategy that emphasizes both compatibility and expansion, directly compatible with and referencing the PCI Express standard configuration space header layout defined by companies such as Intel, using this as the hardware foundation and logical starting point for the entire innovative register system.
[0053] The core purpose and advantages of the above design are:
[0054] 1) By strictly adhering to the standard specifications defined by the PCI-SIG organization, it ensures that hardware devices equipped with this innovative function can be correctly recognized as "standard" PCIe devices by existing standard motherboards, operating systems and BIOSes, and complete the most basic enumeration and initialization process.
[0055] 2) Clear Functional Positioning: This part of the registers (such as Vendor ID, Device ID, Class Code, Status, Command, etc.) defines the most basic identification, basic functions, and status of the device, thus clarifying the fundamental attribute of the device as a "general-purpose I / O device" and positioning its innovation scope as "enhancing memory consistency and topology management capabilities on top of the standard PCIe architecture," rather than creating a completely new and incompatible bus type. For example, the traditional PCIe 3.0 specification header fields are completely retained, such as Vendor ID (0x19E5, corresponding to custom vendors), Device ID (0x2001, corresponding to EX-DCC compatible devices), and the Command register (bit 0 controls device enable).
[0056] 3) Providing a foundation for expansion: The standard PCIe configuration space offers expansion capabilities (such as the Capabilities List mechanism). This design cleverly utilizes this mechanism, mounting subsequently defined innovative register sets (registers starting at offset 08H) as a set of custom, vendor-specific capabilities within the standard configuration space. This allows all innovative features to be organically integrated into the existing PCIe architecture ecosystem, and software can locate and access these new features through the standard PCIe capability discovery mechanism. For example, an 8-bit "EX-DCC flag" (0xAA indicating support for this specification) can be added to the end of the header area. During virtual CPU management initialization, this flag is scanned to quickly filter compliant devices.
[0057] In one embodiment, to further enhance its adaptability to multiple scenarios, problem traceability, and security protection capabilities, it may also include:
[0058] A multi-generation PCIe specification adaptive compatibility mechanism addresses cross-generation compatibility issues in scenarios where multiple generations of PCIe buses (3.0 / 4.0 / 5.0) coexist, enabling "plug-and-play" functionality for the board without manual configuration adjustments. In implementation, a new specification version detection register is added to the compatibility layer. After power-on, the board automatically identifies the PCIe generation of the main CPU bus (based on the Max Link Speed field of the Link Capabilities Register) via the LTSSM signal during the PCIe link training phase and writes the generation identifier to this register. Simultaneously, key fields in the header area are dynamically adjusted based on the detection results: for example, the PCIe 3.0 bus adapts to a maximum load of 128 bytes, and the PCIe 5.0 bus adapts to 512 bytes, ensuring that transmission efficiency matches bus capabilities. Furthermore, a multi-generation compatibility extension capability structure is added to the Capabilities List, declaring the range of versions supported by the board. During main CPU enumeration, the corresponding driver logic can be automatically loaded, avoiding enumeration failures or performance waste.
[0059] Through a multi-generation PCIe specification adaptive compatibility mechanism, the board can adapt to the 3.0-5.0 bus, adapting to mixed deployment scenarios of new and old clusters. At the same time, through dynamic parameter adjustment, it balances compatibility and transmission performance, eliminating the need for customizing multiple versions of hardware.
[0060] Furthermore, in one embodiment, to address the issue of "no access records and difficulty in locating faults" in traditional configuration spaces, a space access log and fault tracing mechanism can be configured. This lightweight log function enables full-link tracing of read and write operations, supporting rapid operation and maintenance troubleshooting.
[0061] In its implementation, the compatibility layer reserves a circular buffer to store the logs of the 16 most recent access operations. Each log entry includes the operation type (read / write), access address, data, and operation source identifier (main CPU / virtual CPU / BIOS, etc.). A "first-in, first-out" (FIFO) mechanism is used, automatically overwriting the oldest record when the buffer is full. Illegal access rules are also defined (such as writing to read-only fields or accessing undefined addresses). When a violation is detected, the current log is immediately latched and overwriting is prohibited. A main CPU interrupt is triggered by setting the Status Register, and an exception notification is simultaneously sent to the virtual CPU management terminal. The main CPU or management terminal can read the latched log to quickly locate the operation source, time, and data, enabling fault tracing.
[0062] This mechanism requires no external debugging tools and can reduce the time for troubleshooting configuration space-related faults from hours to minutes. It also supports operational behavior auditing, clarifies the operational responsibilities of different entities, and reduces maintenance disputes.
[0063] Furthermore, in one embodiment, to address the security risks caused by the traditional "full access" of configuration spaces, the integrity of the configuration space is ensured through hierarchical control of configuration space access permissions, preventing malicious tampering and unauthorized access.
[0064] In terms of implementation, the compatibility layer adds an access control register, dividing configuration space fields into three levels of permissions based on importance: Level 0 (read-only, such as vendor ID / device ID) prohibits all subjects from writing; Level 1 (restricted write, such as the command register) only allows authorized main CPU processes and the virtual CPU management terminal to write; Level 2 (administrator write, such as the extended capability list pointer) requires TPM authentication before writing. When a subject initiates a write operation, it must submit an application frame containing its identity identifier, operation fields, and TPM signature. The compatibility layer verifies the signature validity through the TPM chip on the virtual CPU management terminal. If the permission level matches, the operation is allowed; otherwise, it is rejected and an exception log is recorded. In addition, the virtual CPU management terminal can dynamically update permission configurations (such as temporarily granting permissions during debugging) and retain permission change logs for auditing.
[0065] This mechanism can effectively prevent attacks such as device disabling and identity spoofing, while also taking into account operational flexibility and meeting the dual needs of secure operation and efficient debugging of large-scale clusters.
[0066] The above-described embodiments utilize the industry-standard PCIe standard as a "universal language" and "passport," thereby ensuring device identifiability and interoperability, and laying a solid hardware and protocol foundation for the implementation and access of all subsequent innovative functions, maximizing the use of existing ecosystem resources.
[0067] 3. Standard PCIe capabilities and status register set.
[0068] This register group is the "basic function support layer" of EX-DCC encoding. Its core is "integrated standard capabilities + enhanced status linkage". It fully integrates the key link, device and power status registers in the PCIe standard, and accelerates the acquisition and interrupt linkage through FPGA, so as to provide real-time and accurate low-level status data.
[0069] Specifically, it includes:
[0070] 1) Fully integrated standard registers: Includes three main categories of core registers (128-bit):
[0071] Link Status / Control Register (LNK_STS / LNK_CTL): Records link width (0x08 = 8 channels), speed (0x04 = PCIe 4.0), and error count, used to determine the physical connection status of the device.
[0072] Device Status / Control Register (DEV_STS / DEV_CTL): Records internal device errors (such as ECC errors, checksum errors), and supports error mask configuration.
[0073] Power Management Register (PWR_MGMT): Records current power consumption (in W) and power consumption threshold, and supports switching between high performance and power saving modes (0x01 / 0x02).
[0074] 2) FPGA-accelerated state acquisition: The FPGA's built-in "state machine + DMA" module can, for example, acquire all register data every 10ms and write it directly to the virtual CPU's onboard cache via PCIe DMA, without occupying embedded CPU resources. Differential signal sampling technology is used during data acquisition to reduce data errors caused by electromagnetic interference (error rate ≤ 0.1%).
[0075] 3) Dynamic interrupt linkage mechanism: Different thresholds are set for different registers (e.g., GPU power consumption threshold 0x46 = 70W, NVMe power consumption threshold 0x1E = 30W; link error count threshold 0x05 = 5 times). When the data exceeds the threshold, the FPGA triggers a level interrupt, and the virtual CPU can respond within 1ms (e.g., when the power consumption exceeds the limit, the GPU frequency is reduced to 2.0GHz, and when the link error exceeds the limit, the link is restarted), thereby improving the efficiency of traditional polling response (≥10ms).
[0076] The above embodiments of the present invention achieve real-time acquisition and rapid response of status data through FPGA-accelerated acquisition and dynamic interruption, while providing underlying status input (such as link rate affecting communication latency assessment) for subsequent topology management and affinity calculation. This solves the problem that existing technologies rely heavily on CPU polling to read status, resulting in slow response and high resource consumption.
[0077] 4. Encryptable genetic device DNA identifier register
[0078] This design register is a "secure root of trust" encoded by EX-DCC. It breaks through the limitation of traditional Vendor / Device IDs that can only identify "type". Through "double encryption + genetic design", it gives each device a globally unique, verifiable hardware identity with "lineage" relationship, solving the problems of device counterfeiting and difficulty in tracing.
[0079] Specifically, it includes:
[0080] 1) 128-bit DNA structure design: DNA consists of four parts, ensuring uniqueness and heritability. Specifically, this may include:
[0081] -24-bit vendor root key (0x123456789ABC, stored in TPM 2.0 chip, unreadable).
[0082] -48-bit unique serial number for each device (laser-burned at the factory, globally unique, e.g., 0x100000010001~0x100000010200 corresponds to 200 devices).
[0083] -32-bit family identifier (shared by devices of the same batch / model, such as GPU family identifier 0x00010000, reflecting "lineage").
[0084] -24-bit CRC-32 checksum (calculated from the first 104 bits to ensure data integrity).
[0085] 2) Dual Encryption and Authentication: Employs a dual mechanism of "RSA-2048 asymmetric encryption + AES-256-GCM symmetric encryption".
[0086] - When the device leaves the factory, the DNA is signed using the manufacturer's RSA private key.
[0087] When a device connects to the cluster, the virtual CPU sends a "DNA verification request" → the device encrypts the DNA using an AES key → the virtual CPU decrypts the DNA using the RSA public key in the TPM and verifies the CRC → connection is only allowed after successful verification. Experiments have verified that a single verification takes ≤200μs, and the spoofing detection rate is 100%.
[0088] 3) Genetic Application: The virtual CPU identifies devices in the same batch (such as GPU family identifier 0x00010000) through the "family identifier hash matching algorithm", automatically assigns them the same initial communication affinity weight (0xFFFF), and enables "cluster collaborative optimization" (such as when multi-GPU parallel computing, computing groups are quickly formed based on family identifiers, reducing data synchronization latency by 30%).
[0089] The technical solution in this embodiment ensures identity credibility through double-encrypted DNA and identifies the "lineage" of devices through family identifiers, thereby distinguishing devices of the same model, providing accurate traceability, and providing technical support for accurate operation and maintenance (fault warning for the same batch) and collaborative computing.
[0090] Furthermore, in one embodiment, to address the issue of single-domain TPM authentication failure across multiple management domains (compute / storage / network domains), a cross-domain DNA joint authentication mechanism can also be included. Specifically, this is achieved through a two-level certificate system of "root CA + domain CA," ensuring cross-domain compatibility and authentication redundancy.
[0091] In terms of implementation, a 16-bit "certificate identifier field" is added to the DNA structure. At the time of device shipment, a dual certificate chain is generated: "Root CA certificate (issued by the cluster central control, embedded in all TPMs) - Domain CA certificate (domain-specific, containing domain identifier and RSA public key)," which is stored together with the DNA in the OTP. During cross-domain migration, the target domain management first verifies the validity of the domain CA certificate, then decrypts the DNA signature using the domain CA public key. If single-domain verification fails, a multi-domain joint verification is initiated. If more than 2 / 3 of the domains pass, the verification is deemed valid, increasing the authentication success rate from 99% to 99.99%, with verification time ≤300μs.
[0092] Furthermore, in one embodiment, DNA-equipment lifecycle status binding and traceability can be adopted for the entire lifecycle status (production / in service / maintenance / retirement) of associated equipment, combined with blockchain evidence storage to achieve full-link traceability of "status-DNA-operator", thus making up for the traceability limitations of the original solution.
[0093] Specifically, the original 24-bit CRC checksum of DNA can be split into a 16-bit checksum + a 24-bit "lifecycle status field" (4-bit status type, 8-bit operator ID, and 12-bit timestamp). Status changes require TPM signature from maintenance personnel, and the field is updated after verification by the virtual CPU, generating a "status change record". The record hash value is stored on the blockchain (to prevent tampering), and a "DNA-block mapping table" is built off-chain. For example, after maintenance, it can be traced back to "2024-10-25, maintenance personnel replaced the GPU core at 0x12, and the status changed from in service (0010) to maintenance (0100)", improving the accuracy of fault tracing from the device level to the operation level.
[0094] Furthermore, in one embodiment, to address the issue of identifier conflicts between different manufacturers with the same serial number across clusters, a cluster DNA identifier adaptation and conflict avoidance mechanism can be adopted. Through cluster identifier embedding and conflict detection, cross-cluster uniqueness can be guaranteed.
[0095] Specifically, the implementation is as follows: 8 bits of "cluster identifier field" (e.g., cluster A = 0x01, cluster B = 0x02) are split from the 32-bit DNA family identifier. When the device first connects to the cluster, the virtual CPU uses the cluster AES key to encrypt the DNA twice to generate "cluster adaptation DNA". During migration, the identifier is updated and re-encrypted, while the core fields (manufacturer key / serial number) remain unchanged.
[0096] When the management terminal starts up, it scans all devices for the combination of "vendor key + serial number + cluster identifier" and stores it in a hash table. When a conflict occurs, it compares the original DNA signature. If they do not match, it is determined to be counterfeit. If they match, the serial number is updated (e.g., 0x10000001→0x100000010001). Scanning 200 devices takes ≤10ms. The efficiency of conflict investigation can be improved by hundreds of times compared to manual methods.
[0097] It should be further noted that the solution in this embodiment is a multi-layered, mutually reinforcing system. The design purpose and technical effects of the solution are explained below:
[0098] 1) Building an absolutely trustworthy identity foundation (uniqueness + encryption):
[0099] Beyond Standard Identifiers: Overcomes the limitations of Vendor / Device ID as a "type identifier" and provides a unique "instance identifier" for each physical device.
[0100] Hardware-level authentication: This DNA code can be designed as a digital credential signed with an asymmetric encryption algorithm. The system can verify the signature to authenticate the device's legitimacy, confirming its authenticity and lack of tampering, thereby fundamentally resisting hardware counterfeiting and cloning attacks and laying a solid foundation for system security.
[0101] 2) Enable intelligent collaboration and relationship-aware management (genetic similarity): By building a device relationship network, "genetic similarity" means that devices derived from the same source (such as the same manufacturer, chip, FPGA code) will have a recognizable pattern in their DNA, so that the system can intelligently identify the "bloodline" or "family" relationship between devices.
[0102] Enables advanced features: The system can leverage this relationship network to achieve intelligent collaboration (enable specific optimization modes for AI accelerator cards of the same family), precise operation and maintenance (predict the risk of the same batch of devices based on the failure of one device), and group authorization management (authorize software features to the entire product family).
[0103] 3) Achieve full lifecycle traceability (uniqueness + heritability):
[0104] Enhanced traceability: Combining a unique serial number with embedded manufacturer digital characteristics enables ultimate traceability from the chip level to the system level, which is beneficial for supply chain management, quality control, and targeted recalls. For example, in one embodiment, the unique identifier can be designed as a 24-bit manufacturer root key (0x123456, stored in the virtual CPU management terminal security chip), a 32-bit device unique serial number (factory-programmed, globally unique), and an 8-bit CRC-8 checksum (calculated from the first 56 bits).
[0105] Supporting a secure supply chain: From manufacturing to deployment, each stage allows verification of the device's encrypted DNA and genetic lineage, ensuring a clear and trustworthy origin. For example, the encryption mechanism employed can use the AES-256-GCM encryption algorithm, with the key stored in the TPM 2.0 chip on the virtual CPU management terminal. Each DNA identifier read requires a three-step process: "key verification - data decryption - checksum comparison," preventing identifier tampering. Decryption time can be ≤200μs.
[0106] 4) Supports future trusted computing and extended applications (encryption + uniqueness):
[0107] High trust level: This encrypted DNA can be used as a seed to generate a device-specific encryption key to ensure data security within the device and supports trusted computing paradigms.
[0108] Highly scalable: The massive 128-bit address space provides an identity solution for a vast number of IoT devices, giving them a unique and verifiable identity in future network environments.
[0109] Through the above embodiments, a hardware-level "identity, security, and relationship" system is used to enable devices not only to be identified, but also to be authenticated, trusted, and have their underlying organizational relationships understood by the system, thereby providing technical support for intelligent collaboration, advanced security, precise management, and future applications.
[0110] 5. Device Resource and Topology Management Register Set
[0111] This register group is the "resource and location hub" of EX-DCC encoding. It breaks through the limitations of the traditional PCIe static BAR (base address register) and solves the problems of high resource demand and difficult location of complex topologies for high-performance devices through "enhanced resource declaration + dynamic topology awareness", providing support for resource scheduling of large-scale clusters.
[0112] In one embodiment, it may further include:
[0113] 1) Enhanced resource declaration mechanism: Breaking through the traditional 32-bit BAR 4GB addressing limit, it adopts a "base size + scaling factor" design, including:
[0114] SELF_BARn_BASE_ADDR (32-bit): The physical base address allocated by the storage system (e.g., 0x80000000);
[0115] SELF_BARn_SIZE (16-bit): Base space size (e.g., 0x1000 = 4KB);
[0116] SELF_BAR_SIZE_SCALE (8 bits): scaling factor (e.g., 0x0A = 1024), actual address space = 4KB × 1024 = 4GB, supporting the large memory mapping requirements of devices such as GPUs;
[0117] SELF_BAR_VALID (8-bit bitmap): 1 bit corresponds to 1 BAR. bit0=1 indicates that BAR0 is enabled. The software can determine the BAR status without reading back to detect, which can improve the configuration efficiency by 50%.
[0118] In one embodiment, the enhanced resource declaration mechanism achieves an innovative breakthrough over the 4GB address space limitation of the traditional 32-bit base address register (BAR) by employing a collaborative design architecture of "base size + scaling factor". This architecture mainly consists of three core parts: the base address register set (SELF_BARn_BASE_ADDR), the scalable space size register set (SELF_BARn_SIZE and SELF_BAR_SIZE_SCALE), and the resource validity bitmap (SELF_BAR_VALID). These components together construct a highly flexible and scalable address mapping mechanism, endowing high-performance computing devices with more powerful resource configuration capabilities.
[0119] Furthermore, in one embodiment scenario, the resource validity bitmap may also include a security enhancement module. This module consists of an access control matrix, an abnormal behavior detector, and a security log register. The access control matrix implements a 32×32 permission control table based on the RBAC model, supporting fine-grained authorization for 8 user groups. The abnormal behavior detector has a built-in LSTM neural network inference engine (quantized to 8-bit weights) capable of monitoring abnormal BAR access patterns in real time. The security log register records 64 security event logs, including operation type (4 bits), user ID (8 bits), timestamp (20 bits), and result code (4 bits). It also supports log signature functionality using the SM2 national cryptographic algorithm to ensure that audit information cannot be tampered with.
[0120] Furthermore, in one embodiment scenario, the resource validity bitmap may also include high-speed interfaces and protocol extensions. Specifically, it may include an AXI4-Stream interface, a PCIe MSI-X extension, and a JTAG debug interface. The AXI4-Stream interface implements a 128-bit data width stream interface, supporting DMA batch transfers in BAR state. The PCIe MSI-X extension provides 32 MSI-X interrupt vectors, which can be configured as interrupt sources for events such as BAR state changes and collision detection. The JTAG debug interface supports forced bitmap state setting and single-step tracing in bit debug mode. An integrated SerDes interface hard core achieves a maximum state information transmission rate of 8Gbps to meet the centralized monitoring needs of large systems.
[0121] Furthermore, in one implementation scenario, a cache-optimized resource prefetching strategy can be introduced on top of the existing enhanced resource declaration mechanism. By conducting in-depth analysis of historical resource access patterns and combining this with the current system's operating status, resources that may be accessed in the future can be predicted, and these resources can be pre-loaded into the cache.
[0122] 2) Dynamic Topology Awareness and Verification: Constructing a three-dimensional coordinate system of "domain-hierarchy-ordination", including:
[0123] CUR_TIER_IN_DOMAIN (8 bits): Device domain (0x02 = second domain);
[0124] SELF_NUM_IN_TIER (8 bits): Domain level number (0x05 = fifth level);
[0125] The virtual CPU updates the coordinates hourly using a "topology scan algorithm" (traversing the downstream ports of the PCIe switch) and compares them with the topology snapshot in the onboard ZFS file system. Simultaneously, by analyzing the difference between CUR_TIER_DEV_RECOG_CNT (number of physical devices) and CUR_TIER_SYS_RECOG_CNT (system enumeration count) (e.g., difference = 1), unenumerated devices can be located within 50ms, thereby improving fault location efficiency.
[0126] Furthermore, in one embodiment, dynamic topology awareness and verification leverages self-localization and environmental awareness capabilities to achieve device collaboration and system-level management. By constructing a three-dimensional coordinate system of "domain-hierarchy-ordination," devices are logically assigned unique identifiers and clear hierarchical relationships within the topology. This coordinate system acts like a precise map, determining a unique position for each device within the system, facilitating unified management and scheduling. Specifically, it may include: a topology location register group (CUR_TIER_IN_DOMAIN, SELF_NUM_IN_TIER) and an environmental status register group (CUR_TIER_DEV_RECOG_CNT, CUR_TIER_SYS_RECOG_CNT, CUR_TIER_MAX_DEV).
[0127] First, the CUR_TIER_IN_DOMAIN register in the topology location register group serves as the core configuration unit, its main function being to accurately identify the domain number to which a device belongs. In complex multi-rooted or partitioned system architectures, this register provides crucial support for the rapid location of device logical subtrees, thus ensuring the efficiency of system resource management and scheduling in multi-domain environments. Simultaneously, the SELF_NUM_IN_TIER register, by assigning a unique sequence number to each device at its respective level, effectively achieves accurate differentiation and identification between devices at the same level, significantly reducing the probability of conflicts during device identification. The globally unique coordinate system (composed of domain number, hierarchy information, and device ordinal number) constructed through the collaborative work of these register groups lays the foundation for accurate allocation of system resources, efficient implementation of fault isolation mechanisms, and deep optimization of Non-Unified Memory Access (NUMA) architectures, thereby improving the reliability and performance of large-scale computing systems.
[0128] Secondly, in one embodiment, the CUR_TIER_DEV_RECOG_CNT register in the environment status register group aims to accurately count and report the total number of physical devices actually identified in the current level from a hardware perspective, truly reflecting the "physical truth" of the system's underlying structure. CUR_TIER_SYS_RECOG_CNT, on the other hand, records the number of devices in the same level identified by the system software during the enumeration process. Furthermore, CUR_TIER_MAX_DEV is used to explicitly define the maximum number of devices that the current level can support on the physical hardware, i.e., the so-called "physical capacity" upper limit. Through this register group, the system can provide each device with a comprehensive view of its local operating environment, allowing it to clearly understand the status of its level. By comparing the number of physical devices with the number enumerated by the system in real time, the system can achieve rapid and consistent diagnostics and effective health assessment. The maximum device value further clarifies the theoretical limit of the environment in terms of device expansion, providing an important basis for resource planning and dynamic adjustment.
[0129] 3) Dynamic resource reconfiguration: Supports real-time adjustment of BAR in FPGA partial reprogramming scenarios: When the memory space needs to be expanded after FPGA reprogramming, the virtual CPU updates BAR_SIZE and SCALE through "SM4 encrypted instructions" without restarting the device. The configuration time is ≤100ms, which meets the needs of dynamic computing power expansion.
[0130] The above-described embodiments meet large resource requirements through scalable resource declarations, achieve rapid topology positioning through three-dimensional coordinates and snapshot verification, support dynamic reconfiguration, and adapt to flexible resource scheduling for large-scale clusters.
[0131] 6. Equipment historical status and reliability management register group
[0132] The Equipment Historical Status and Reliability Management Register Group is the "health management core" in the EX-DCC coding system. By integrating error recording, performance benchmarking, and historical configuration backup functions through a comprehensive strategy based on "history recording + benchmark comparison + fault recovery", this solution overcomes the shortcomings of traditional technologies that can only view the current status, are difficult to predict faults, and have low recovery efficiency, and realizes reliability management throughout the entire life cycle of the equipment.
[0133] In this embodiment, the technical solution mainly includes the following six key parts:
[0134] 1) Error logging and trend analysis module:
[0135] The system employs a combination of a circular buffer and a weighted counter. The SELF_ERR_CREDIT (an 8-bit register) is used as an error counter, with different weights assigned to different error types (e.g., ECC error weight is 0.6, and communication timeout weight is 0.4). When the weighted cumulative value reaches the preset threshold 0x0A (i.e., 10 points), a system warning is triggered.
[0136] The HIST_ERR_TYPE (64-bit register) stores an error type code in 16-bit increments, storing a maximum of the four most recent error types (e.g., 0x01 indicates an ECC error, 0x02 indicates a communication timeout). New errors overwrite the oldest records, supporting root cause analysis (e.g., three consecutive ECC errors can be identified as a hardware failure).
[0137] The virtual CPU runs a "weighted moving average (WMA) algorithm" to calculate a reliability score (range 0-100) based on the most recent 128 error records. When the score is below or equal to 60, it is marked as "high risk," with an early warning accuracy rate of no less than 95%.
[0138] 2) Performance benchmark and loss quantification module:
[0139] LINK_LATENCY (32-bit register) is used for the ideal communication latency of the storage device (e.g., 200ns for GPUs, 150ns for NVMe). Baseline values are updated monthly via a "benchmark calibration algorithm" (tested 3 times under no-load conditions and the average value is taken) to ensure data accuracy.
[0140] During runtime, the actual latency is compared with the baseline value to quantify the performance loss (e.g., if the actual latency is 300ns, the loss is 50%), providing a basis for performance optimization (e.g., cleaning up link interference when the loss exceeds 30%).
[0141] 3) Historical configuration backup and recovery module:
[0142] Back up topology configuration information such as LAST_SELF_TIER (level) and LAST_MASTER_TIER (upper level) and store it in onboard non-volatile memory (NVM).
[0143] When the number of current configuration errors reaches or exceeds 3 (e.g., hierarchical anomaly), the system automatically loads the backup configuration, and the recovery time is no more than 200ms, which is 99% more efficient than the traditional restart recovery method (which usually takes more than 5 minutes).
[0144] 4) Environmental Adaptability Monitoring Module:
[0145] The operating environment of equipment has a significant impact on its performance and reliability. This module utilizes multiple sensors to monitor the environmental parameters of the equipment in real time. Temperature sensors accurately measure the temperature around the equipment, humidity sensors monitor ambient humidity, and a vibration sensor is also included to detect abnormal vibrations affecting the equipment. Excessive temperature may lead to decreased equipment performance or even damage, abnormal humidity may cause short circuits, and abnormal vibrations may indicate unstable installation or loose internal components.
[0146] These sensors transmit data to the system in real time, which compares these environmental parameters with preset suitable ranges. If an environmental parameter exceeds the normal range, the system immediately issues a warning and takes appropriate measures based on the specific situation. For example, when the temperature is too high, the system can automatically adjust the equipment's workload or activate the cooling system to enhance heat dissipation; when the humidity is abnormal, it can activate the dehumidification function or issue a notification to staff for handling. Through real-time environmental monitoring and timely response, the reliability and stability of the equipment in different environments are greatly improved.
[0147] 5) Data integrity verification module:
[0148] Data integrity is paramount during device operation. This module employs advanced verification algorithms to perform real-time verification of the data stored and transmitted by the device. For stored data, cyclic redundancy check (CRC) is performed periodically to ensure that the data has not been corrupted or lost during storage. During data transmission, hash algorithms are used to encrypt and verify the data, preventing tampering or errors during transmission.
[0149] When data anomalies are detected during verification, the system immediately marks the problematic data and attempts to repair it. If repair is unsuccessful, the system will promptly notify the administrator for further processing. Simultaneously, this module records historical data verification information, including verification time and results, facilitating subsequent data analysis and troubleshooting. Through rigorous data integrity verification, the accuracy and reliability of device data are guaranteed, preventing equipment failures and business interruptions caused by data errors.
[0150] 6) Intelligent resource scheduling module:
[0151] During equipment operation, different tasks have varying resource requirements. This module achieves intelligent resource scheduling by monitoring and analyzing the equipment's resource usage in real time. The system dynamically allocates resources such as CPU, memory, and bandwidth based on task priority, resource requirements, and the equipment's current resource status.
[0152] For example, for high-priority tasks, the system allocates sufficient resources to ensure efficient operation; while low-priority tasks can be processed when resources are idle. Simultaneously, this module adjusts resource allocation strategies based on device performance degradation and error logs. When a portion of a device's resources experiences performance degradation or frequent errors, the system reduces the use of those resources and allocates tasks to other, higher-performing resources. Through intelligent resource scheduling, the overall performance and resource utilization of the device are improved, reducing device failures and performance bottlenecks caused by unreasonable resource allocation.
[0153] This technical solution achieves predictive maintenance through error trend analysis, quantifies performance loss through benchmark comparison, and quickly restores system status through backup configuration. At the same time, it combines environmental adaptability monitoring, data integrity verification, and intelligent resource scheduling to significantly improve the reliability and maintainability of the equipment.
[0154] 7. Device Communication Affinity and Behavior Measurement Register Set
[0155] This group serves as the "intelligent collaboration hub" of the EX-DCC encoding. Its core breakthrough lies in overcoming the limitations of traditional static communication management. Through "quantified behavior + machine learning optimization," it achieves dynamic perception and optimization of device communication relationships, solving the problems of high communication latency and resource waste in large-scale clusters. The specific implementation method is as follows:
[0156] 1) Quantitative measurement of communication behavior: A "multi-dimensional indicator + sliding window" design is adopted.
[0157] SELF_VISIT_WEIGHT / OTHER_VISIT_WEIGHT (16 bits): Quantify the priority of the device "initiating access" and "being accessed" respectively (e.g., GPU access to NVMe weight 0xFFFF, access to network card weight 0x8000).
[0158] SELF_VISIT_QUTIY / OTHER_VISIT_QUTIY (32-bit): 1-minute sliding window (50% overlap, reducing statistical error) statistics on traffic (in KB), such as GPU-NVMe traffic of 100,000 KB.
[0159] VISIT_WEIGHT_SCALE (8 bits): Weight scaling factor (e.g., 0x02 = 2 times), adapting to different business priority requirements (e.g., increasing GPU weights in AI training scenarios).
[0160] 2) GNN+Q-Learning Intelligent Optimization:
[0161] Every 5 minutes, the virtual CPU constructs a communication affinity graph (node = device, edge weight = affinity score) based on weight, traffic, and topological distance (calculated in three-dimensional coordinates).
[0162] Run the Q-Learning algorithm to optimize affinity: With the goal of "lowest cluster communication latency", dynamically adjust the weight (e.g., if the communication volume is ≥100MB / min for 5 consecutive minutes, the weight increases by 10%), and adjust the step size ≤5% to avoid fluctuations;
[0163] By allocating high-affinity devices (such as GPUs and NVMe) to the same PCIe switch based on the graph, communication latency can be reduced from 500ns to 300ns, and the overall cluster transmission efficiency can be improved by 40%.
[0164] 3) Traffic load balancing linkage: Affinity data is synchronized to the traffic management module of the PCIe Switch, and more bandwidth is allocated among high affinity devices (such as 40%), thereby avoiding link congestion and effectively improving bandwidth utilization.
[0165] Compared to existing technologies that statically configure communication weights without intelligent optimization, the embodiments of this invention quantify communication behavior in multiple dimensions, use GNN+Q-Learning to achieve dynamic affinity optimization, link traffic management, significantly reduce communication latency, and improve cluster collaboration efficiency.
[0166] In summary, this solution achieves a comprehensive breakthrough in traditional PCIe management technology through the collaborative design of the seven register groups of EX-DCC. Experimental results show that: in terms of compatibility, the underlying layer fully complies with the PCIe standard, with a 100% device enumeration success rate, and can be deployed without modifying the system kernel. In terms of security, the encrypted DNA combined with the TPM chip provides 100% protection against device spoofing, and fault tracing accuracy reaches the single-device level. In terms of management efficiency, topology location time is reduced from 30 minutes to 50ms, fault warning accuracy is ≥95%, and device recovery time is reduced from 5 minutes to 200ms. In terms of performance, communication latency is reduced by 40%, GPU computing power utilization is increased by 20%, and bandwidth utilization is increased by 25%. Overall, it can support the stable operation of 200+ GPU / NVMe devices, fully meeting the secure, efficient, and intelligent management requirements of large-scale computing clusters.
[0167] Furthermore, in one embodiment, the virtual CPU management terminal has a built-in embedded CPU and onboard extended memory. The embedded CPU manages devices according to a preset strategy by reading or writing PCIe board-side device information and operation logs on the bus into the onboard extended register.
[0168] In the above embodiments, the virtual CPU management terminal is an independent board conforming to the PCIe 5.0 specification (bus width x16, bandwidth 64GB / s). Its core adopts a heterogeneous architecture of "embedded CPU + FPGA + hierarchical storage," with clearly defined hardware parameters and functional divisions, ensuring efficient collaboration between data acquisition, computation, and storage. The built-in devices are described in detail below:
[0169] 1. Embedded CPU module
[0170] Specifically, an industrial-grade ARM Cortex-A78AE processor (4 cores, 2.8GHz, supporting ECC memory verification) can be selected. It features high reliability and low power consumption, and is specifically responsible for strategy algorithm execution, device configuration instruction generation, and system interaction. Its core technologies include:
[0171] - Supports heterogeneous deployment of dual operating systems: embedded Linux (responsible for policy management) and RTX real-time system (responsible for interrupt response, real-time performance ≤1ms), avoiding resource contention through hardware isolation (MMU memory partition).
[0172] -Integrated hardware encryption engine (supports SM4, AES-256-GCM). All configuration instructions must be encrypted by the engine before being transmitted to the onboard extended registers to prevent the instructions from being tampered with.
[0173] - Communicates with the FPGA via the AXI-4 bus (10GB / s bandwidth), with a data interaction latency of ≤500ns, ensuring that the acquired data is quickly transmitted to the CPU for algorithm processing.
[0174] 2. Onboard extended memory module
[0175] In one embodiment, a "three-tiered storage architecture" can be adopted to adapt to different data access frequencies and storage requirements. Specific parameters and functional examples are as follows:
[0176]
[0177]
[0178] 3. Onboard Extended Register Design
[0179] The register set addresses are mapped to the local address space of the virtual CPU management terminal (0x90000000-0x9000FFFF), and can be divided into three categories of function registers, supporting direct read and write operations via PCIe DMA (avoiding CPU interrupt overhead). The specific definitions are as follows:
[0180] - Device Information Register (0x90000000-0x90007FFF): Each device is allocated a 32-byte register. Bits 0-7 store the device type (0x01 = GPU, 0x02 = NVMe), bits 8-15 store the real-time temperature (in °C), and bits 16-31 store the bandwidth utilization (%). It is updated by the FPGA every 10ms via DMA.
[0181] - Policy Configuration Register (0x90008000-0x9000DFFF): Stores preset policy parameters (such as temperature threshold, error count threshold). Bits 0-15 are the GPU temperature threshold (default 0x55 = 85℃), and bits 16-31 are the error count threshold (default 0x0A = 10 times). Only the CPU can write to this register (encryption verification required).
[0182] - Interrupt Status Register (0x9000E000-0x9000FFFF): bits 0-19 correspond to the interrupt flags of 200 devices (1 = abnormal, 0 = normal). When the FPGA detects that the data exceeds the threshold, the corresponding bit is set, triggering the CPU to respond in real time.
[0183] Furthermore, in one embodiment scenario, device management according to a preset strategy may include a virtual CPU management terminal analyzing and processing collected data through the preset strategy to monitor device status, provide fault warnings, and manage configuration. The preset strategy is designed based on "multi-dimensional data fusion + machine learning algorithms" and can be divided into three main modules: device status monitoring, fault warnings, and configuration management. Each module clearly defines the algorithm logic, considerations, and parameter thresholds to ensure the strategy is executable and optimizable. Specific details are as follows:
[0184] 1. Equipment status monitoring strategy
[0185] The core objective is to quantify the health status of equipment in real time, avoiding the one-sidedness of traditional "single-indicator monitoring," as explained below:
[0186] 1) Data input dimensions: Five core indicators were selected, and each indicator was assigned a differentiated weight (determined based on the Analytic Hierarchy Process, AHP):
[0187] ① Hardware temperature (weight 0.3): GPU threshold 85℃, NVMe threshold 70℃.
[0188] ②PCIe link bandwidth utilization (weight 0.4): threshold 90% (if exceeded, the link is congested).
[0189] ③ Error count (weight 0.15): Cumulative value of ECC errors and communication timeouts.
[0190] ④ Power consumption (weight 0.1): GPU threshold 300W, NVMe threshold 30W.
[0191] ⑤ Register read / write latency (weight 0.05): threshold 500ns (if exceeded, hardware response will be slow).
[0192] 2) Health Value Calculation Algorithm: This algorithm uses a "weighted summation and normalization" method for quantitative evaluation. The specific calculation formula is: Health Value = Σ(Actual Value of Indicator / Indicator Threshold × Weight) × 100. During the calculation, firstly, for each indicator, its actual measured value is compared with a preset threshold to obtain the completion ratio of that indicator, which is then multiplied by the corresponding weight coefficient to reflect its importance. Subsequently, the weighted results of all indicators are summed and standardized by multiplying by 100, ensuring the final result falls within the score range of 0-100. This score is further mapped to health status according to the following rules: 80-100 points indicate normal system operation, 60-80 points trigger a warning state, and below 60 points is judged as a fault state, thus achieving an intuitive and hierarchical evaluation of the system's health status.
[0193] 3) Execution process: The CPU reads 5 types of indicators from the cache layer every 50ms, runs the algorithm to calculate the health value, and writes it to the device information register and NVMe log layer in real time. Administrators can view the health value curve in real time through the cluster platform.
[0194] 2. Fault Early Warning Strategy
[0195] Overcoming the lag of traditional "threshold triggering," this solution employs "LSTM neural network + historical data training" to achieve predictive early warning. The solution is described below:
[0196] - Training data preparation: Extract fault data from the past 6 months (including 1000+ GPU ECC errors and NVMe disk failure cases) from the NVMe log layer. Each data entry contains "time series sequence of 5 types of indicators in the previous 10 minutes + fault type" to build the training dataset.
[0197] -LSTM model parameters: input layer dimension 5 (corresponding to 5 categories of indicators), hidden layer 2 layers (64 neurons per layer), output layer dimension 3 (normal / warning / fault probability), training iterations 1000 times, loss function uses cross-entropy, model accuracy reaches over 95%.
[0198] - Early warning execution logic: Every 2 minutes, the CPU inputs the time series sequence of the indicators from the most recent 10 minutes into the LSTM model. If "fault probability > 80%", then:
[0199] 1) Set the corresponding bit in the interrupt status register to trigger a response from the RTX system within 1ms;
[0200] 2) Generate early warning information (including device coordinates, predicted fault type, and suggested measures), encrypt it using SM4, and send it to the cluster management platform;
[0201] 3) Automatically back up the device's current configuration to NVM to prepare for subsequent recovery.
[0202] 3. Configuration Management Strategy
[0203] The "template-based + dynamic adjustment" mechanism solves the problem of low efficiency in the traditional "device-by-device configuration" method. Details are as follows:
[0204] - Configuration template design: Pre-set templates for different device types (such as GPU computing template, NVMe storage template), the templates include:
[0205] 1) Basic parameters: PCIe link speed (GPU is set to PCIe 4.0x16, NVMe is set to PCIe 4.0x4).
[0206] 2) Performance parameters: GPU frequency (compute template 2.4GHz, power saving template 1.8GHz), NVMe queue depth (storage template 256).
[0207] 3) Security parameters: Register access permissions (only virtual CPU can write).
[0208] - Dynamic adjustment logic: When the monitoring policy detects that "GPU bandwidth utilization is <30% for 5 consecutive minutes", the CPU will automatically execute:
[0209] 1) Read the "GPU power saving template" from NVM, encrypt it with SM4, and write it to the policy configuration register.
[0210] 2) Generate configuration modification instructions and write them to the GPU's device resource and topology management registers via PCIe DMA.
[0211] 3) Record modification logs (including parameters before and after modification, execution time) to NVMe for easy auditing and traceability.
[0212] - Batch configuration function: Supports selecting "same-domain devices" (such as 12 GPUs in the second domain) through the cluster platform. The CPU simultaneously distributes the configuration through the "broadcast write" mechanism, reducing the batch configuration time from the traditional 30 minutes to 2 minutes.
[0213] To enable those skilled in the art to better understand the embodiments provided by the present invention, the system operation and coordination process will be further described below.
[0214] 1. Initialization phase (within 10 seconds after system startup):
[0215] - The FPGA scans 200 devices in the cluster via the PCIe bus, obtains the DNA identifier of each device, and writes it to the cache layer via DMA.
[0216] - The CPU reads the preset strategy template and initial configuration from NVM, loads the LSTM model into memory, and completes hardware and software initialization.
[0217] 2. Real-time operation phase:
[0218] - The FPGA collects device metrics every 10ms and writes them to the device information register and cache layer via DMA.
[0219] The CPU executes a cycle strategy of "monitoring (50ms interval) → alerting (2-minute interval) → configuration adjustment (on demand)," and data interaction is efficiently completed through the AXI-4 bus and PCIe DMA.
[0220] 3. Exception handling phase:
[0221] - If a fault is detected (health value < 60 or fault probability > 80%), the CPU will immediately trigger an alert and load the backup configuration in NVM to restore the device. The restoration time is ≤ 200ms.
[0222] The above-described solution, through its heterogeneous architecture and multi-dimensional strategies in the virtual CPU management terminal, demonstrates significant superiority over traditional PCIe management solutions in several aspects, as tested in experiments: First, at the hardware level, the embedded CPU + FPGA collaboration reduces data acquisition latency from 50ms to 10ms, and onboard tiered storage improves log read / write speed by 3 times. Second, at the strategy level, LSTM alerts detect faults 5-10 minutes earlier than traditional threshold triggers, with an accuracy rate exceeding 95%, and templated configuration improves batch operation efficiency by 93%. Third, in terms of resource utilization, the virtual CPU independently undertakes management tasks, reducing the main CPU's computing power utilization from the traditional 25% to below 5%, ensuring that cluster computing power is concentrated on business operations. Overall, it can support stable operation of 200+ devices 24 / 7, fully meeting the efficient, intelligent, and secure management requirements of large-scale computing clusters.
[0223] Furthermore, in one embodiment, the PCIe board independently runs the encoding management system, and manages and operates the standard PCIe peripherals on the bus according to encoding rules through the encoding management system. Specific details are as follows:
[0224] I. The PCIe board (i.e., the board where the virtual CPU management terminal is located) ensures the autonomous operation of the encoding management system through a "three-layer independent architecture," including:
[0225] 1. Hardware Independence: Onboard independent power supply module (12V / 5A), independent of the main CPU power supply. Integrated PCIe 5.0 Switch independent channel (x8 bandwidth 32GB / s), dedicated to communication with bus peripherals, avoiding the occupation of the main CPU's PCIe bandwidth resources.
[0226] 2. System independence: It runs a customized embedded Linux system (kernel version 5.15, with redundant modules removed), stored on an onboard 8GB eMMC (read-only partition to prevent tampering), with a boot time of ≤15s, and system stability is ensured by a hardware watchdog (timeout threshold of 60s).
[0227] 3. Software independence: The coding management system adopts a "modular architecture" which includes three major modules: device management (parsing EX-DCC rules), policy execution (calling early warning / configuration logic), and log auditing (storing operation records). The modules communicate with each other through IPC (latency ≤1ms) and do not depend on any process of the main CPU.
[0228] II. The specific process for management according to coding rules is as follows:
[0229] The system strictly follows the EX-DCC coding rules and implements full lifecycle management of peripherals in three steps:
[0230] 1. Device discovery and identity verification: The FPGA has a built-in PCIe enumeration engine that scans the bus peripherals every 30 seconds. It filters compliant devices by reading the device's "EX-DCC identifier" (0xAA) and verifies the peripheral's DNA identifier (RSA-2048 decryption). The verification time is ≤200μs / device, ensuring that only compliant devices are managed.
[0231] 2. Data Acquisition and Rule Matching: For peripherals that have passed verification, data such as temperature and bandwidth are acquired via DMA according to the EX-DCC register address mapping (e.g., device status register offset 0x48) (acquisition cycle 10ms). The data is matched against the threshold in the encoding rules in real time (e.g., GPU temperature ≤85℃). Data that exceeds the rules is marked as "pending processing".
[0232] 3. Command Issuance and Result Feedback: For "pending" data, the system generates EX-DCC compliant commands (such as frequency reduction commands), which are encrypted by the TPM 2.0 chip SM4 and issued through an independent PCIe channel. After the peripheral device executes the command, it writes the result to the feedback register. After the system receives the result, it records it to the onboard 4TB NVMe SSD (ZFS file system). The command response time is ≤50ms, and the entire process does not occupy the main CPU resources.
[0233] Through the above embodiments, the core of "hardware resource isolation + software stack independence + full-process coding drive" is achieved, breaking through the limitations of traditional reliance on main CPU management and realizing complete decoupling of peripheral management and business computing power.
[0234] Furthermore, in one embodiment, the PCIe card is logically recognized by the main CPU as a standard PCIe peripheral. The PCIe card (the card housing the virtual CPU management terminal) enables the main CPU to recognize it as a "normal peripheral" by precisely matching the PCIe standard device attributes.
[0235] 1. Device type definition: In the header area of the standard PCIe configuration space, set the "Class Code" to 0x0108 (standard storage controller class). This is a device type natively supported by the main CPU (such as Intel Xeon series) and can be recognized without additional drivers.
[0236] 2. Core Identifier Configuration: The Vendor ID uses 0x19E5 (custom compliant vendor code), and the Device ID is set to 0x2002 (standard storage controller ID), matching the device list in the main CPU driver library to ensure that it is classified as "identifiable storage peripheral" during enumeration.
[0237] 3. Basic Capability Declaration: In the standard PCIe capability register, only the "necessary capabilities of the storage controller" (such as SATA protocol compatibility and NVMe protocol support) are exposed, while core functions such as FPGA acceleration and EX-DCC management are hidden to avoid the main CPU triggering alarms for unknown devices.
[0238] Furthermore, to prevent the main CPU from interfering with the board's encoding management functions, an isolation mechanism for identification and management is implemented through "hardware address filtering + function permission division":
[0239] 1. Address Space Partitioning: The board's address space is divided into a "Main CPU Visible Area" (containing only the standard memory controller configuration space, address 0x0000-0x0FFF) and a "Management Dedicated Area" (EX-DCC register, policy algorithm memory, address 0x1000-0xFFFF). The FPGA has a built-in address filter, and the main CPU automatically returns an "invalid address" when accessing the "Management Dedicated Area".
[0240] 2. Driver interaction limitation: The main CPU can only read and write to the board's 4TB NVMe SSD (such as reading management log backups) through the native storage driver, and cannot access the processes and registers of the encoding management system to ensure that the management function runs independently.
[0241] 3. Enumeration process adaptation: When the main CPU starts up, the board simulates the enumeration response of the standard storage controller (such as sending the configuration space read / write ACK signal). The enumeration time is ≤10ms, which is consistent with traditional storage peripherals and does not affect the startup efficiency of the main CPU.
[0242] Through the standard peripheral logical profile and functional isolation of the above embodiments of the present invention, the following can be achieved: First, the main CPU does not need to install any custom drivers, and the native recognition rate is 100%, avoiding the stability risks caused by modifying the operating system kernel; Second, the board's encoding management function is completely hidden, and the main CPU can only perform basic storage interaction, which not only ensures management independence, but also integrates into the cluster as an "ordinary peripheral", adapting to various mainstream server hardware and significantly reducing the deployment difficulty of large-scale clusters.
[0243] Furthermore, in one embodiment, the PCIe board encodes standard PCIe peripherals according to the preset standard specification and writes the encoding into the onboard memory. This achieves standardization, encryption, and persistence of peripheral encoding through a "code generation + hierarchical storage writing mechanism based on the EX-DCC specification," solving the problems of traditional PCIe devices lacking unified encoding and easily lost identifiers. Specific details are as follows:
[0244] I. Encoding Generation Process Based on EX-DCC Specification
[0245] The PCIe board (the board where the virtual CPU management terminal is located) uses EX-DCC as the default standard and generates a unique peripheral code in three steps:
[0246] 1. Information Acquisition: The FPGA uses PCIe DMA technology to read the core information of the seven registers of the peripheral EX-DCC, including a 64-bit encrypted DNA identifier (device unique serial number), 16-bit topological coordinates (domain + hierarchical number), and 8-bit device type code (0x01 = GPU / 0x02 = NVMe). The acquisition time is ≤50μs / unit, ensuring data real-time performance.
[0247] 2. Encoding and Encryption: The embedded CPU combines the collected information into a 128-bit code according to the EX-DCC rule (field allocation: DNA identifier 64 bits + topology coordinates 16 bits + device type 8 bits + CRC32 check 32 bits + reserved 8 bits), and then encrypts it through the onboard hardware encryption engine (supporting SM4 algorithm) to generate an "encrypted encoded string" to prevent the code from being tampered with.
[0248] 3. Encoding Verification: The encrypted encoding is double-verified—first, the integrity of the encoding is verified using the CRC32 algorithm (error rate ≤ 0.001%), and then compared with the data after decryption of the external device DNA identifier to ensure that the encoding corresponds one-to-one with the device. If the verification fails, the data is re-collected and generated.
[0249] II. Encoding and writing to onboard memory can employ a "three-level storage collaboration" mechanism to adapt to different access requirements:
[0250] 1. Cache storage: The encrypted code that passes the verification is first written to 8GB DDR5 cache (address range 0x90010000-0x9001FFFF) to respond to the real-time query of the management system (access latency ≤100ns) and meet the requirements of high-frequency encoding calls.
[0251] 2. Hot data persistence: Every 5 minutes, the newly added / updated code in DDR5 is synchronized to 1TB SLC NAND (100,000 erase / write cycles) through the "incremental synchronization algorithm". Redundancy is built using RAID 5 array technology to prevent code loss.
[0252] 3. Cold data archiving: Every day at 2:00 AM, the encoding of SLC NAND exceeding 7 days is compressed (LZ4 algorithm, compression ratio 2:1) and archived to 8TB QLC NAND, retaining 30 days of archived records and supporting historical encoding traceability.
[0253] Through the EX-DCC standardized encoding and hierarchical storage writing of the above-described embodiments, the following technical effects are achieved: First, the 128-bit encrypted encoding ensures that each peripheral device identifier is unique and tamper-proof, solving the problem of dynamic address changes in traditional BDF; second, the three-level storage collaboration reduces the encoding access latency to within 100ns, and the archived retention is 30 days, meeting the needs of real-time management and traceability; third, the entire process of generation and writing is automated, requiring no manual intervention, with a 100% encoding generation success rate, adapting to large-scale cluster management of 200+ devices, improving the efficiency and security of encoding management.
[0254] Furthermore, in one embodiment, the encoding management system includes a code-behavior bidirectional mapping memory mechanism to coordinate and organize the allocation of device functions. This embodiment is described below:
[0255] I. The core architecture of the bidirectional mapping memory mechanism constructs a closed loop based on "encoding as the index, behavior as the variable, and memory as the link," specifically including:
[0256] 1. Mapping Table Construction: A bidirectional mapping table between 128-bit EX-DCC encoding and device behavior characteristics is established in the onboard 8GB DDR5 cache. The table entries contain "encoding ID (primary key), behavior feature vector (communication affinity, resource utilization, and other 5-dimensional data), function allocation weight (0-100), and update timestamp". An index is built using a hash algorithm (SHA-256), and the query time for a single table entry is ≤50ns.
[0257] 2. Hardware-triggered mapping: When the behavior of a peripheral device changes (such as a sudden 20% increase in GPU communication), its behavior measurement register (offset 88H) is automatically set to the "change flag bit". The FPGA on the virtual CPU management side captures this signal through a PCIe interrupt (1ms response) and triggers the mapping table update.
[0258] 3. Memory storage tiers: Short-term memory (mapping relationship of the most recent 1 hour) is stored in DDR5 cache, supporting high-frequency read and write; medium-term memory (mapping relationship of the most recent 7 days) is compressed and stored in 1TB SLC NAND; long-term memory (30 days) is archived in 8TB QLC NAND, and "hot sorting" is used to retain the mapping relationship of high-frequency interaction and eliminate low-frequency data (access volume <1 time / day).
[0259] II. The two-way interaction and function coordination process is as follows:
[0260] 1. Encoding-driven behavior: When a new task (such as AI training) is assigned, the system queries the function weight (90) corresponding to "GPU encoding (0x10000001...)" in the mapping table according to the task requirements (high GPU computing power + high NVMe bandwidth), and prioritizes scheduling high-weight devices, which can improve task startup efficiency by 30%.
[0261] 2. Behavioral feedback coding: During device operation, the LSTM neural network analyzes behavioral characteristics in real time (such as GPU bandwidth utilization >90% for 3 consecutive times), automatically increases its functional allocation weight (+10), and synchronously updates the "functional level field" of the mapping table and the device DNA identifier register, so that subsequent tasks are given priority in resource allocation.
[0262] 3. Anomaly Correction Mechanism: If the mapping relationship deviates from the actual behavior by more than 20% (e.g., the encoding indicates high affinity but the communication delay is high), the system triggers "mapping calibration", re-collects data to generate a new mapping, and the calibration time is ≤1s to ensure the accuracy of the mapping.
[0263] Through the bidirectional dynamic association between encoding and behavior in the above embodiments, device function allocation is upgraded from "static preset" to "intelligent adaptation," thereby improving task scheduling accuracy, reducing communication latency of high-affinity device groups, and increasing cluster resource utilization compared to traditional static allocation. Meanwhile, the hierarchical memory design balances real-time performance and storage efficiency, perfectly adapting to the large-scale cluster collaboration needs of 200+ devices.
[0264] In summary, the large-scale PCIe device coding management system provided by the above embodiments of the present invention can dynamically assign a unique identity code to specific device instances across the entire domain by configuring global device group coding on the PCIe board, thereby achieving rapid and accurate identification of device identities in large-scale, high-performance computing scenarios.
[0265] While numerous embodiments of the invention have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Many modifications, alterations, and alternatives will occur to those skilled in the art without departing from the spirit and essence of the invention. It should be understood that various alternatives to the embodiments of the invention described herein may be employed in the practice of the invention. The appended claims are intended to define the scope of protection of the invention and therefore cover equivalents or alternatives within the scope of these claims.
Claims
1. A large-scale PCIe device coding management system, characterized in that, The encoding management system runs on a PCIe board, which includes a virtual CPU management terminal and a PCIe board terminal. The PCIe board terminal is a standard PCIe device that conforms to a preset standard specification. The virtual CPU management terminal is installed on a standard PCIe bus and reads and writes to the PCIe board terminal through the computer bus.
2. The system according to claim 1, characterized in that, The PCIe board is configured as onboard memory, and the virtual CPU management terminal is configured as a board that conforms to the PCIe specification.
3. The system according to claim 1, characterized in that, The preset standard specification is a group coding system for all devices.
4. The system according to claim 3, characterized in that, The global device group coding rules include: Device declaration area, standard PCIe configuration space header area compatibility layer, standard PCIe capability and status register group, encrypted heritable device DNA identifier register, device resource and topology management register group, device history status and reliability management register group, device communication affinity and behavior measurement register group.
5. The system according to claim 1, characterized in that, The virtual CPU management terminal has a built-in embedded CPU and onboard extended memory. The embedded CPU reads or writes PCIe board-side device information and operation logs on the bus into the onboard extended register to manage devices according to a preset strategy.
6. The system according to claim 5, characterized in that, The device management according to the preset strategy includes the virtual CPU management terminal analyzing and processing the collected data through the preset strategy in order to monitor device status, provide fault warnings, and manage configurations.
7. The system according to claim 1, characterized in that, The PCIe board runs the encoding management system independently and manages and operates the standard PCIe peripherals on the bus according to the encoding rules through the encoding management system.
8. The system according to claim 1, characterized in that, The PCIe card is logically recognized by the main CPU as a standard PCIe peripheral.
9. The system according to claim 1, characterized in that, The PCIe board encodes the standard PCIe peripherals according to the preset standard specifications and writes the encoding into the onboard memory.
10. The system according to claim 1, characterized in that, The encoding management system includes a bidirectional encoding-behavior mapping memory mechanism to coordinate and organize the allocation of device functions.