An AI supernode interconnection exchange chip port dynamic grouping management method and system

CN122640370BActive Publication Date: 2026-09-29SHENZHEN NANFEI MICROELECTRONICS CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202611113477.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-27
Publication Date
2026-09-29
Estimated Expiration
2046-07-27

AI Technical Summary

Technical Problem

但该方法存在以下不足:一是软件栈复杂、部署成本较高,需要配套包含多个组件的管理框架,难以适用于轻量化设备或资源受限的嵌入式场景;二是需要与硬件平台深度绑定,依赖特定芯片内部架构,无法应用到其他互联交换芯片,通用性差;三是缺乏高效的跨域合并操作,无法满足动态调度对快速响应的要求

Benefits of technology

[0038]依据上述实施例的一种AI超节点互联交换芯片端口动态分组管理方法及系统,通过双向映射数据结构建立端口与计算域之间的双向索引,实现低时间复杂度的快速双向查询;通过默认计算域设计,利用AI超节点互联芯片上电后端口通信权限默认全互通的硬件特性,在软件初始化时不修改任何硬件寄存器,实现初始化零开销;通过引入并查集来管理端口之间的集合关系,支持路径压缩和按秩进行合并优化,实现计算域的快速合并,可大幅降低时间复杂度,提高了资源调度的灵活性;且仅依赖少量数组和基本的并查集操作,无需复杂的独立软件组件,易于集成到嵌入式环境或现有系统中,实现软件栈轻量化;基于通用的端口通信权限配置设计,可适配任何支持端口级隔离的互联交换芯片,提升了端口分组管理的通用性和可移植性,进而实现了完整、高效且可移植的AI超节点互联交换芯片端口动态分组管理。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122640370B_ABST
    Figure CN122640370B_ABST
Patent Text Reader

Abstract

The application discloses an AI supernode interconnection exchange chip port dynamic grouping management method and system, and belongs to the technical field of network communication. The application obtains a bidirectional mapping table, the bidirectional mapping table is used for representing bidirectional indexes between ports and computing domains, so as to determine a computing domain to which each port belongs and a port list corresponding to each computing domain; a union-find set is obtained and searched, and the union-find set is used for maintaining a set relationship of the ports in different computing domains; and according to a received user operation instruction, the bidirectional mapping table and / or the union-find set are updated or queried, so as to perform port dynamic grouping management. The bidirectional mapping data structure is used to realize low-time-complexity fast bidirectional query; the default computing domain is used to realize initialization zero cost; the union-find set is used to realize fast merging of the computing domains and software stack light weight; and based on port communication permission configuration, the universality and portability of the grouping management are improved, and efficient AI supernode interconnection exchange chip port dynamic grouping management is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of network communication technology, specifically to a method and system for dynamic group management of AI supernode interconnection switching chip ports. Background Technology

[0002] In the fields of artificial intelligence and high-performance computing, large model training relies on clusters built from massive amounts of GPUs or AI accelerators. Among them, scale-up networks interconnect dozens or even hundreds of computing units (GPUs) through interconnection switching chips to build a supercomputing node, generally called a supernode or AI supernode, so that it appears as a unified supercomputing unit to meet the needs of high bandwidth and low latency AI training.

[0003] AI supernode interconnect switching chips support on-network computing (that is, some computing tasks executed by GPUs are offloaded to the interconnect switching chip, such as some Reduce communication operations). With the expansion of cluster scale and the emergence of scenarios such as multi-tenant sharing, on-demand task partitioning, and fault isolation, AI supernode interconnect switching chips are also required to support dynamic port grouping management in order to allocate computing resources.

[0004] Existing methods create one or more computing domains (i.e., user partitions), with partition configuration managed through a software stack, achieving grouping and isolation of interconnect switch chip ports through hardware and software collaboration. However, this method has the following shortcomings: First, the software stack is complex and has high deployment costs, requiring a management framework containing multiple components, making it difficult to apply to lightweight devices or resource-constrained embedded scenarios; second, it requires deep integration with the hardware platform, depends on the specific chip's internal architecture, and cannot be applied to other interconnect switch chips, resulting in poor versatility; third, it lacks efficient cross-domain merging operations, failing to meet the rapid response requirements of dynamic scheduling. Summary of the Invention

[0005] The main technical problem this invention addresses is how to achieve efficient dynamic group management of AI supernode interconnection switching chip ports.

[0006] According to the first aspect, one embodiment provides a method for dynamic group management of AI supernode interconnection switching chip ports, including:

[0007] Obtain a bidirectional mapping table, which is used to represent the bidirectional index between ports and computing domains; the bidirectional mapping table includes a first mapping table and a second mapping table, the first mapping table is used to determine the computing domain to which each port belongs; the second mapping table is used to determine the port list corresponding to each computing domain;

[0008] Obtain a disjoint-set data structure, which is used to maintain the set relationships of ports in different computing domains;

[0009] Based on the received user operation instructions, the bidirectional mapping table and / or disjoint-set data structure are updated or queried to perform dynamic port grouping management. The user operation instructions include query instructions and port ownership modification instructions. When the received user operation instruction is a port ownership modification instruction, the bidirectional mapping table and / or disjoint-set data structure are updated according to the port ownership modification instruction, and the communication permission configuration corresponding to the port whose port ownership has changed is updated to respond to the port ownership change. When the received user operation instruction is a query instruction, the query result is obtained according to the bidirectional mapping table and / or disjoint-set data structure.

[0010] In one embodiment, the disjoint-set data structure includes a root mapping table and set information of all port sets corresponding to all computational domains; wherein, the set information includes the rank of the port set and the parent node corresponding to each port; the root mapping table is used to characterize the mapping relationship between each root node in the disjoint-set data structure and its respective computational domain;

[0011] For any computing domain: if the computing domain is not the default computing domain, the computing domain corresponds to a set of ports; if the computing domain is the default computing domain, each port in the default computing domain corresponds to a set of ports, the rank of each set of ports is a preset initial value, each port is the root node of its corresponding set of ports, and the parent node of the port is the port itself.

[0012] In one embodiment, the port ownership modification instruction includes instructions for creating a computing domain, merging computing domains, destroying a computing domain, adding a port to a specified computing domain, or deleting a port from a specified computing domain; the query instruction includes instructions for querying the computing domain to which a target port belongs, querying ports in a target computing domain, or determining whether multiple ports belong to the same domain.

[0013] Specifically, when the query instruction is a multi-port same-domain judgment instruction, the root node of the port set corresponding to each port is obtained, and the port is judged to belong to the same computing domain based on the root node of the port set corresponding to each port, so as to obtain the corresponding query result.

[0014] In one embodiment, before updating the bidirectional mapping table and / or disjoint-set data structure according to the received user operation instruction (which is a port ownership modification instruction), a validity check of the port and / or computing domain is required. If the validity check fails, the execution of the port ownership modification instruction is stopped; if the validity check passes, the port ownership modification instruction is executed. This includes:

[0015] When the port ownership modification instruction is an instruction to create a computing domain, if the computing domain identifier corresponding to the created computing domain is not occupied and all ports to be added to the computing domain belong to the default computing domain, the legality verification is considered to have passed;

[0016] When the port ownership modification instruction is a merge computing domain instruction, if none of the computing domains to be merged are the default computing domains, the legality check is considered to have passed.

[0017] When the port ownership modification instruction is an instruction to destroy a computing domain, if the destroyed computing domain is not the default computing domain, the legality check is considered to have passed.

[0018] When the port ownership modification instruction is an instruction to add a port to a specified computing domain, if the specified computing domain exists and the port to be added belongs to the default computing domain, the legality verification is considered to have passed.

[0019] When the port ownership modification instruction is an instruction to delete a port from a specified computing domain, if the specified computing domain exists and the deleted port belongs to the specified computing domain, the legality verification is considered to have passed.

[0020] In one embodiment, updating the bidirectional mapping table and / or disjoint-set data structure according to the port ownership modification instruction includes:

[0021] Update the bidirectional mapping table according to the port ownership modification instruction;

[0022] If the port ownership modification instruction changes the root node in the disjoint-setup, the root mapping table is updated;

[0023] When the port ownership modification instruction is an instruction to create a computing domain, an instruction to merge computing domains, or an instruction to add a port to a specified computing domain, it also includes a port merging operation based on a disjoint-set data structure to update the disjoint-set data structure. The port merging operation includes:

[0024] Obtain the ports to be merged required for the port merging operation; determine the rank of the port set corresponding to each port to be merged; determine the first computation domain based on the rank of the port set corresponding to each port to be merged, and use another computation domain as the second computation domain; merge the port set corresponding to the second computation domain into the port set corresponding to the first computation domain to obtain the merged port set.

[0025] In one embodiment, obtaining the ports to be merged required for the port merging operation includes:

[0026] When the port ownership modification instruction is an instruction to create a computing domain, a port list consisting of all ports allocated to the computing domain is obtained, a target port is determined according to the port list, and the target port is merged with other ports in the port list. In each port merging operation, the target port and any other port other than the target port are used as the port to be merged for the port merging operation.

[0027] When the port ownership modification instruction is an instruction to add a port to a specified computing domain, and a port needs to be added to the specified computing domain, the port to be added is a port to be merged, and any port in the specified computing domain is another port to be merged required for the port merging operation.

[0028] When the port ownership modification instruction is a merge computing domain instruction, for the two computing domains to be merged: any port in each computing domain is used as the port to be merged for the port merging operation.

[0029] In one embodiment, merging the port set corresponding to the second computing domain into the port set corresponding to the first computing domain to obtain the merged port set includes:

[0030] Obtain the root node of the port set corresponding to each port to be merged; merge the port set corresponding to the second computing domain and the port set corresponding to the first computing domain according to the root node corresponding to each port to be merged, including: updating the parent node of the root node corresponding to the second computing domain to the root node corresponding to the first computing domain, and updating the rank corresponding to the merged port set.

[0031] In one embodiment, obtaining the root node of the port set corresponding to each port to be merged includes:

[0032] For any port to be merged: obtain the parent node of the port to be merged, and determine whether the parent node of the port to be merged is the port itself; if so, the port to be merged is the root node; otherwise, the port to be merged is not the root node, and continue to traverse upwards to determine whether the parent node is the root node, until a port is found whose parent node is the port itself. At this time, the obtained port is taken as the root node of the set of ports corresponding to the port to be merged.

[0033] In one embodiment, after obtaining the root node, the method further includes: performing path compression optimization on each port, wherein the parent node of each port in the traversal path is updated to the root node, so that the parent node corresponding to each port is the root node.

[0034] In one embodiment, in the initial bidirectional mapping table, all ports belong to the default computing domain; wherein, in the initial first mapping table, the computing domain to which each port belongs is the default computing domain; and in the initial second mapping table, the port list corresponding to the default computing domain contains all ports.

[0035] According to the second aspect, one embodiment provides an AI supernode interconnection switching chip port dynamic group management system, including:

[0036] Memory, used to store programs;

[0037] And a processor, used to implement the above-described AI supernode interconnection switching chip port dynamic grouping management method by executing the program.

[0038] According to the above embodiments, an AI supernode interconnect switching chip port dynamic grouping management method and system establish a bidirectional index between ports and computing domains through a bidirectional mapping data structure, realizing fast bidirectional queries with low time complexity. Through a default computing domain design, leveraging the hardware characteristic of the AI ​​supernode interconnect chip's default full interoperability of port communication permissions after power-on, no hardware registers are modified during software initialization, achieving zero initialization overhead. By introducing a disjoint-set data structure to manage the set relationships between ports, path compression and rank-based merging optimization are supported, enabling rapid merging of computing domains, significantly reducing time complexity and improving resource scheduling flexibility. Furthermore, it relies only on a small number of arrays and basic disjoint-set operations, requiring no complex independent software components, making it easy to integrate into embedded environments or existing systems, achieving a lightweight software stack. Based on a universal port communication permission configuration design, it can be adapted to any interconnect switching chip that supports port-level isolation, improving the universality and portability of port grouping management, thus realizing complete, efficient, and portable AI supernode interconnect switching chip port dynamic grouping management. Attached Figure Description

[0039] Figure 1 A flowchart illustrating a method for dynamic group management of ports on an AI supernode interconnection switching chip;

[0040] Figure 2 This is a schematic diagram of a two-way mapping table;

[0041] Figure 3 This is a schematic diagram of the tree structure corresponding to the disjoint-set data structure in its initial state.

[0042] Figure 4 Here is a flowchart of the port merging operation;

[0043] Figure 5 A schematic diagram illustrating the merging of two computational domains;

[0044] Figure 6 This is a system block diagram of a dynamic grouping management system for the interconnection and switching chip ports of an AI supernode. Detailed Implementation

[0045] The present invention will now be described in further detail with reference to specific embodiments and accompanying drawings. Similar elements in different embodiments are referred to by associated similar element reference numerals. In the following embodiments, many details are described to facilitate a better understanding of the invention. However, those skilled in the art will readily recognize that some features may be omitted in different situations, or may be replaced by other elements, materials, or methods. In some cases, certain operations related to the present invention are not shown or described in the specification. This is to avoid obscuring the core parts of the invention with excessive description. For those skilled in the art, detailed description of these related operations is not necessary; they can fully understand the related operations based on the description in the specification and general technical knowledge in the art.

[0046] Furthermore, the features, operations, or characteristics described in the specification can be combined in any suitable manner to form various embodiments. At the same time, the steps or actions in the method description can be rearranged or adjusted in a manner obvious to those skilled in the art. Therefore, the various orders in the specification and drawings are only for the clear description of a particular embodiment and do not imply a necessary order, unless otherwise stated that a particular order must be followed.

[0047] The serial numbers assigned to components in this document, such as "first" and "second," are used only to distinguish the described objects and have no sequential or technical meaning. The terms "connection" and "linkage" used in this invention, unless otherwise specified, include both direct and indirect connections (linkages).

[0048] AI supernode: refers to a high-performance computing unit formed by tightly interconnecting multiple GPUs or accelerators through high-speed interconnection switching chips in artificial intelligence training or inference scenarios.

[0049] Scale-up networks (vertical scaling networks) refer to network architectures that tightly couple multiple computing units through high-speed interconnect switching chips, making them appear to the outside world as a single supercomputing unit.

[0050] Computation domain: also known as a computing group, is a logical group of ports divided on a physical interconnect switching chip. Ports within the same computing domain can communicate with each other, while communication between ports in different computing domains is prohibited, thereby achieving hardware-level isolation.

[0051] Port communication permission configuration: refers to a set of control information (port permission configuration information) maintained by each port, which indicates which target ports the port is allowed to communicate with.

[0052] In scale-up networks, the AI ​​supernode interconnect switch chip is the core hub connecting all computing units. For security, performance isolation, and independent resource scheduling, the physical ports on the interconnect switch chip need to be divided into multiple computing domains. This allows ports within the same computing domain to communicate with each other, while ports in different computing domains are hardware isolated. However, as computing clusters grow in scale, scenarios such as multi-tenant sharing, on-demand task partitioning, and fault domain isolation emerge, placing higher demands on the flexible management of computing resources.

[0053] Existing port management solutions typically rely on the coordinated operation of hardware and software, such as NVLink port partitioning technology. The interconnect switching chip uses a fully interconnected cross-connect matrix, meaning any input port can directly connect to any output port. In group management, one or more user partitions (i.e., computing domains) can be created. Each partition contains a set of ports, and a port can only belong to one partition at a time. Ports in different partitions are isolated at the data link layer and cannot communicate directly. At the management architecture level, partition configuration is uniformly managed through a dedicated management software stack running on the host, allowing for dynamic creation and deletion of partitions, as well as querying the current partition status. However, existing technical solutions have the following main drawbacks:

[0054] 1) Existing technical solutions have overly complex software stacks and high deployment costs. Port grouping management typically requires a complete software management framework containing multiple independent components, relying on a complete operating system and runtime environment. This results in high deployment costs and limited applicability of the technical solution in resource-sensitive scenarios such as lightweight devices or dedicated supernodes.

[0055] 2) Existing technical solutions rely on dedicated hardware, which has poor versatility and portability. Their management logic usually depends on the interconnect switching chips and supporting proprietary interconnect protocols of specific manufacturers. For example, NVLink port partitioning technology needs to be deeply bound to the hardware design of NVSwitch and tightly coupled with the NVLink interconnect protocol. It cannot be reused on interconnect switching chips of other manufacturers and is difficult to adapt to open interconnect protocols (such as UALink).

[0056] 3) Existing technical solutions lack efficient cross-domain merging operations. In actual scheduling, it is often necessary to merge two or more independent computing domains into one to integrate scattered computing power to meet the needs of large-scale training tasks. However, existing solutions do not provide direct cross-domain merging operations. Usually, it is necessary to delete the original group configuration first, then recreate the new computing domain and migrate the configuration port by port. This involves a large number of register read and write operations, which is inefficient and difficult to meet the requirements of rapid response in dynamic scheduling scenarios.

[0057] In summary, existing technical solutions are inadequate in terms of lightweight design, portability, and integration efficiency.

[0058] This invention provides a completely open software management algorithm that does not depend on any specific hardware vendor by introducing bidirectional mapping data structures, default computing domain design, and disjoint-setup acceleration merging techniques. Its core data structure and management process can be directly implemented by those skilled in the art, achieving complete, efficient, and portable dynamic group management of AI supernode interconnection switching chip ports. Specifically, a bidirectional index is established between ports and computing domains through a bidirectional mapping data structure, enabling fast bidirectional queries with low time complexity. A default computing domain design leverages the hardware characteristic of fully interconnected ports by default after power-on of the AI ​​supernode interconnect switching chip, achieving zero initialization overhead by not modifying any hardware registers during software initialization. The introduction of a disjoint-set data structure to manage the set relationships between ports supports path compression and rank-based merging optimization, relying only on a small number of arrays and basic disjoint-set operations. The code is concise, requiring no complex independent software components, and is easily integrated into embedded environments (bare metal environments) or existing systems, achieving a lightweight software stack. Furthermore, the disjoint-set data structure enables rapid merging of computing domains, with a time complexity far lower than traditional traversal schemes, improving resource scheduling flexibility. In addition, based on a universal port communication permission configuration design, it is adaptable to any interconnect switching chip that supports port-level isolation, such as the UALink interconnect switching chip, enhancing the versatility and portability of port group management.

[0059] Please refer to Figure 1 Some embodiments provide a method for dynamic group management of AI supernode interconnection switching chip ports, including:

[0060] By combining three data structures—bidirectional mapping tables, disjoint-set data structures, and root mapping tables—the system responds uniformly to query operation commands and any changes in port ownership through the following general steps, while synchronizing hardware communication permission configurations.

[0061] Step S100: Obtain a bidirectional mapping table, which is used to represent the bidirectional index between the port and the computing domain.

[0062] The bidirectional mapping table in this embodiment consists of two logically inverse mappings, essentially two sides of the same management structure; specifically, it includes a first mapping table and a second mapping table, wherein the first mapping table is used to determine the computing domain to which each port belongs. For example, the bidirectional mapping table in this embodiment is as follows: Figure 2As shown, the first mapping table can be set as an array indexed by port number. Each value in the array is the computing domain identifier (computing domain ID) of the computing domain to which the corresponding port belongs. Therefore, the computing domain identifier corresponding to the computing domain to which the port belongs is determined based on the port number of each port. For example, the i-th value in the array is the computing domain identifier of the computing domain to which port i belongs. The first mapping table can determine which computing domain a port belongs to in O(1) time. The second mapping table is used to determine the list of ports contained in each computing domain. For example, the second mapping table can be set as an array indexed by computing domain identifier. This array stores all ports (port list) contained in each computing domain. Therefore, the port number corresponding to all ports contained in the computing domain is determined based on the computing domain identifier corresponding to each computing domain. The second mapping table can answer which ports are contained in a computing domain in O(1) time.

[0063] In this embodiment, the first mapping table and the second mapping table are always kept consistent. When a port moves from computing domain A to computing domain B, the computing domain identifier corresponding to the port in the first mapping table is modified to the computing domain identifier corresponding to computing domain B. The port is deleted from the port list corresponding to computing domain A in the second mapping table and added to the port list corresponding to computing domain B. The bidirectional indexing mechanism ensures that state queries in any direction do not require traversal, thereby achieving constant-level query efficiency. In this embodiment, by maintaining the first mapping (port → computing domain) and the second mapping (computing domain → port list) simultaneously, both querying the domain to which the port belongs and obtaining the port list within the domain can achieve O(1) time complexity. The bidirectional indexing mechanism enables efficient querying. If the mapping in either direction is missing, a certain type of query will have to be completed by traversal.

[0064] It should be noted that this embodiment utilizes the hardware characteristic of the AI ​​supernode interconnect chip, which allows for full interoperability of port communication permissions by default after power-on, to design a default computing domain (with a corresponding computing domain identifier of 0). In the initial bidirectional mapping table, all ports belong to the default computing domain, and the computing domain identifier corresponding to the default computing domain is set to 0. At this time, in the first mapping table, the computing domain identifiers corresponding to all ports are 0, and in the second mapping table, the port list corresponding to the default computing domain contains all ports. Since all ports belong to the default computing domain in the initial state, the software data structure is completely consistent with the full interoperability state of the hardware. Therefore, no hardware registers are modified during software initialization, which can directly achieve zero hardware overhead in the initialization process and also provide a natural starting point for the subsequent dynamic creation of computing domains.

[0065] For example, if the total number of system ports is N, a first mapping table Port2Group[N] is created. The subscript of this array corresponds to the physical port number. Each entry is used to store the computing domain identifier corresponding to the computing domain to which the corresponding port belongs. At this time, the initial values ​​of all entries in the first mapping table are all 0. At the same time, a second mapping table Group2Ports is created. In this table, a corresponding entry Group2Ports[0] is created for the default computing domain, and all port numbers [0, 1, 2, ..., N-1] are filled into it. Thus, a bidirectional index between the port and the default computing domain is established, and the initial bidirectional mapping table is obtained.

[0066] Furthermore, the first and second mapping tables can be implemented using different data structures. For example, when the number of ports changes dynamically or the port numbers are not consecutive, a linked list can be used instead of an array, but the query performance is not as good as that of an array.

[0067] Step S110: Obtain a disjoint set, which is used to maintain the set relationship of ports in different computing domains.

[0068] Due to dynamic changes in task scale, tenant isolation, faults, and scheduling strategies, the requirements for GPU interoperability may continuously change. For example, to isolate different tasks / users, or to open a batch of high-speed GPU interconnect channels on demand, the communication relationships between ports may change. Furthermore, factors such as the variable computing power requirements of AI tasks, hardware sharing and reuse, automated scheduling, and fault self-healing will lead to continuous adjustments to the boundaries of the interconnected GPU set. Therefore, in practical applications, it is necessary to modify port communication permissions and / or adjust computing domain boundaries accordingly based on the actual situation. This embodiment introduces a disjoint-set data structure to manage the set relationships between ports, utilizing its Fin... The d and Union operations enable efficient merging of computation domains and determination of port co-domains. Among them, the disjoint set is a tree-shaped data structure used to maintain the query and merging of disjoint sets. In this invention, each port is regarded as an element. All ports in the same computation domain (non-default computation domain) are structurally represented as a tree. All ports contained in this tree constitute a port set. In the default computation domain, the parent node of each port is itself. At this time, each port forms a single-node tree. The ports contained in each single-node tree also constitute a port set. That is, each port in the default computation domain corresponds to a port set.

[0069] In this embodiment, the disjoint-set data structure includes a root mapping table and set information of all port sets corresponding to all computing domains. The set information includes the rank of the port set and the parent node corresponding to each port. The root mapping table is used to characterize the mapping relationship between each root node in the disjoint-set data structure and its corresponding computing domain. It is an auxiliary data structure that maps the root nodes in the disjoint-set data structure to the corresponding computing domain identifiers and is used to connect the disjoint-set data structure and the bidirectional mapping.

[0070] For example, the parent node corresponding to each port can be represented by a Parent array, the number of values ​​stored in it (i.e., the array length) is the total number of ports of the interconnect switching chip, where Parent[i] represents the port number corresponding to the parent node of port i. In this embodiment, if the parent node corresponding to a port is itself, that is, Parent[i] equals i, then port i is considered to be the root node, and port i can represent the entire port set; if Parent[i] is not equal to i, then port i is considered not to be the root node.

[0071] The rank of all port sets in the disjoint-set data structure can be set as a Rank array, the number of values ​​stored in it (i.e., the array length) is the total number of ports of the interconnect switching chip, where Rank[i] represents the height (i.e., rank) of the tree rooted at port i. However, this array is only meaningful for the root node, and the Rank values ​​of non-root nodes are undefined. In other words, in this embodiment, it is only necessary to use the root node to represent the entire port set, and then determine the rank corresponding to the port set (or computation domain) according to the value stored in the Rank array corresponding to the root node, without having to establish a correspondence between each port in the port set and the rank, thereby saving the space cost required for storage.

[0072] It should be noted that the Parent and Rank arrays in a disjoint-set data structure can be replaced by other equivalent data structures, such as using linked lists to represent port sets or using balanced trees to maintain set relationships. The Rank array can also store the size of different port sets, allowing for optimization during subsequent merging based on the size of the corresponding port set. However, these alternatives are generally less efficient in terms of time complexity or space compared to a disjoint-set data structure implemented using arrays.

[0073] In this embodiment, the disjoint-set data structure supports efficient computation domain merging and same-domain determination through two core operations. One core operation, Find(x), searches for the root node of the port set containing port x to determine the root node that can represent the port set (i.e., the representative port). This allows the direction of computation domain merging to be determined (i.e., for two computation domains to be merged, which computation domain's ports need to be adjusted). When determining whether multiple ports belong to the same computation domain (same-domain determination), if the root node of a port can be directly determined, the determination is made based on whether the root nodes are the same. When the root nodes are different, the determination can be further made based on whether the computation domains to which the root nodes belong are the same. For example, if two ports have different root nodes, but both their corresponding root nodes belong to the default computation domain, these two ports still belong to the same computation domain. Compared to determining the computation domain to which each port belongs by traversal, determining the same domain of multiple ports by using the root node can significantly reduce the time complexity of query operations and improve query efficiency. This advantage is particularly significant in scenarios where the number of ports on a supernode chip is large.

[0074] Since the ownership of ports can change, the root node in the set of ports corresponding to a computing domain can also change. Therefore, the root node cannot usually be determined directly. In this embodiment, the root node can be queried through the Parent array. For example, after determining that port i is not the root node, we record the parent node corresponding to port i as port a, and continue to traverse upwards to obtain the parent node of port a. By judging whether the parent node of port a is itself, we can determine whether port a is the root node. By tracing upwards, we can find a port whose parent node is itself, and thus obtain the root node.

[0075] Another core operation is the port merging operation Union(x, y), which merges the sets containing port x and port y. For example, suppose there are two independent computing domains that respectively carry two isolated GPU resources, and there is a hardware block between their respective ports, preventing cross-domain communication. However, when the business experiences insufficient computing power and needs to integrate idle computing power to meet the demand for large-scale model training, the existing distributed computing power cannot meet the needs of parallel communication of the model. At this time, it is necessary to update the port isolation rules through the interconnection switching chip to integrate the two into a single interconnected computing domain, so as to achieve unimpeded collaborative computing of all GPUs.

[0076] In this embodiment, when merging computational domains, the rank of the set of ports corresponding to each port is determined by the root node corresponding to that port. Trees with smaller ranks are merged into trees with larger ranks to avoid excessive growth of tree height and ensure that the tree height is controlled within O(logn). This ensures the efficiency of subsequent search operations. Even if a large number of sets are merged repeatedly, the average time complexity of each Find operation can still be kept at O(αn) (αn is the inverse function of the Ackerman function, which can be regarded as a constant for the actual number of ports). Here, n in the time complexity represents the scale variable and is not a fixed number.

[0077] Since the root node in the disjoint-set data structure corresponds to the port number, if we want to know which computing domain this root node represents, in addition to traversing the bidirectional mapping table, we can also maintain an additional root mapping table (Root2Group) to quickly determine the computing domain to which the root node belongs.

[0078] For example, assuming r is the root node of a port set, then Root2Group[r] represents the computing domain identifier corresponding to the root node. By mapping the physical root nodes in the disjoint-setup to the logical computing domain, the root node corresponding to the port can be quickly located when querying the computing domain to which the port belongs or merging computing domains, and then the corresponding computing domain identifier can be determined through the root mapping table.

[0079] It should be noted that in the initial disjoint-set data structure, an initial root mapping table and set information (Rank array and Parent array) are created. At this time, the initial array length of both arrays is the total number of ports. Since only the default computation domain exists in the initial state, and each port forms its own single-node tree in the default computation domain, for any port i, its parent node Parent[i] = i, and the corresponding Rank[i] = 0. Since each port is the root node at this time, in the initial disjoint-set data structure, each port is regarded as an independent port set, and the root mapping table maps the root nodes of these sets to the default computation domain. At this time, for port i, the value (i.e., the computation domain identifier) ​​corresponding to it in the initial root mapping table is 0. If the port's affiliation changes later, for example, if an independent port (root node) in the default computation domain becomes a child node of another port, causing the port to no longer be the root node, the root mapping table needs to be updated to set the corresponding value of the port in the root mapping table to an invalid value.

[0080] In this embodiment, for any computing domain: if the computing domain is not the default computing domain, the computing domain corresponds to a set of ports; if the computing domain is the default computing domain, each port in the computing domain corresponds to a set of ports, and the rank of each set of ports is a preset initial value. For example, the preset initial value can be set to 0, and each port is the root node of the corresponding set of ports. At this time, the parent node of the port is the port itself.

[0081] Typically, supernode interconnect switching chips have multiple ports (e.g., 72 or 144), each of which can connect to a GPU or accelerator. Here, we'll use an 8-port example. The initial disjoint-set data structure is as follows: Figure 3 As shown, assuming the disjoint-set data structure initially contains ports {0, 1, 2, 3, 4, 5, 6, 7}, besides the default computation domain, there are two other computation domains, Group1 and Group2. Group1 corresponds to the port set {0, 1, 2}, with port 0 as the root node, and both port 1 and port 2 have port 0 as their parent node, resulting in a tree height (rank) of 1. Group2 corresponds to the port set {3, 4, 5}, with port 3 as the root node, and both port 4 and port 5 have port 3 as their parent node, resulting in a tree height (rank) of 1. Ports 6 and 7 belong to the default computation domain (Group0), with a computation domain identifier of 0. Since each port in the default computation domain forms its own single-node tree, the parent node of both ports 6 and 7 is themselves, resulting in a tree height (rank) of 0 for both.

[0082] In the entire disjoint-set data structure, the Parent array for each port is [0, 0, 0, 3, 3, 3, 6, 7], the Rank array is [1, 0, 0, 1, 0, 0, 0, 0], and the root mapping table is [1, -, -, 2, -, -, 0, 0]. Since ports 1, 2, 4, and 5 are not root nodes, their corresponding values ​​in the root mapping table are invalid. For example, invalid values ​​can be set to all 0xFF.

[0083] Step S120: Update or query the bidirectional mapping table and / or disjoint set according to the received user operation command to perform dynamic port group management.

[0084] In this embodiment, user operation instructions include query instructions and port ownership modification instructions. Port ownership modification instructions are instructions that will cause changes in port ownership relationships, including modification instructions such as creating a computing domain, adding a port to a specified computing domain, deleting a port from a specified computing domain, merging computing domains, and destroying a computing domain. Query instructions include instructions such as querying the computing domain to which the target port belongs, querying ports in the target computing domain, and determining whether multiple ports are in the same domain.

[0085] It should be noted that before updating the bidirectional mapping table and / or disjoint-set data structure according to the received user operation command (port ownership modification command), a validity check of the port and / or computing domain is required to prevent the system from executing unexpected erroneous actions on invalid operation objects. Specifically, if the validity check fails, the execution of the port ownership modification command is stopped; if the validity check passes, the port ownership modification command is executed, including:

[0086] When the port ownership modification command is to create a compute domain, the validity check is considered successful if the compute domain identifier corresponding to the created compute domain is not occupied and all ports to be added to the compute domain belong to the default compute domain; when the port ownership modification command is to merge compute domains, the validity check is considered successful if none of the compute domains to be merged belong to the default compute domain; when the port ownership modification command is to destroy a compute domain, the validity check is considered successful if the destroyed compute domain is not the default compute domain; when the port ownership modification command is to add a port to a specified compute domain, the validity check is considered successful if the specified compute domain exists and the port to be added belongs to the default compute domain; when the port ownership modification command is to delete a port from a specified compute domain, the validity check is considered successful if the specified compute domain exists and the deleted port belongs to the specified compute domain.

[0087] After the legality verification is passed, the bidirectional mapping table and / or disjoint set are updated according to the port ownership modification instruction, and the communication permission configuration corresponding to the port whose port ownership has changed is updated to respond to the port ownership change; the port ownership modification instruction in this embodiment includes one or more of the following: instruction to create a computing domain, instruction to merge computing domains, instruction to destroy a computing domain, instruction to add a port to a specified computing domain, or instruction to delete a port from a specified computing domain.

[0088] When the user operation command is a port ownership modification command, this embodiment updates the bidirectional mapping table according to the port ownership modification command. For ports whose computing domain ownership has changed, it is necessary not only to modify the corresponding computing domain information in the first mapping table, that is, to change the computing domain identifier corresponding to the port to the computing domain identifier after the ownership change, but also to modify the port list information of the two computing domains involved before and after the ownership change in the second mapping table. That is, to delete the port from the port list of the computing domain before the ownership change and add the port to the port list of the computing domain after the ownership change.

[0089] If the port ownership modification instruction also changes the root node in the disjoint-setup, the root mapping table needs to be updated. For example, if a port is no longer a root node, the value corresponding to that port in the root mapping table needs to be set to an invalid value. If the port is still a root node but its computing domain ownership has changed, the computing domain identifier corresponding to that port in the root mapping table needs to be updated. If a port changes from not being a root node to becoming a root node, for example, after the computing domain is destroyed, the port in that computing domain no longer corresponds to any computing domain other than the default computing domain. The computing domain corresponding to that port can be set to the default computing domain, and the port becomes the root node in the default computing domain. In this case, the value corresponding to that port in the root mapping table needs to be set to 0 (corresponding to the default computing domain).

[0090] Furthermore, when the port ownership modification instruction is an instruction to create a computing domain, merge computing domains, or add a port to a specified computing domain, in addition to modifying the aforementioned bidirectional mapping table and the root mapping table when the root node changes, it also includes a port merging operation based on a disjoint-set data structure, updating the disjoint-set data structure (including set information corresponding to different port sets). The port merging operation in this embodiment is as follows: Figure 4 As shown, it includes the following steps:

[0091] Step S121: Obtain the ports to be merged required for the port merging operation.

[0092] In this embodiment, when the port ownership modification instruction is an instruction to create a computing domain, a port list consisting of all ports allocated to the computing domain is obtained, the target port is determined according to the port list, and the target port is merged with other ports in the port list. In each port merge operation, the target port and any other port other than the target port are used as the port to be merged for the port merge operation.

[0093] The determination of the target port is not specifically required; it can be any port in the port list. However, since the port list is stored in array form in this embodiment, the first port can be used as the target port. In other words, when performing port merging operations, the first port in the port list is used as the root node of the port set corresponding to the computing domain, and the other ports are merged in turn as child nodes of the root node.

[0094] When the port ownership modification instruction is an instruction to add a port to a specified computing domain, and a port needs to be added to the specified computing domain, the port to be added is a port to be merged. Any port in the specified computing domain is used as another port to be merged for the port merging operation. Subsequently, the root node corresponding to the specified computing domain is determined through any port in the specified computing domain, and the port to be added is used as a child node of the obtained root node, thereby adding it to the specified computing domain.

[0095] When the port ownership modification instruction is a merge computation domain instruction, for the two computation domains to be merged: any port in each computation domain is used as the port to be merged in the port merge operation. Since merging a computation domain with a larger rank into a computation domain with a smaller rank requires a longer search path and incurs more time costs, the specific circumstances of different computation domains need to be considered during the actual merge, rather than merging the two computation domains directly. In this embodiment, the root node corresponding to the computation domain is determined by any port in the computation domain, and the merge direction is determined by the rank of the obtained root node stored in the Rank array, in order to improve the merge efficiency.

[0096] Step S122: Determine the rank of the port set corresponding to each port to be merged.

[0097] In this embodiment, the rank of a port set corresponds to the root node in the port set. Therefore, when obtaining the rank of a port set, it is first necessary to obtain the root node of the port set corresponding to each port to be merged. That is, the Find operation is called to determine the root node corresponding to each port to be merged. Then, the rank of the port set is determined according to the value of the root node corresponding to the port set in the Rank array, so as to determine the merging direction during subsequent merging.

[0098] Since the parent node of a root node is itself, this embodiment determines whether the current port is the root node based on the parent node of the current port. Specifically, for any port to be merged: obtain the parent node of the port to be merged, and determine whether the parent node of the port to be merged is the port itself; if so, the port to be merged is the root node; otherwise, the port to be merged is not the root node, and continue to determine whether the obtained parent node is the root node. For example, assuming that the parent node corresponding to the port to be merged is port a, then continue to obtain the parent node of port a, and determine whether port a is the root node based on whether the obtained parent node is port a itself. If the parent node corresponding to the port to be merged (i.e., port a) is not the root node, then continue to traverse upwards and continue to determine whether the parent node of port a is the root node, and so on, until a port is found whose parent node is the port itself. At this time, the obtained port is taken as the root node of the port to be merged, and then the rank of the port set corresponding to the port to be merged is determined based on the root node.

[0099] After obtaining the root node, path compression optimization can be further performed on each port through iteration or recursion. The parent node of each port in the traversed path is updated to the obtained root node, so that the parent node of each port is the root node. Optimization through path compression means that after all ports involved in the query path are directly pointed to the root node, subsequent queries on that port and other ports on the path can directly reduce the search path length to 1, thereby reducing the time complexity to O(1), achieving fast same-domain queries and obtaining higher performance. It should be noted that without path compression optimization, relying solely on rank merging can also ensure that the same-domain judgment has an acceptable time complexity in engineering.

[0100] It should be noted that, generally speaking, creating a computing domain involves assigning ports that do not belong to other computing domains (in this embodiment, the ports that are in the default computing domain) to a new computing domain. Therefore, each port in its port list is essentially a root node. Since a port can only belong to one computing domain at a time, the port to be added is often an unassigned port in the default computing domain. That is, the port to be added can also be regarded as a root node. Therefore, the corresponding rank can be obtained directly through the Rank array.

[0101] Step S123: Determine the first computation domain based on the rank of the port set corresponding to each port to be merged, and use the other computation domain as the second computation domain.

[0102] In this embodiment, the first computing domain is used to receive ports that need to be merged, excluding those in this computing domain. When merging, the first computing domain is determined according to a preset merging rule. Specifically, merging can be performed by rank, that is, merging the tree with the smaller rank into the tree with the larger rank to avoid excessive growth in tree height. In this case, the computing domain corresponding to the first computing domain can be used as the computing domain to which the merged port set belongs. If the two trees have the same rank, since the time cost required for searching is the same regardless of which tree they are merged into, either of the two computing domain identifiers can be retained and used as the identifier of the corresponding computing domain after merging. In this case, the first computing domain can be further determined according to the size of the computing domain identifier, for example, the one with the larger computing domain identifier can be used as the first computing domain.

[0103] It should be noted that when creating a compute domain, after confirming the port's validity, the target port can be added to the newly created compute domain as the root node. At this time, the rank of the target port is the same as the rank of the port in the default compute domain, but the compute identifier of the newly created compute domain is greater than that of the default compute domain. Therefore, ports that are currently in the default compute domain but belong to the newly created compute domain are merged into the newly created compute domain. In the subsequent merging process, since the rank of the newly created compute domain is greater than the rank of the ports in the default compute domain (i.e., ports that have not yet been merged into the newly created compute domain), other ports in the port list will continue to be merged into the newly created compute domain. Therefore, the newly created compute domain is the first compute domain, and the default compute domain corresponds to the second compute domain. The port set corresponding to the second compute domain includes all ports assigned to the newly created compute domain.

[0104] When adding a port to a specified computing domain, since the rank of the port to be added is generally less than the rank of the port set corresponding to the specified computing domain, the specified computing domain is the first computing domain, and the default computing domain where the port to be added is located can be regarded as the second computing domain. At this time, the port set corresponding to the second computing domain only includes the port to be added.

[0105] When merging computational domains, the computational domain with the larger rank is designated as the first computational domain, and the computational domain with the smaller rank is designated as the second computational domain, based on the corresponding ranks of the two computational domains. If the ranks of the two computational domains are equal, the one with the larger identifier is designated as the first computational domain, and the other is designated as the second computational domain.

[0106] It should be noted that since merging two computation domains (both with non-zero identifiers) is essentially merging port sets, the time cost is the same regardless of which side is merged into when the rank is the same. Therefore, the one with the smaller identifier can be taken as the first computation domain.

[0107] For example, in the case of the above 8 ports, if it is necessary to merge the port sets containing port 0 and port 4, i.e., Union(0, 4), the root nodes corresponding to these two ports to be merged are first obtained through the Find operation. Since Parent[0] = 0, the root node corresponding to port 0 is itself, i.e., Find(0) = 0. For port 4, the Parent array is accessed first to determine the parent node corresponding to port 4. At this time, Parent[4] = 3, so port 4 is not the root node. Parent[3] is then obtained. Since Parent[3] = 3, the root node corresponding to port 4 is... Port 3, i.e., Find(4) = 3; then determine the corresponding rank based on the obtained root node. Since Rank[0] = 1 and Rank[3] = 1, and neither of these two computation domains is the default computation domain, either one can be selected as the first computation domain and the other as the second computation domain. For example, the one with the smaller computation domain identifier is selected as the first computation domain. The first computation domain is set as Group1 and the second computation domain is set as Group2. That is, the port set {3, 4, 5} corresponding to Group2 is merged into the port set {0, 1, 2} corresponding to Group1. The schematic diagram of merging the two computation domains is as follows. Figure 5 As shown.

[0108] Step S124: Merge the port set corresponding to the second computing domain into the port set corresponding to the first computing domain to obtain the merged port set.

[0109] First, obtain the root node of the port set corresponding to each port to be merged. At this time, the root node corresponding to each port to be merged has been obtained in the process of determining the rank. Then, merge the port set corresponding to the second computing domain and the port set corresponding to the first computing domain according to the root node corresponding to each port to be merged. This includes updating the parent node of the root node corresponding to the second computing domain to the root node corresponding to the first computing domain, and updating the rank corresponding to the merged port set.

[0110] For example, when merging the port set in Group2 into the port set corresponding to Group1 for the above 8 ports, it is only necessary to update the parent node corresponding to the root node of Group2 to the root node of Group1, that is, Parent[3]=0, and update Rank[0] to 2.

[0111] Furthermore, since the port sets of Group1 and Group2 have the same rank, the modification cost required to merge them into either group is the same. Therefore, the selection of the value corresponding to the computing domain identifier of the merged port set does not affect the final merging result. In other words, either computing domain identifier of the two computing domains can be retained and used as the computing domain to which the final merged port set belongs. For example, after merging the port set in Group2 into the port set corresponding to Group1, the computing domain identifier corresponding to the merged port set can be the computing domain identifier corresponding to Group2. Correspondingly, in the root mapping table, the computing domain identifier corresponding to port 0 (root node) is changed from the computing domain identifier corresponding to Group1 to the computing domain identifier corresponding to Group2. At this time, Root2Group[0]=2. Since the computing domain to which the merged port set belongs is Group2, it is necessary to update the ports whose computing domain identifier is not the computing domain identifier corresponding to Group2 (i.e., the ports that originally belonged to Group1).

[0112] It should be noted that after the merge is complete, the Parent array can be updated proactively to make the parent node of each port point to the root node, or the parent node of the corresponding port can be updated when the port or the root node is determined.

[0113] If the user's operation command is to delete a port from a specified computing domain or destroy the computing domain, it is only necessary to map the deleted port or the port in the destroyed computing domain to the default computing domain, without modifying the disjoint-set data structure. In this process, a lazy root mapping update strategy can be used, that is, when deleting a port from a specified computing domain or destroying the computing domain, the port involved will be transformed into a root node, and its parent node is the port itself. At this time, the expired entries in the root mapping table can be temporarily not cleaned up, and only corrected in subsequent queries or merges, so as to further reduce the overhead of real-time operations.

[0114] In addition, when the port list of a computing domain becomes empty, the computing domain can be automatically destroyed, or the empty computing domain structure can be retained for later reuse and explicitly deleted by the user.

[0115] Finally, the hardware configuration is synchronized according to the changed ownership relationship, enabling ports within the same computing domain to communicate with each other, thus satisfying communication permissions, while isolating ports between different computing domains, meaning ports in different computing domains cannot communicate with each other. This embodiment recalculates the communication permission configuration of all affected ports based on the changed ownership relationship, updating the communication permission configuration corresponding to ports whose ownership has changed, and updating the hardware registers through batch writing. Based on a universal port communication permission configuration mechanism, it does not rely on proprietary hardware features of specific vendors and can be adapted to any interconnect switching chip that supports port-level isolation, exhibiting strong portability.

[0116] For example, the communication permission configuration for each port is implemented through a set of 0 and 1 binary numbers. The number of bits in the binary number corresponding to each port is equal to the total number of ports. Each bit of the binary number corresponding to the current port is used to represent the connection status between the current port and other ports. All ports belonging to the same computing domain are obtained according to the second mapping table of bidirectional mapping. For any port in the computing domain: if the port can communicate with a port in the computing domain (i.e., the destination port), then in the communication permission configuration of the port, the binary number corresponding to the destination port is set to 1.

[0117] When the received user operation instruction is a query instruction, the query result is obtained according to the bidirectional mapping table and / or disjoint set; the query instruction in this embodiment includes one or more of the following: an instruction to query the computing domain to which the target port belongs (such as querying through the first mapping table), an instruction to query the port in the target computing domain (such as querying through the second mapping table), or an instruction to determine the same domain of multiple ports.

[0118] When the query command is a command to determine whether multiple ports belong to the same computing domain, the computing domain identifier corresponding to each port can be obtained. This can be done by obtaining the computing domain identifier corresponding to each port according to the bidirectional mapping table. In this case, it is necessary to traverse the first mapping table to determine the computing domain identifier corresponding to each port, and then determine whether the ports belong to the same computing domain based on whether the computing domain identifiers corresponding to each port are the same. Alternatively, the computing domain identifier corresponding to one of the ports can be obtained according to the first mapping table, and the port list of the corresponding computing domain in the second mapping table can be used to determine whether these ports belong to the same computing domain.

[0119] Preferably, the root node of the port set corresponding to each port can also be obtained. That is, for any port: trace upwards from the parent node of the port in the disjoint-set until the root node corresponding to the port is found. Determine whether the ports belong to the same computing domain based on whether the root nodes of the port sets corresponding to each port are the same, and obtain the corresponding query result. Alternatively, the computing domain identifier of the computing domain to which the root node of the port set corresponding to each port belongs can be determined based on the root mapping table. Determine whether the ports belong to the same computing domain and which computing domain they belong to based on whether the computing domain identifiers corresponding to each port are the same, and obtain the corresponding query result.

[0120] Taking the merged port set as an example, to query whether port 2 and port 5 belong to the same computing domain, the root node corresponding to port 2 and port 5 is determined by calling the Find operation. Since Parent[2]=0 and Parent[0]=0, the root node corresponding to port 2 is port 0. Since Parent[5]=3 and Parent[3]=0, the root node corresponding to port 5 is port 0. Since the root nodes are the same, port 2 and port 5 belong to the same computing domain. At this time, the computing domain to which port 2 and port 5 belong can be further determined according to the root mapping table. At the same time, Parent[5]=0 is set to perform path compression, so that the original query path changes from "port 5 to port 3 and then to the root node" to "port 5 directly to the root node". When dynamically managing a large number of port groups in the supernode interconnection switching chip, the efficiency of merging, querying and other operations in this embodiment is much higher than that of the traditional traversal-based method.

[0121] This embodiment achieves complete, efficient, and portable dynamic grouping management of AI supernode interconnection and switching chip ports by introducing bidirectional mapping data structures, default computation domain design, and disjoint-set data structure for accelerated merging. By establishing a bidirectional index between ports and computation domains through the bidirectional mapping data structure, fast bidirectional queries with low time complexity are achieved. This makes the time complexity of querying the computation domain to which a port belongs and retrieving all ports within that computation domain O(1), far superior to traditional traversal schemes. This advantage is particularly significant in scenarios with a large number of supernode chip ports. By introducing a disjoint-set data structure to manage the set relationships between ports and employing path compression and rank-based merging optimization, relying only on a small number of arrays and basic disjoint-set operations, the time complexity of merging two computation domains is reduced from O(n) of the traditional traversal scheme to O(αn) (approximately O(1)), improving efficiency. The flexibility of resource scheduling makes rapid resource reorganization possible in dynamic scheduling scenarios. Furthermore, since it eliminates the need for complex independent software components, it is easier to integrate into embedded environments (bare metal environments) or existing systems, achieving a lightweight software stack. Based on a universal port communication permission configuration design, it can be adapted to any interconnect switching chip that supports port-level isolation, improving the universality and portability of port group management. Utilizing the hardware characteristic that port communication permissions are fully interconnected by default after chip power-on, a default computing domain is designed, and the software data structure directly matches the hardware state, eliminating the need for any initial write operations to hardware registers and removing initialization overhead.

[0122] Please refer to Figure 6 Some embodiments provide a dynamic grouping management system for AI supernode interconnection switching chip ports, which includes:

[0123] Memory 200 is used to store programs;

[0124] And processor 210, used to implement the above-described AI supernode interconnection switching chip port dynamic group management method by executing the program.

[0125] After the supernode host CPU issues user operation commands, it performs resource validity verification to determine the port ownership adjustment scheme and encapsulates and generates a configuration message. This configuration message is then sent to the internal CPU (processor within the interconnect switching chip). The internal CPU immediately calls the program stored in memory and responds to the user operation command based on a bidirectional mapping table and / or a disjoint-set data structure. Furthermore, when port ownership changes, the internal CPU also reads the existing communication permission configuration within the interconnect switching chip and generates a new communication permission configuration for each port based on the port ownership change result, thus updating the logical ownership and interconnection isolation rules of all ports at once. After the configuration is written, the communication permission rules take effect. Finally, the internal CPU sends the execution result of this configuration back to the supernode host CPU via a reverse message, thus completing the entire operation loop and result verification.

[0126] In the above embodiments, implementation can be achieved, in whole or in part, by software, hardware, firmware, or any combination thereof. Furthermore, as those skilled in the art will understand, the principles herein can be reflected in a computer program product on a computer-readable storage medium pre-loaded with computer-readable program code. Any tangible, non-transitory computer-readable storage medium may be used, including magnetic storage devices (hard disks, floppy disks, etc.), optical storage devices (CDs, DVDs, Blu-ray discs, etc.), flash memory, and / or the like. These computer program instructions can be loaded onto a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to form a machine, such that instructions executing on the computer or other programmable data processing apparatus can generate means for performing a specified function. These computer program instructions can also be stored in a computer-readable storage medium that can instruct the computer or other programmable data processing apparatus to operate in a particular manner, such that instructions stored in the computer-readable storage medium can form an article of manufacture, including means for implementing the specified function. The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to perform a series of operational steps on the computer or other programmable apparatus to produce a computer-implemented process, such that instructions executing on the computer or other programmable apparatus can provide steps for implementing the specified function.

[0127] This document describes various exemplary embodiments with reference to them. However, those skilled in the art will recognize that changes and modifications can be made to the exemplary embodiments without departing from the scope of this document. For example, various operational steps and components for performing operational steps can be implemented in different ways depending on the specific application or considering any number of cost functions associated with the operation of the system (e.g., one or more steps can be deleted, modified, or combined with other steps).

[0128] While the principles herein have been illustrated in various embodiments, numerous modifications to the structures, arrangements, proportions, elements, materials, and components, particularly suited to specific environments and operational requirements, may be used without departing from the principles and scope of this disclosure. These modifications and other alterations or alterations will be included within the scope of this document. Those skilled in the art will recognize that many changes can be made to the details of the above embodiments without departing from the fundamental principles of the invention.

Claims

1. A method for dynamic group management of AI supernode interconnection switching chip ports, characterized in that, include: Obtain a bidirectional mapping table, which is used to represent the bidirectional index between the port and the computing domain; The bidirectional mapping table includes a first mapping table and a second mapping table. The first mapping table is used to determine the computing domain to which each port belongs; the second mapping table is used to determine the port list corresponding to each computing domain. A disjoint-set data structure is obtained, which is used to maintain the set relationship of ports in different computing domains; the disjoint-set data structure includes a root mapping table and set information of the port sets corresponding to all computing domains; wherein, the set information includes the rank of the port set and the parent node corresponding to each port; the root mapping table is used to represent the mapping relationship between each root node in the disjoint-set data structure and its computing domain. Based on the received user operation instructions, the bidirectional mapping table and / or disjoint set are updated or queried to perform dynamic port grouping management; the user operation instructions include query instructions and port attribution modification instructions; when the received user operation instruction is a port attribution modification instruction, the bidirectional mapping table and / or disjoint set are updated according to the port attribution modification instruction, including: updating the bidirectional mapping table according to the port attribution modification instruction; if the port attribution modification instruction changes the root node in the disjoint set, the root mapping table is updated; when the port attribution modification instruction is an instruction to create a computing domain, an instruction to merge computing domains, or an instruction to add a port to a specified computing domain, it also includes a port merging operation based on the disjoint set, updating the disjoint set; It updates the communication permission configuration corresponding to the port whose port ownership has changed in response to the port ownership change; when the received user operation command is a query command, it obtains the query result according to the bidirectional mapping table and / or disjoint set.

2. The AI ​​supernode interconnection switching chip port dynamic group management method as described in claim 1, characterized in that, include: For any computing domain: if the computing domain is not the default computing domain, the computing domain corresponds to a set of ports; If the computing domain is the default computing domain, each port in the default computing domain corresponds to a port set, the rank of each port set is a preset initial value, each port is the root node of its corresponding port set, and the parent node of the port is the port itself.

3. The AI ​​supernode interconnection switching chip port dynamic group management method as described in claim 2, characterized in that, include: The port ownership modification instructions include instructions for creating a computing domain, merging computing domains, destroying a computing domain, adding a port to a specified computing domain, or deleting a port from a specified computing domain. The query instructions include: instructions to query the computing domain to which the target port belongs, instructions to query ports in the target computing domain, or instructions to determine whether multiple ports are in the same domain. Specifically, when the query instruction is a multi-port same-domain judgment instruction, the root node of the port set corresponding to each port is obtained, and the port is judged to belong to the same computing domain based on the root node of the port set corresponding to each port, so as to obtain the corresponding query result.

4. The AI ​​supernode interconnection switching chip port dynamic group management method as described in claim 3, characterized in that, Before updating the bidirectional mapping table and / or disjoint-set data structure according to the received user operation instruction of port ownership modification, it is necessary to perform a validity check on the port and / or computing domain. If the validity check fails, the execution of the port ownership modification instruction will be stopped. If the validity check passes, the port ownership modification instruction is executed; including: When the port ownership modification instruction is an instruction to create a computing domain, if the computing domain identifier corresponding to the created computing domain is not occupied and all ports to be added to the computing domain belong to the default computing domain, the legality verification is considered to have passed; When the port ownership modification instruction is a merge computing domain instruction, if none of the computing domains to be merged are the default computing domains, the legality check is considered to have passed. When the port ownership modification instruction is an instruction to destroy a computing domain, if the destroyed computing domain is not the default computing domain, the legality check is considered to have passed. When the port ownership modification instruction is an instruction to add a port to a specified computing domain, if the specified computing domain exists and the port to be added belongs to the default computing domain, the legality verification is considered to have passed. When the port ownership modification instruction is an instruction to delete a port from a specified computing domain, if the specified computing domain exists and the deleted port belongs to the specified computing domain, the legality verification is considered to have passed.

5. The AI ​​supernode interconnection switching chip port dynamic group management method as described in claim 1, characterized in that, The port merging operation includes: Obtain the ports to be merged required for the port merging operation; determine the rank of the port set corresponding to each port to be merged; determine the first computation domain based on the rank of the port set corresponding to each port to be merged, and use another computation domain as the second computation domain; merge the port set corresponding to the second computation domain into the port set corresponding to the first computation domain to obtain the merged port set.

6. The AI ​​supernode interconnection switching chip port dynamic group management method as described in claim 5, characterized in that, The process of obtaining the ports to be merged required for the port merging operation includes: When the port ownership modification instruction is an instruction to create a computing domain, a port list consisting of all ports allocated to the computing domain is obtained, a target port is determined according to the port list, and the target port is merged with other ports in the port list. In each port merging operation, the target port and any other port other than the target port are used as the port to be merged for the port merging operation. When the port ownership modification instruction is an instruction to add a port to a specified computing domain, and a port needs to be added to the specified computing domain, the port to be added is a port to be merged, and any port in the specified computing domain is another port to be merged required for the port merging operation. When the port ownership modification instruction is a merge computing domain instruction, for the two computing domains to be merged: any port in each computing domain is used as the port to be merged for the port merging operation.

7. The AI ​​supernode interconnection switching chip port dynamic group management method as described in claim 5, characterized in that, The step of merging the port set corresponding to the second computing domain into the port set corresponding to the first computing domain to obtain the merged port set includes: Obtain the root node of the port set corresponding to each port to be merged; merge the port set corresponding to the second computing domain and the port set corresponding to the first computing domain according to the root node corresponding to each port to be merged, including: updating the parent node of the root node corresponding to the second computing domain to the root node corresponding to the first computing domain, and updating the rank corresponding to the merged port set.

8. The AI ​​supernode interconnection switching chip port dynamic group management method as described in claim 7, characterized in that, The step of obtaining the root node of the port set corresponding to each port to be merged includes: For any port to be merged: obtain the parent node of the port to be merged, and determine whether the parent node of the port to be merged is the port itself; if so, the port to be merged is the root node; otherwise, the port to be merged is not the root node, and continue to traverse upwards to determine whether the parent node is the root node, until a port is found whose parent node is the port itself. At this time, the obtained port is taken as the root node of the set of ports corresponding to the port to be merged.

9. The AI ​​supernode interconnection switching chip port dynamic group management method as described in claim 3 or 7, characterized in that, After obtaining the root node, the process also includes: performing path compression optimization on each port, wherein the parent node of each port in the traversed path is updated to the root node, so that the parent node of each port is the root node.

10. The AI ​​supernode interconnection switching chip port dynamic group management method as described in claim 1, characterized in that, In the initial bidirectional mapping table, all ports belong to the default computing domain; wherein, in the initial first mapping table, the computing domain to which each port belongs is the default computing domain; and in the initial second mapping table, the port list corresponding to the default computing domain contains all ports.

11. A dynamic grouping management system for AI supernode interconnection and switching chip ports, characterized in that, include: Memory, used to store programs; And a processor for implementing the AI ​​supernode interconnect switching chip port dynamic group management method as described in any one of claims 1 to 10 by executing the program.

Citation Information

Patent Citations

  • Troubleshooting method, device and equipment based on system component and switch chip resource configuration mapping and storage medium

    CN120675963A

  • System and method for providing port resiliency and artificial intelligence (AI) recommendations for failed ports at an information handling system

    US20260119347A1