Parameter plane network configuration method, apparatus, device, medium and product

By dynamically configuring the parameter plane network and configuring parameter plane information for the AI ​​server based on VPC information, the error problem caused by manual static configuration is solved, the isolation and communication between tenants are realized, and the configuration accuracy and training efficiency are improved.

CN120880912BActive Publication Date: 2026-01-23CHINA MOBILE (SUZHOU) SOFTWARE TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511374680.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-25
Publication Date
2026-01-23
Estimated Expiration
2045-09-25

AI Technical Summary

Technical Problem

In existing technologies, the configuration of parametric surface networks relies on manual static configuration, which is prone to errors, leading to a decrease in configuration accuracy and failing to meet the scalability and security requirements of large-scale intelligent computing resource pools.

Method used

By receiving network configuration requests, the system dynamically configures parameter plane information based on the configuration information of the Virtual Private Cloud (VPC) to which the AI ​​server belongs. This includes VLANs, parameter plane network segments, and NPU parameter plane addresses (IP addresses). It also orchestrates tenant isolation information to achieve dynamic isolation and communication between tenants, thus avoiding configuration errors.

Benefits of technology

It improves the accuracy of parameter surface network configuration, ensures the connectivity and interaction of parameter surface networks among the same tenants, prevents information interference between different tenants, avoids information leakage, and optimizes training efficiency and scheduling capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120880912B_ABST
    Figure CN120880912B_ABST
Patent Text Reader

Abstract

The present disclosure relates to the technical field of computer, and particularly provides a parameter plane network configuration method, device, equipment, medium and product. In the present disclosure, a network configuration request is received, and the network configuration request is used to request to configure parameter plane information of an artificial intelligence (AI) server for a parameter plane tenant; based on configuration information of a virtual private cloud (VPC) to which the AI server belongs, the parameter plane information of the AI server is configured; the parameter plane information includes a virtual local area network (VLAN), a parameter plane network segment, and a parameter plane address (IP) of a neural network processor (NPU) associated with the AI server; based on the parameter plane information, tenant isolation information is arranged; and the tenant isolation information is issued to a parameter plane network. Through the VPC configuration information of different tenants, the parameter plane information of the corresponding AI server is configured by using a dynamic configuration method, thereby avoiding configuration errors caused by pre-planning configuration. Therefore, the accuracy of parameter plane network configuration can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure belongs to the technical field of computers, and particularly relates to a parameter plane network configuration method, device, equipment, medium and product. BACKGROUND

[0002] The parameter plane network is a core network plane connecting computing power cards in AI (Artificial Intelligence) large model training, and is a network architecture component of a wisdom computing data center. The core function of the parameter plane network is to support parameter exchange and parallel computing between computing power cards. The parameter plane network isolation can realize independent network environments of different tenants / services, and support rapid deployment and safe switching of development, test and production environments.

[0003] In the related art, the parameter plane network information of a server uplink switch is pre-planned and configured in a manual static configuration manner, so as to realize the isolation of the parameter plane network. However, such a configuration method relies on manual work, and is prone to configuration errors. With the expansion of the scale and the increasing complexity of the wisdom computing resource pool parameter plane network, the accuracy of the parameter plane network configuration is reduced. SUMMARY

[0004] The present disclosure is proposed in view of the above problems. The present disclosure provides a parameter plane network configuration method, device, equipment, medium and product, which can improve the accuracy of the parameter plane network configuration.

[0005] According to one aspect of the present disclosure, a parameter plane network configuration method is provided, applied to a parameter plane network controller, and the method comprises:

[0006] receiving a network configuration request, the network configuration request being used to request to configure parameter plane information of an artificial intelligence (AI) server for a parameter plane tenant;

[0007] configuring parameter plane information of the AI server based on configuration information of a virtual private cloud (VPC) to which the AI server belongs; the parameter plane information comprising: a virtual local area network (VLAN), a parameter plane network segment, and a parameter plane address (IP) of a neural network processor (NPU) associated with the AI server;

[0008] based on the parameter plane information, compiling tenant isolation information;

[0009] issuing the tenant isolation information to a parameter plane network.

[0010] Optionally, the configuring parameter plane information of the AI server based on the configuration information of the VPC to which the AI server belongs comprises:

[0011] judging whether the VPC to which the AI server belongs is configured;

[0012] If the VPC to which the AI server belongs is configured, it is judged whether the parameter plane network segment of the VPC meets the configuration requirement of the parameter plane IP.

[0013] If the parameter plane network segment of the VPC meets the configuration requirement of the parameter plane IP, the parameter plane IP is configured for the NPU from the remaining network segment of the parameter plane network segment of the VPC; the VLAN and the parameter plane network segment of the VPC are unchanged.

[0014] Optionally, the method further comprises:

[0015] If the VPC to which the AI server belongs is not configured, or if the parameter plane network segment of the VPC does not meet the configuration requirement of the parameter plane IP, a new VLAN is allocated for the VPC.

[0016] The parameter plane network segment is allocated for the VPC from a pre-allocated parameter plane address resource pool.

[0017] The parameter plane IP is configured for the NPU from the parameter plane network segment.

[0018] Optionally, the judgment of whether the parameter plane network segment of the VPC meets the configuration requirement of the parameter plane IP comprises:

[0019] It is judged whether the number of remaining addresses of the parameter plane network segment of the VPC is greater than or equal to the number of the NPUs.

[0020] If the number of remaining addresses is greater than or equal to the number of the NPUs, it is determined that the parameter plane network segment of the VPC meets the configuration requirement of the parameter plane IP.

[0021] Optionally, the network configuration request carries one or more of the following information: elastic network card information of the AI server, physical address of the NPU, management address of the upper-connection leaf switch, and interface position of the NPU access switch.

[0022] Optionally, the tenant isolation information comprises one or more of the following: the VLAN corresponding to the AI server, and an access control list (ACL).

[0023] Optionally, the tenant isolation information is issued into the parameter plane network, comprising:

[0024] The tenant isolation information is sent to the leaf switch.

[0025] Optionally, the AI server comprises a bare metal server.

[0026] Optionally, the network configuration request is from the bare metal server.

[0027] The method further comprises:

[0028] The parameter plane information is fed back to the bare metal server.

[0029] Optionally, the parameter plane network is a two-layer networking architecture formed by a spine switch and a leaf switch; the spine switch and the leaf switch realize network full connection through a multi-track connection mode.

[0030] The leaf switch is connected with the AI server.

[0031] Optionally, the parameter plane network comprises at least one tenant group.

[0032] Any one tenant group corresponds to one or more AI servers, and any one AI server is associated with multiple NPUs.

[0033] Any one NPU directly communicates with the leaf switch and indirectly communicates with the spine switch or other NPUs through the leaf switch.

[0034] According to still another aspect of the present disclosure, a parameter plane network configuration device is provided, comprising:

[0035] A receiving module is configured to receive a network configuration request, wherein the network configuration request is used to request to configure parameter plane information of an artificial intelligence (AI) server for a parameter plane tenant;

[0036] A configuration module is configured to configure parameter plane information of the AI server based on configuration information of a virtual private cloud (VPC) to which the AI server belongs; the parameter plane information comprises a virtual local area network (VLAN), a parameter plane network segment, and a parameter plane address (IP) of a neural network processing unit (NPU) associated with the AI server.

[0037] An arrangement module is configured to arrange tenant isolation information based on the parameter plane information.

[0038] A sending module is configured to send the tenant isolation information to a parameter plane network.

[0039] According to another aspect of the present disclosure, a parameter plane network is provided, comprising a spine switch, a leaf switch, an AI server, an NPU, and a parameter plane network controller.

[0040] The parameter plane network controller is configured to implement the above-mentioned parameter plane network configuration method.

[0041] According to still another aspect of the present disclosure, an electronic device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the above-mentioned parameter plane network configuration method.

[0042] According to another aspect of the present disclosure, a computer readable storage medium is provided, having stored thereon a computer program, the program being executed by a processor to implement the above-mentioned parameter plane network configuration method.

[0043] According to still another aspect of the present disclosure, a computer program product is provided, comprising computer readable code, or a non-transitory computer readable storage medium carrying computer readable code, when the computer readable code is run in a processor of an electronic device, the processor in the electronic device executes to implement the above-mentioned parameter plane network configuration method.

[0044] In the present disclosure, a network configuration request is received, the network configuration request being used to request to configure parameter plane information of an artificial intelligence (AI) server for a parameter plane tenant; based on configuration information of a virtual private cloud (VPC) to which the AI server belongs, parameter plane information is configured for the AI server; the parameter plane information comprises a virtual local area network (VLAN), a parameter plane network segment, and a parameter plane address (IP) of a neural network processor (NPU) associated with the AI server; based on the parameter plane information, tenant isolation information is orchestrated; and the tenant isolation information is issued to a parameter plane network. By using the VPC configuration information of different tenants, the parameter plane information of the corresponding AI server is configured in a dynamic configuration manner, thereby avoiding configuration errors caused by pre-planning configuration. By orchestrating the tenant isolation information, dynamic isolation of the parameter plane network between tenants is achieved, so that the parameter plane networks between the same tenants are interconnected and interacted, the information between different tenants does not interfere with each other, and information leakage is avoided. Therefore, the accuracy of parameter plane network configuration can be improved.

[0045] It is to be understood that both the foregoing general description and the following detailed description are exemplary, and are intended to provide further explanation of the subject technology. BRIEF DESCRIPTION OF DRAWINGS

[0046] The foregoing and other objects, features, and advantages of the present disclosure will become more apparent from the following detailed description, which proceeds with reference to the accompanying drawings. The accompanying drawings are provided to illustrate embodiments of the present disclosure and to provide a further understanding of the present disclosure, and constitute a part of this specification, together with the description, to explain the present disclosure, and do not constitute a limitation of the present disclosure. In the drawings, the same reference numerals generally refer to the same components or steps throughout the drawings.

[0047] Figure 1 A parameter plane communication schematic diagram in an intelligent computing scenario in the related art.

[0048] Figure 2 A flowchart of a parameter plane network configuration method provided by the present disclosure.

[0049] Figure 3 A parameter plane network architecture diagram provided by the present disclosure.

[0050] Figure 4 Another flowchart is provided for a parametric surface network configuration method provided in this disclosure.

[0051] Figure 5 This is a schematic diagram of the structure of a parametric surface network configuration device provided in this disclosure.

[0052] Figure 6 This is a schematic diagram of the structure of a parametric surface network provided in this disclosure.

[0053] Figure 7 This is a hardware block diagram of an electronic device provided in this disclosure.

[0054] Figure 8 This is a schematic diagram of a computer program product provided in this disclosure. Detailed Implementation

[0055] To enable those skilled in the art to better understand the technical solution of this application, the application scenario of this application will be described first below.

[0056] Training a large model typically requires numerous NPUs (Neural Network Processing Units) or GPUs (Graphics Processing Units). For example, GPT 3.5 (Chat Generative Pre-trained Transformer 3.5) is estimated to have used over 30,000 GPUs. Large models are evolving towards clusters with tens of thousands of GPUs. The deployment of such clusters is not simply a matter of piling up computing power, but rather a pursuit of highly efficient collaborative work among tens of thousands of NPUs, much like "supercomputers." Distributed computing is employed within these clusters to collaboratively complete the same task. During the training of large model tasks, tens of thousands of NPUs or GPUs are required to train collaboratively, with data streams transmitted between the various AI servers via parametric networks.

[0057] Typical communication in large-scale intelligent computing scenarios with over 10,000 calories is as follows: Figure 1 As shown, Figure 1 This is a schematic diagram of parameter plane communication in intelligent computing scenarios in related technologies.

[0058] For training large models, data can be processed in parallel across AI servers by dividing the training samples into multiple mini-batches and running them on multiple AI servers. AI server construction requires synchronization of model parameters and gradients, and the communication model primarily uses AllReduce (full reduction). Figure 1The curved arrows indicate data parallelism. Pipelines can run in parallel across AI servers, with different Transformer Layers of the model running on each server. This mainly involves the synchronization of intermediate computation results between nodes, which can primarily be achieved through P2P (peer-to-peer) communication. Figure 1 The black double-headed arrows in the middle represent pipelined parallelism. Tensor parallelism exists within each AI server, meaning multiple cards within the AI ​​server can run in parallel. Each layer of the model is broken down into multiple sub-layers, each running on a different computing chip. This mainly involves synchronizing intermediate computation results; the communication model primarily uses AllReduce. Figure 1 The medium gray biphasic arrows indicate tensor parallelism.

[0059] The parameter plane network is a network architecture in data centers used to support large-scale distributed computing tasks (such as deep learning training). It is the core network plane connecting computing cards in AI large-scale model training. It primarily handles communication between parameter servers and computing nodes, supports parameter exchange and parallel computing between computing cards, and ensures efficient and reliable parameter updates and synchronization during training. The parameter plane network supports the RDMA (Remote Direct Memory Access) protocol, possessing ultra-high bandwidth and low latency capabilities. For high-performance, ultra-large data computing and transmission needs such as large-scale AI training, the traditional Ethernet TCP / IP network protocol (Transmission Control Protocol / Internet Protocol) cannot meet the requirements. Parameter plane network isolation enables independent network environments for different tenants / services, supporting rapid deployment and secure switching between development, testing, and production environments. As the scale and complexity of the parameter plane network in intelligent computing resource pools continue to increase, the requirements for large-scale model training efficiency and the scalability of parameter plane network isolation become even higher, making reasonable parameter plane network configuration particularly important.

[0060] In related technologies, the parameter plane network information of the switch connected to the AI ​​server is usually pre-planned and configured manually to achieve parameter plane network isolation. However, this configuration method of pre-planning the overall resource pool relies on manual labor, which is costly and prone to configuration errors. It also cannot achieve real-time isolation and reduces the accuracy of parameter plane network configuration.

[0061] To address the aforementioned technical problems, this disclosure provides an inventive concept: configuring parameter plane information for the corresponding AI server based on the configuration information of the VPC to which the AI ​​server belongs, i.e., the VPC configuration information of different tenants. This allows for dynamic configuration of parameter plane information based on the current VPC configuration, avoiding configuration errors caused by pre-planned static configuration methods. Furthermore, by orchestrating tenant isolation information, dynamic isolation of parameter plane networks between tenants is achieved, enabling interconnection and interaction between parameter plane networks within the same tenant while preventing information interference between different tenants and avoiding information leakage. Therefore, the accuracy of parameter plane network configuration can be improved.

[0062] To make the objectives, technical solutions, and advantages of this disclosure more apparent, exemplary embodiments according to this disclosure will now be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of this disclosure, and not all embodiments of this disclosure. It should be understood that this disclosure is not limited to the exemplary embodiments described herein.

[0063] Figure 2 This is a flowchart illustrating a parametric surface network configuration method provided in this disclosure. Figure 2 As shown, this method is applied to a parametric surface network controller and specifically includes:

[0064] S201: Receive network configuration request.

[0065] Specifically, users can subscribe to any number of AI servers after ordering VPC resources to complete training tasks. In cloud clusters with over 10,000 AI servers, many tenants may subscribe to AI servers, each training their own tasks. The NPUs on each AI server and the various switches form a parameter plane network. During AI task training, the NPUs synchronize gradient results through the parameter plane network. At this point, the training tasks of different tenants need to be independent; tenant A cannot access the training results from the NPUs on tenant B's server. However, all NPUs subscribed to by a single tenant must be interconnected, and the training results between NPUs must be synchronized. This requires the NPUs to send network configuration requests to the parameter plane network controller to configure the parameter plane network for each NPU, achieving data synchronization within the same tenant and data isolation between different tenants. Specifically, the network configuration request is used to request the configuration of the AI ​​server's parameter plane information for the parameter plane tenant.

[0066] S202: Configure parameter plane information for the AI ​​server based on the configuration information of the Virtual Private Cloud (VPC) to which the AI ​​server belongs.

[0067] Specifically, the parameter plane network controller stores the historical configuration information of the VPC for parameter plane tenants. In this embodiment, the AI ​​server is used to complete the training task for the corresponding parameter plane tenant. To facilitate data synchronization for the same parameter plane tenant, parameter plane information can be configured for the AI ​​server based on the historical configuration information. The parameter plane information includes the Virtual Local Area Network (VLAN), the parameter plane network segment, and the parameter plane address (IP) of the Neural Processing Unit (NPU) associated with the AI ​​server.

[0068] S203: Arrange tenant isolation information based on parameter plane information.

[0069] Specifically, in the relevant technologies, before being ordered, AI servers belong to the intelligent computing platform's resources. The AI ​​servers are bound to the platform's VPC, and they communicate with each other on the parameter plane network without tenant isolation. Tenant isolation can be implemented for AI servers ordered by parameter plane tenants. This isolation information can be orchestrated based on parameter plane information to achieve connectivity across all parameter plane network segments of the same tenant. Furthermore, VLANs can be used to isolate different parameter plane tenants, enabling elastic scaling of computing resources with parameter plane isolation.

[0070] S204: Send tenant isolation information to the parameter plane network.

[0071] Specifically, tenant isolation information is distributed to the parameter plane network. The parameter plane network then configures its network based on this information to enable interconnection among all parameter plane networks of the same tenant, while ensuring isolation between different parameter plane tenants. Dynamic configuration of the parameter plane network achieves complete decoupling between the parameter plane network and the service plane network. The parameter plane network orchestration process involves network resource allocation throughout, ensuring that the management of the parameter plane network does not affect the orchestration of the service plane network.

[0072] This disclosure involves receiving a network configuration request, which requests the configuration of parameter plane information for an AI server for a parameter plane tenant. Based on the configuration information of the Virtual Private Cloud (VPC) to which the AI ​​server belongs, parameter plane information is configured for the AI ​​server. This parameter plane information includes: Virtual Local Area Network (VLAN), parameter plane network segment, and the parameter plane address (IP) of the Neural Processing Unit (NPU) associated with the AI ​​server. Tenant isolation information is orchestrated based on this parameter plane information, and the tenant isolation information is distributed to the parameter plane network. By using the VPC configuration information of different tenants, parameter plane information for the corresponding AI server is configured dynamically, avoiding configuration errors caused by pre-planned configuration. Furthermore, by orchestrating tenant isolation information, dynamic isolation between tenants is achieved, enabling interconnection and interaction between parameter plane networks within the same tenant while preventing information interference between different tenants and avoiding information leakage. Therefore, the accuracy of parameter plane network configuration can be improved.

[0073] In one possible implementation, an exemplary method for configuring parameter plane information for the AI ​​server based on the configuration information of the Virtual Private Cloud (VPC) to which the AI ​​server belongs includes:

[0074] Determine if the VPC to which the AI ​​server belongs has been configured; if the VPC to which the AI ​​server belongs has been configured, determine if the parameter plane network segment of the VPC meets the configuration requirements of the parameter plane IP; if the parameter plane network segment of the VPC meets the configuration requirements of the parameter plane IP, configure the parameter plane IP for the NPU from the remaining network segment of the parameter plane network segment of the VPC.

[0075] Specifically, the parameter plane network controller stores the historical configuration information of the VPCs of the parameter plane tenants. It determines whether the VPCs mentioned by the AI ​​server have been configured, i.e., whether historical configuration information exists, based on the relevant information carried in the network configuration request. Furthermore, the network configuration request carries one or more of the following information: the AI ​​server's elastic network interface card information, the NPU's physical address, the management address of the uplink leaf switch, and the interface location of the NPU access switch.

[0076] In this embodiment, it can be determined whether the AI ​​server's elastic network interface card (NIC) has been configured by checking its information. If it has been configured, it can be further determined whether the VPC's parameter plane network segment meets the configuration requirements of the parameter plane IP.

[0077] Furthermore, exemplary methods for determining whether a VPC's parameter plane network segment meets the configuration requirements of the parameter plane IP include:

[0078] Determine if the number of remaining addresses in the VPC's parameter plane segment is greater than or equal to the number of NPUs; if the number of remaining addresses is greater than or equal to the number of NPUs, then the VPC's parameter plane segment meets the parameter plane IP configuration requirements.

[0079] For example, if the number of NPUs in the AI ​​server is 8, it is necessary to determine whether the number of remaining addresses in the parameter plane network segment is greater than or equal to 8. If so, the configuration requirements are met.

[0080] Next, the target number of IPs are selected from the remaining network segments to allocate parameter plane IPs to the NPUs of the AI ​​server. Assuming the parameter plane IP configuration requirements are met, the VPC's VLANs and parameter plane network segments remain unchanged; that is, the AI ​​server's NPUs are automatically assigned to the same VLAN and parameter plane network segment.

[0081] By configuring dynamic parameter plane information, network configuration is automated, significantly reducing the need for manual intervention. The system automatically and uniformly manages network parameters, ensuring the efficiency and consistency of network configuration and maintenance.

[0082] In one possible implementation, the method further includes:

[0083] If the VPC to which the AI ​​server belongs has not been configured, or if the parameter plane network segment of the VPC does not meet the parameter plane IP configuration requirements, allocate a new VLAN for the VPC; allocate a parameter plane network segment for the VPC from the pre-allocated parameter plane address resource pool; configure parameter plane IP for the NPU from the parameter plane network segment.

[0084] Specifically, if the VPC of the AI ​​server has not been configured, meaning there is no assigned VLAN and parameter plane segment in the parameter plane network controller; or, the number of remaining addresses in the parameter plane segment is less than the number of NPUs, meaning the AI ​​server's NPUs cannot be assigned to the same parameter plane segment, then a new VLAN needs to be automatically assigned, and a corresponding parameter plane segment needs to be allocated from the pre-allocated parameter plane address resource pool.

[0085] In one possible implementation, tenant isolation information includes one or more of the following: the VLAN corresponding to the AI ​​server, and the Access Control List (ACL).

[0086] Specifically, the VLAN corresponding to the AI ​​server is used to isolate the parameter plane network between different tenants by dividing virtual LANs, and the access control list (ACL) is used to connect the parameter plane network segments of different AI servers of the same tenant, so as to realize the isolation and communication between tenants.

[0087] In one possible implementation, the AI ​​server includes a bare metal server.

[0088] Furthermore, the network configuration request originates from the bare metal server; the method also includes:

[0089] Feedback of parameter plane information to the bare metal server.

[0090] Specifically, when the AI ​​server is a bare metal server, after configuring the parameter plane information for the bare metal server, the parameter plane information is fed back to the corresponding bare metal server. The bare metal server can configure the parameter plane IP for each NPU through the driver.

[0091] In one possible implementation, the parameter plane network is a two-layer network architecture consisting of spine switches and leaf switches; the spine switches and leaf switches achieve full network connectivity through a multi-track connection method; the leaf switches are connected to the AI ​​server.

[0092] Furthermore, the parameter plane network includes at least one tenant group; any tenant group corresponds to one or more AI servers, any AI server is associated with multiple NPUs; any NPU communicates directly with the leaf switch, and indirectly communicates with the spine switch or other NPUs through the leaf switch.

[0093] Specifically, Figure 3 The parametric network architecture diagram provided in this disclosure is as follows: Figure 3 As shown. In this embodiment, a parameter plane network can be constructed using a two-layer networking architecture consisting of spine switches and leaf switches. The spine switches and leaf switches achieve full network connectivity through a multi-track connection method; the leaf switches are connected to the AI ​​server. The AI ​​server serves as the training port for training large models.

[0094] For example, an AI server includes 8 NPUs, and a leaf switch has 64 interfaces, of which 32 interfaces are used for uplink communication with the spine switch, and 32 interfaces communicate with the NPUs on the AI ​​server for downlink communication. One leaf switch can connect 4 AI servers. This embodiment uses a multi-track connection method, that is, the 8 NPUs on the AI ​​server are connected to 8 leaf switches. The uplink-to-downlink convergence ratio of the leaf switches is 1:1, ensuring network stability. Compared with single-track networking, this architecture can significantly reduce the number of spine switches required for communication between network ports, facilitating both network scalability and performance optimization.

[0095] In one possible implementation, the tenant isolation information is distributed to the parameter plane network, including sending the tenant isolation information to the leaf switch.

[0096] Specifically, the NPUs on each AI server and each switch form a parameter plane network, which sends tenant isolation information to the leaf switches corresponding to the AI ​​servers to achieve isolation and communication between tenants.

[0097] Furthermore, during large-scale model training, training tasks can be submitted with a single click, and continuous, efficient training can last for half a month with normal model convergence. However, in reality, large-scale model training often involves frequent task restarts due to various hardware and configuration issues, with continuous training time typically not exceeding one day. A single training interruption can add up to four hours of training time. When a pod is interrupted during training and needs to resume training from the breakpoint, the intelligent computing platform performs affinity scheduling. The parameter plane network controller can generate topology information for parameter plane tenants based on the NPU's material address, uplink switch management address, parameter plane IP, and the corresponding NPU access switch interface location—that is, dynamically managed parameter plane information. This provides the upper-layer intelligent computing platform with precise resource training scheduling capabilities. Based on the topology information, the intelligent computing platform can place the training tasks of the tenant under the same leaf switch as much as possible, achieving affinity scheduling for training, effectively reducing parameter plane communication latency loss, improving training efficiency by approximately 8%, and optimizing training and scheduling capabilities.

[0098] Figure 4 This is another flowchart illustrating a parametric surface network configuration method provided in this disclosure. Figure 4 As shown, the method includes:

[0099] S401: Receive network configuration requests from bare metal servers.

[0100] Specifically, tenants can create VPC resources and subscribe to any number of bare metal servers under these VPC resources to implement training tasks. Typically, a bare metal server will have 8 NPU cards installed for training, so one bare metal server corresponds to 8 parameter plane NPU cards. Parameter plane IPs are created for the 8 NPU cards of each bare metal server. The bare metal server calls the northbound interface of the parameter network controller to send a network configuration request carrying the elastic network interface information of the bare metal server, the physical address of the NPU, the management address of the uplink leaf switch, and the interface location of the NPU access switch to the parameter plane network controller.

[0101] S402: Based on VPC configuration information, assign parameter plane IPs to the NPUs in the bare metal server.

[0102] Specifically, parameter plane IPs are assigned to the NPUs in the bare metal server based on the historical configuration information of the VPC.

[0103] S403: Determine whether the VPC of the bare metal server has been configured.

[0104] Specifically, you can determine whether the same elastic network information has been configured before based on the bare metal server's elastic network information. If it has been configured, you can execute S404; if it has not been configured, you can execute S405.

[0105] S404: Determine whether the number of remaining addresses in the configured parameter plane network segment meets the requirements.

[0106] Specifically, if a VPC has been configured, it is necessary to determine whether there are 8 remaining parameter plane IPs in the configured parameter plane network segment. In this embodiment, multiple NPUs of the same bare metal server need to be assigned to the same parameter plane network segment. If there are 8 or more remaining, S406 can be executed; if there are fewer than 8 remaining, S405 can be executed.

[0107] S405: Assign a new VLAN and a parameter plane segment.

[0108] Specifically, if the VPC has not been allocated before, or the remaining number of parameter plane IPs does not meet the requirements, a new VLAN needs to be automatically allocated, and a parameter plane network segment with 24 subnets needs to be allocated from the pre-allocated parameter plane address resource pool.

[0109] S406: Configure the parameter plane IP for the NPU from the parameter plane network segment.

[0110] Specifically, eight parameter plane IPs are selected from the parameter plane network segments of the 24 subnets and assigned to the corresponding NPUs in sequence.

[0111] S407: Arrange tenant isolation information based on parametric plane information.

[0112] Specifically, the parameter plane information includes the Virtual LAN (VLAN), parameter plane network segment, and the parameter plane IP address of the NPU associated with the bare metal server. Based on the parameter plane information, tenant isolation information is orchestrated, which includes at least one of the following: the VLAN corresponding to the bare metal server and the Access Control List (ACL).

[0113] S408: Sends parameter plane information to the bare metal server.

[0114] Specifically, the parameter plane information is fed back to the corresponding bare metal server, which can configure the parameter plane IP for each NPU through the driver.

[0115] S409: Send tenant isolation information to the leaf switch.

[0116] Specifically, tenant isolation information is sent to the leaf switch to achieve tenant isolation and interoperability.

[0117] Figure 5 This is a schematic diagram of a parametric surface network configuration device provided in this disclosure. Figure 5 As shown, the device 500 includes: a receiving module 510, a configuration module 520, an arrangement module 530, and a sending module 540.

[0118] The receiving module 510 is used to receive a network configuration request, wherein the network configuration request is used to request the configuration of the parameter plane information of the AI ​​server for the parameter plane tenant;

[0119] The configuration module 520 is used to configure parameter plane information for the AI ​​server based on the configuration information of the virtual private cloud (VPC) to which the AI ​​server belongs; the parameter plane information includes: virtual local area network (VLAN), parameter plane network segment, and parameter plane address (IP) of the neural network processor (NPU) associated with the AI ​​server;

[0120] The orchestration module 530 is used to orchestrate tenant isolation information based on the parameter plane information;

[0121] The sending module 540 is used to send the tenant isolation information to the parameter plane network.

[0122] Optionally, the configuration module is used for:

[0123] Determine whether the VPC to which the AI ​​server belongs has been configured.

[0124] If the VPC to which the AI ​​server belongs has been configured, determine whether the parameter plane network segment of the VPC meets the configuration requirements of the parameter plane IP.

[0125] If the parameter plane network segment of the VPC meets the parameter plane IP configuration requirements, configure the parameter plane IP for the NPU from the remaining network segments of the parameter plane network segment of the VPC; the VLAN and parameter plane network segment of the VPC remain unchanged.

[0126] Optionally, the device further includes:

[0127] If the VPC to which the AI ​​server belongs has not been configured, or if the parameter plane network segment of the VPC does not meet the parameter plane IP configuration requirements, a new VLAN shall be assigned to the VPC.

[0128] Allocate the parameter plane network segment to the VPC from the pre-allocated parameter plane address resource pool;

[0129] Configure the parameter plane IP for the NPU from the parameter plane network segment.

[0130] Optionally, the device further includes:

[0131] Determine whether the number of remaining addresses in the parameter plane network segment of the VPC is greater than or equal to the number of NPUs;

[0132] If the number of remaining addresses is greater than or equal to the number of NPUs, it is determined that the parameter plane network segment of the VPC meets the parameter plane IP configuration requirements.

[0133] Optionally, the sending module is used to:

[0134] Send the tenant isolation information to the leaf switch.

[0135] Optionally, the network configuration request originates from the bare metal server; the apparatus further includes:

[0136] The parameter plane information is fed back to the bare metal server.

[0137] Figure 6 This is a schematic diagram of the structure of a parametric surface network provided in this disclosure. For example... Figure 6 As shown, the parameter plane network includes: a spine switch 610, a leaf switch 620, an AI server 630, an NPU 640, and a parameter plane network controller 650.

[0138] The parameter plane network controller is used to implement the above parameter plane network configuration method.

[0139] This application also provides an electronic device for performing the above-described parameter plane network configuration method. Please refer to... Figure 7It illustrates a schematic diagram of an electronic device provided by some embodiments of this application. For example... Figure 7 As shown, the electronic device 7 includes: a processor 700, a memory 701, a bus 702, and a communication interface 703. The processor 700, the communication interface 703, and the memory 701 are connected via the bus 702. The memory 701 stores a computer program that can run on the processor 700. When the processor 700 runs the computer program, it executes the parameter plane network configuration method provided in any of the foregoing embodiments of this application.

[0140] The memory 701 may include high-speed random access memory (RAM) or non-volatile memory, such as at least one disk storage device. Communication between the device network element and at least one other network element is achieved through at least one communication interface 703 (which can be wired or wireless), such as the Internet, wide area network, local area network, metropolitan area network, etc.

[0141] Bus 702 can be an ISA bus, PCI bus, or EISA bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. The memory 701 is used to store programs. After receiving an execution instruction, the processor 700 executes the program. The parameter plane network configuration method disclosed in any of the foregoing embodiments of this application can be applied to the processor 700, or implemented by the processor 700.

[0142] The processor 700 may be an integrated circuit chip with signal processing capabilities. In implementation, each step of the above method can be completed by the integrated logic circuitry in the hardware of the processor 700 or by instructions in software form. The processor 700 may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), an off-the-shelf programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly embodied in the execution of a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software modules may reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The storage medium is located in memory 701. Processor 700 reads the information in memory 701 and, in conjunction with its hardware, completes the steps of the above method.

[0143] The electronic device provided in this application embodiment and the parameter plane network configuration method provided in this application embodiment are based on the same inventive concept and have the same beneficial effects as the methods they adopt, operate or implement.

[0144] This application also provides a computer-readable storage medium corresponding to the parameter plane network configuration method provided in the foregoing embodiments. The computer-readable storage medium shown can be an optical disc, on which a computer program is stored. When the computer program is run by a processor, it executes the parameter plane network configuration method provided in any of the foregoing embodiments.

[0145] It should be noted that examples of the computer-readable storage medium may also include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other optical and magnetic storage media, which will not be elaborated here.

[0146] The computer-readable storage medium provided in the above embodiments of this application and the parameter plane network configuration method provided in the embodiments of this application are based on the same inventive concept and have the same beneficial effects as the methods adopted, run or implemented by the applications stored therein.

[0147] This application also provides a computer program product 800, such as... Figure 8 As shown. This computer program product carries a computer program 801, the instructions of which can be used to execute the steps of the parameter plane network configuration method described in the above method embodiments. For details, please refer to the above method embodiments, which will not be repeated here.

[0148] The aforementioned computer program product can be implemented through hardware, software, or a combination thereof. In one optional embodiment, the computer program product is specifically embodied in a computer storage medium; in another optional embodiment, the computer program product is specifically embodied in a software product, such as a software development kit (SDK), etc.

[0149] The basic principles of this disclosure have been described above with reference to specific embodiments. However, it should be noted that the advantages, benefits, and effects mentioned in this disclosure are merely examples and not limitations, and should not be considered as essential features of each embodiment of this disclosure. Furthermore, the specific details disclosed above are for illustrative and facilitative purposes only, and are not limitations. These details do not limit the scope of this disclosure to the necessity of employing the aforementioned specific details for implementation.

[0150] The block diagrams of devices, apparatuses, devices, and systems disclosed herein are merely illustrative examples and are not intended to require or imply that they must be connected, arranged, or configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, devices, and systems can be connected, arranged, and configured in any manner. Words such as “comprising,” “including,” “having,” etc., are open-ended terms meaning “including but not limited to,” and are used interchangeably with them. The terms “or” and “and” as used herein refer to the terms “and / or,” and are used interchangeably with them unless the context clearly indicates otherwise. The term “such as” as used herein refers to the phrase “such as but not limited to,” and is used interchangeably with it.

[0151] Additionally, as used herein, the "or" used in a list of items beginning with "at least one" indicates a separate list, such that a list of, for example, "at least one of A, B, or C" means A or B or C, or AB or AC or BC, or ABC (i.e., A and B and C). Furthermore, the word "exemplary" does not imply that the described example is preferred or better than other examples.

[0152] It should also be noted that in the systems and methods of this disclosure, the components or steps can be decomposed and / or recombined. These decompositions and / or recombinations should be considered as equivalent solutions to this disclosure.

[0153] Various changes, substitutions, and modifications can be made to the technology described herein without departing from the teachings defined by the appended claims. Furthermore, the scope of the claims of this disclosure is not limited to the specific aspects of the processes, machines, manufactures, events, means, methods, and actions described above. Currently existing or later-developed processes, machines, manufactures, events, means, methods, or actions that perform substantially the same function or achieve substantially the same result as the corresponding aspects described herein can be utilized. Therefore, the appended claims include such processes, machines, manufactures, events, means, methods, or actions within their scope.

[0154] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use this disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects without departing from the scope of this disclosure. Therefore, this disclosure is not intended to be limited to the aspects shown herein, but rather to be carried out within the widest scope consistent with the principles and novel features disclosed herein.

[0155] The above description has been given for purposes of illustration and description. Furthermore, this description is not intended to limit the embodiments of this disclosure to the forms disclosed herein. Although numerous exemplary aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations therein.

Claims

1. A method for configuring parametric surface networks, characterized in that, Applied to a parametric plane network controller, the method includes: Receive a network configuration request, the network configuration request being used to request the configuration of the parameter plane information of the AI ​​server for the parameter plane tenant; Based on the configuration information of the Virtual Private Cloud (VPC) to which the AI ​​server belongs, parameter plane information is configured for the AI ​​server; the parameter plane information includes: Virtual Local Area Network (VLAN), parameter plane network segment, and parameter plane address IP of the Neural Processing Unit (NPU) associated with the AI ​​server; Based on the parameter plane information, arrange tenant isolation information; The tenant isolation information is then sent to the parameter plane network; The configuration parameter plane information for the AI ​​server based on the configuration information of the Virtual Private Cloud (VPC) to which the AI ​​server belongs includes: Determine whether the VPC to which the AI ​​server belongs has been configured. If the VPC to which the AI ​​server belongs has not been configured, or if the parameter plane network segment of the VPC does not meet the parameter plane IP configuration requirements, a new VLAN shall be assigned to the VPC. Allocate the parameter plane network segment to the VPC from the pre-allocated parameter plane address resource pool; Configure the parameter plane IP for the NPU from the parameter plane network segment.

2. The method according to claim 1, characterized in that, The configuration parameter plane information for the AI ​​server based on the configuration information of the Virtual Private Cloud (VPC) to which the AI ​​server belongs includes: Determine whether the VPC to which the AI ​​server belongs has been configured. If the VPC to which the AI ​​server belongs has been configured, determine whether the parameter plane network segment of the VPC meets the configuration requirements of the parameter plane IP. If the parameter plane network segment of the VPC meets the parameter plane IP configuration requirements, configure the parameter plane IP for the NPU from the remaining network segments of the parameter plane network segment of the VPC; the VLAN and parameter plane network segment of the VPC remain unchanged.

3. The method according to claim 2, characterized in that, The determination of whether the parameter plane network segment of the VPC meets the configuration requirements of the parameter plane IP includes: Determine whether the number of remaining addresses in the parameter plane network segment of the VPC is greater than or equal to the number of NPUs; If the number of remaining addresses is greater than or equal to the number of NPUs, it is determined that the parameter plane network segment of the VPC meets the parameter plane IP configuration requirements.

4. The method according to any one of claims 1-3, characterized in that, The network configuration request carries one or more of the following information: the AI ​​server's elastic network interface card information, the NPU's physical address, the management address of the uplink leaf switch, and the interface location of the NPU access switch.

5. The method according to any one of claims 1-3, characterized in that, The parameter plane network is a two-layer network architecture consisting of a spine switch and a leaf switch; the spine switch and the leaf switch achieve full network connectivity through a multi-track connection method. The leaf switch is connected to the AI ​​server.

6. The method according to claim 5, characterized in that, The parameter plane network includes at least one tenant group; Each tenant group corresponds to one or more AI servers, and each AI server is associated with multiple NPUs; Each of the NPUs communicates directly with the leaf switch and indirectly with the spine switch or other NPUs through the leaf switch.

7. A parametric surface network configuration device, characterized in that, include: The receiving module is used to receive network configuration requests, which are used to request the configuration of parameter plane information of the AI ​​server for the parameter plane tenant; The configuration module is used to configure parameter plane information for the AI ​​server based on the configuration information of the virtual private cloud (VPC) to which the AI ​​server belongs; The parameter plane information includes: Virtual Local Area Network (VLAN), parameter plane network segment, and parameter plane address (IP) of the Neural Processing Unit (NPU) associated with the AI ​​server; The orchestration module is used to orchestrate tenant isolation information based on the parameter plane information; The sending module is used to send the tenant isolation information to the parameter plane network; The configuration module is used for: Determine whether the VPC to which the AI ​​server belongs has been configured. If the VPC to which the AI ​​server belongs has not been configured, or if the parameter plane network segment of the VPC does not meet the parameter plane IP configuration requirements, a new VLAN shall be assigned to the VPC. Allocate the parameter plane network segment to the VPC from the pre-allocated parameter plane address resource pool; Configure the parameter plane IP for the NPU from the parameter plane network segment.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, The processor executes the computer program to implement the method as described in any one of claims 1-6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, The program is executed by a processor to implement the method as described in any one of claims 1-6.

10. A computer program product, characterized in that, Includes computer-readable code, or a non-volatile computer-readable storage medium carrying computer-readable code, wherein when the computer-readable code is run in a processor of an electronic device, the processor in the electronic device performs the method as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Tenant isolation management method and device, equipment and storage medium

    CN117176720A

  • Tenant isolation method, tenant isolation device and related equipment

    CN119520020A