AI cluster simulation method and device, electronic equipment and storage medium
Through container technology, the creation of simulation node containers on the host is solved, and the problems of waste of AI cluster verification resources and uncontrollable nodes in the existing technology are achieved, efficient AI cluster simulation is achieved, and hardware simulation and cluster scale adjustment are supported in the GPU chip development stage.
Patent Information
- Application Number
- CN202510570281.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-07-11
AI Technical Summary
In the prior art, verifying the effectiveness of the AI cluster environment requires the real construction of an AI cluster, which leads to waste of resources and time-consuming and labor-intensive, and cannot be effectively verified during the development stage of GPU chips, and the number of cluster nodes is uncontrollable, which affects the verification cost and difficulty of the large model training process.
Through container technology, we create simulation node containers on a small number of hosts, and use host lists, network configuration information and cluster configuration information to automatically create target AI simulation clusters, including multiple simulation node containers created on the host, realizing resource isolation and sharing, and dynamically adjusting the cluster size.
It effectively saves physical hardware resources, improves the efficiency of AI cluster simulation, can perform hardware simulation in the GPU chip development stage, flexibly adjusts the cluster scale, reduces the difficulty of cluster simulation, and improves the development and testing efficiency of large-scale training models.
Smart Images

Figure CN120301781A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technologies, and in particular, to an AI cluster simulation method, an apparatus, an electronic device, and a storage medium. Background Art
[0002] With the continuous growth of deep learning model parameters, training large models requires huge computing resources and complex cluster environments, and usually requires a large number of Graphics Processing Unit (GPU) servers and corresponding software stacks to support. However, in the prior art, to verify the effectiveness of these cluster environments, it is often necessary to actually build an Artificial Intelligence (AI) cluster, which is not only time-consuming and laborious, but also may cause waste of resources. Summary of the Invention
[0003] The present disclosure provides a technical solution for an AI cluster simulation method, an apparatus, an electronic device, and a storage medium.
[0004] According to one aspect of the present disclosure, there is provided an AI cluster simulation method, including: determining configuration information of a target AI simulation cluster, where the configuration information includes: a list of host machines, network configuration information, simulation node configuration information, and cluster configuration information; creating a container network on each host machine included in the list of host machines according to the network configuration information; creating a plurality of simulation node containers on each host machine according to the container network and the simulation node configuration information, and determining simulation node information of each simulation node container; creating the target AI simulation cluster according to the cluster configuration information and the simulation node information of each simulation node container, where the target AI simulation cluster includes a plurality of simulation node containers created on each host machine.
[0005] In a possible implementation manner, the list of host machines includes: connection information of each host machine; the method further includes: checking whether each host machine can be normally connected according to the connection information of each host machine.
[0006] In a possible implementation manner, the network configuration information includes: a network name, a subnet mask, a gateway, a network type, and an IP address range; the creating a container network on each host machine included in the list of host machines according to the network configuration information includes: for any one host machine, creating the container network in the host machine according to the network name, the subnet mask, the gateway, the network type, and the IP address range.
[0007] In a possible implementation, the list of host machines includes: the number of target simulation nodes corresponding to each host machine; the simulation node configuration information includes: the container image address, the hardware configuration information; creating multiple simulation node containers on each host machine according to the container network and the simulation node configuration information, and determining the simulation node information of each simulation node container includes: for any one host machine, performing the simulation node creation operation multiple times on this host machine according to the IP address range, the container image address, and the hardware configuration information until the number of simulation node containers of the target simulation nodes is created on this host machine.
[0008] In a possible implementation, the simulation node information of each simulation node container includes: the node name of this simulation node container, the node IP address; the simulation node creation operation includes: randomly generating a string, and determining this string as the node name of the target simulation node container, where the target simulation node container is any one of the multiple simulation node containers that need to be created in this host machine; selecting an idle IP address from the IP address range, and determining this idle IP address as the node IP address of the target simulation node container; obtaining the target container image from the container image address, and creating the target simulation node container by using the target container image; performing corresponding hardware simulation in the target simulation node container according to the hardware configuration information.
[0009] In a possible implementation, the hardware configuration information includes: the GPU model, the number of GPUs, the number of CPUs, the memory size.
[0010] In a possible implementation, the method further includes: performing connectivity detection on each simulation node container according to the simulation node information of each simulation node container.
[0011] In a possible implementation, the cluster configuration information includes: the cluster version number; creating the target AI simulation cluster according to the cluster configuration information and the simulation node information of each simulation node container includes: installing the target software in each simulation node container according to the simulation node information of each simulation node container, and performing time zone synchronization processing on each simulation node container; determining the target cluster installation tool corresponding to the cluster version number according to the cluster version number; using the target cluster installation tool to create the target AI simulation cluster including each simulation node container, where the target AI simulation cluster corresponds to the cluster version indicated by the cluster version number.
[0012] In a possible implementation, the cluster configuration information includes: a network plugin, a component image address; creating the target AI simulation cluster including each simulation node container by using the target cluster installation tool includes: obtaining a target component image from the component image address by using the target cluster installation tool; installing a target component in each simulation node container to add each simulation node container to the target AI simulation cluster; and installing the network plugin in each simulation node container.
[0013] In a possible implementation, the method further includes: determining connection authentication information of the target AI simulation cluster; logging in to the target AI simulation cluster according to the connection authentication information of the target AI simulation cluster; and performing a functional test on the target AI simulation cluster by executing a target AI training task.
[0014] According to one aspect of the present disclosure, there is provided an AI cluster simulation device, including: a simulation configuration module, configured to determine configuration information of a target AI simulation cluster, where the configuration information includes: a list of host machines, network configuration information, simulation node configuration information, cluster configuration information; a network configuration module, configured to create a container network on each host machine included in the list of host machines according to the network configuration information; a node simulation module, configured to create a plurality of simulation node containers on each host machine according to the container network and the simulation node configuration information, and determine simulation node information of each simulation node container; and a cluster formation module, configured to create the target AI simulation cluster according to the cluster configuration information and the simulation node information of each simulation node container, where the target AI simulation cluster includes a plurality of simulation node containers created on each host machine.
[0015] According to one aspect of the present disclosure, there is provided an electronic device, including: a processor; a memory for storing instructions executable by the processor; wherein the processor is configured to call the instructions stored in the memory to execute the above method.
[0016] According to one aspect of the present disclosure, there is provided a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the above method is implemented.
[0017] In the embodiments of the present disclosure, the configuration information of the target AI simulation cluster is determined. The configuration information includes: a list of host machines, network configuration information, simulation node configuration information, and cluster configuration information. A container network is created on each host machine included in the list of host machines according to the network configuration information. According to the container network and the simulation node configuration information, a plurality of simulation node containers are created on each host machine, and the simulation node information of each simulation node container is determined. According to the cluster configuration information and the simulation node information of each simulation node container, a target AI simulation cluster including the plurality of simulation node containers created on each host machine is created. Through container technology, it is possible to automatically create a specified number of simulation node containers on a small number of host machines according to the configuration, thereby effectively realizing AI cluster simulation by using a small amount of physical hardware resources, effectively saving physical hardware resources, and improving the efficiency of AI cluster simulation.
[0018] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and do not limit the present disclosure. Other features and aspects of the present disclosure will become clear from the following detailed description of exemplary embodiments with reference to the accompanying drawings. Brief Description of the Drawings
[0019] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification. These drawings illustrate embodiments consistent with the present disclosure and, together with the specification, are used to explain the technical solutions of the present disclosure.
[0020] Figure 1 A flowchart showing an AI cluster simulation method according to an embodiment of the present disclosure.
[0021] Figure 2 A block diagram showing an AI cluster simulation device according to an embodiment of the present disclosure.
[0022] Figure 3 A block diagram showing an electronic device according to an embodiment of the present disclosure. Detailed Description of the Embodiments
[0023] The following will detail various exemplary embodiments, features, and aspects of the present disclosure with reference to the accompanying drawings. The same reference numerals in the drawings denote elements having the same or similar functions. Although various aspects of the embodiments are shown in the drawings, the drawings are not necessarily drawn to scale unless otherwise specified.
[0024] The special term "exemplary" herein means "serving as an example, embodiment, or illustration". Any embodiment described as "exemplary" here does not necessarily need to be construed as superior to or better than other embodiments.
[0025] In this text, the term "and / or" is merely a description of the relationship between associated objects, indicating that there can be three relationships. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. Additionally, the term "at least one" in this text means any one of multiple types or any combination of at least two of multiple types. For example, including at least one of A, B, and C can mean including any one or more elements selected from the set composed of A, B, and C.
[0026] In addition, to better illustrate the present disclosure, numerous specific details are provided in the following specific embodiments. Those skilled in the art should understand that the present disclosure can still be implemented without certain specific details. In some instances, methods, means, elements, and circuits well-known to those skilled in the art are not described in detail to highlight the gist of the present disclosure.
[0027] In the prior art, to verify the upper-layer components of large-scale model training, it is usually necessary to build a real AI cluster. However, there are multiple problems in building and maintaining an actual AI cluster. 1. High consumption of hardware resources: Building a real AI cluster requires a large number of GPU servers and hardware resources. Compared with ordinary servers, GPU servers are expensive. Especially in the GPU development stage, these hardware resources are usually not sufficiently supplied, thus affecting the development and verification of upper-layer components. 2. Limitations in the R & D stage of GPU chips: In the R & D stage of new GPU chips, physical GPUs may not have been produced or have limited resources, resulting in the inability to verify the functions of GPU chips in a real AI cluster environment and delaying the R & D progress. 3. Uncontrollable number of cluster nodes: The existing verification solutions for AI clusters lack flexible control over the number of cluster nodes and are difficult to dynamically adjust the cluster scale according to different training tasks, restricting the verification and testing of different scenarios. These problems greatly increase the verification cost and difficulty of upper-layer components in the large model training process.
[0028] To solve the above technical problems, the embodiments of the present disclosure provide an AI cluster simulation method. Through container technology, it is possible to automatically create a specified number of simulation node containers on a small number of host machines according to the configuration, thereby effectively realizing AI cluster simulation using a small amount of physical hardware resources, effectively saving physical hardware resources, and improving the efficiency of AI cluster simulation. Moreover, through node simulation using container technology, it is possible to better achieve resource isolation and sharing, and physical resources can be more dynamically and flexibly allocated according to the situation of the simulation load to adjust the cluster scale. The AI cluster simulation method provided by the embodiments of the present disclosure will be described in detail below.
[0029] Figure 1The flowchart shows an AI cluster simulation method according to an embodiment of the present disclosure. This method can be executed by an electronic device such as a terminal device or a server. The terminal device can be a user equipment (UE), a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (PDA), a handheld device, a computing device, a vehicle-mounted device, a wearable device, etc. This method can be implemented by a processor calling computer-readable instructions stored in a memory. Alternatively, the server can execute this method. As Figure 1 shown, the method includes:
[0030] In step S11, determine the configuration information of the target AI simulation cluster, where the configuration information includes: a list of host machines, network configuration information, simulation node configuration information, and cluster configuration information.
[0031] The configuration information of the target AI simulation cluster can be flexibly configured by the user who needs to perform AI cluster simulation according to actual requirements.
[0032] In one example, the user constructs a configuration file according to actual requirements. The configuration file includes the configuration information of the target AI simulation cluster, and then passes the configuration file to the simulation software. The simulation software refers to the software product corresponding to the AI cluster simulation method provided in the embodiment of the present disclosure. The appropriate simulation software can be selected according to actual scenario requirements, and the present disclosure does not make specific limitations in this regard.
[0033] The simulation software includes a simulation configuration module. The simulation configuration module is responsible for receiving the configuration file and parsing and validating the configuration file to obtain the configuration information for constructing the target AI simulation cluster.
[0034] In one example, the simulation configuration module can check whether each configuration item of the configuration information conforms to a preset specification. If it does not conform to the preset specification, a prompt message can be generated and displayed.
[0035] In addition to the list of host machines, network configuration information, simulation node configuration information, and cluster configuration information, the configuration information of the target AI simulation cluster can also include other specific configuration information according to actual requirements. The present disclosure does not make specific limitations in this regard.
[0036] In step S12, create a container network on each host machine included in the list of host machines according to the network configuration information.
[0037] The configuration information of the target AI simulation cluster includes: a list of host machines. The list of host machines is used to indicate the actual physical hardware resources for constructing the target AI simulation cluster. In one example, the list of host machines is used to indicate multiple ordinary servers.
[0038] The configuration information of the target AI simulation cluster includes: network configuration information.
[0039] The simulation software includes a network configuration module. The network configuration module is responsible for connecting to each host included in the host list, initializing the network according to the network configuration information, and configuring the container network on each host to ensure that the simulation node containers created on each subsequent host can be connected to each other.
[0040] The following will, in combination with possible implementation manners of the present disclosure, describe in detail the specific process of creating a container network on each host included in the host list according to the network configuration information, which will not be elaborated here.
[0041] In step S13, according to the container network and the simulation node configuration information, multiple simulation node containers are created on each host, and the simulation node information of each simulation node container is determined.
[0042] The configuration information of the target AI simulation cluster includes: simulation node configuration information.
[0043] The simulation software includes a node simulation module. The node simulation module is responsible for creating multiple simulation node containers on each host according to the container network and the simulation node configuration information to implement node simulation. After each simulation node container is successfully started, the simulation node information of each simulation node container can be determined.
[0044] The following will, in combination with possible implementation manners of the present disclosure, describe in detail the specific process of creating multiple simulation node containers on each host according to the container network and the simulation node configuration information, which will not be elaborated here.
[0045] In step S14, according to the cluster configuration information and the simulation node information of each simulation node container, a target AI simulation cluster is created, where the target AI simulation cluster includes multiple simulation node containers created on each host.
[0046] The configuration information of the target AI simulation cluster includes: cluster configuration information.
[0047] The simulation software includes a cluster component module. The cluster component module is responsible for forming the multiple simulation node containers created on each host into a target AI simulation cluster through a cluster installation tool according to the cluster configuration information and the simulation node information of each simulation node container.
[0048] Without building an actual AI cluster, relevant components supporting large model training can be efficiently verified according to the simulated target AI simulation cluster.
[0049] The specific process of creating a target AI simulation cluster based on the cluster configuration information and the simulation node information of each simulation node container will be described in detail in combination with possible implementation manners of the present disclosure later, and will not be elaborated here.
[0050] According to an embodiment of the present disclosure, the configuration information of a target AI simulation cluster is determined. The configuration information includes: a list of host machines, network configuration information, simulation node configuration information, and cluster configuration information. A container network is created on each host machine included in the list of host machines according to the network configuration information. Multiple simulation node containers are created on each host machine according to the container network and the simulation node configuration information, and the simulation node information of each simulation node container is determined. A target AI simulation cluster including the multiple simulation node containers created on each host machine is created according to the cluster configuration information and the simulation node information of each simulation node container.
[0051] Through container technology, it is possible to fully automatically create a specified number of simulation node containers on a small number of host machines according to the configuration, so as to effectively implement AI cluster simulation by using a small amount of physical hardware resources, effectively saving the physical hardware resources and improving the efficiency of AI cluster simulation. Moreover, by performing node simulation through container technology, it is possible to better achieve resource isolation and sharing, and physical resources can be allocated more dynamically and flexibly according to the simulation load situation to adjust the cluster scale.
[0052] In a possible implementation manner, the list of host machines includes: connection information of each host machine. The method further includes: checking whether each host machine can be normally connected according to the connection information of each host machine.
[0053] The list of host machines includes: connection information of each host machine. After receiving the configuration information of the target AI simulation cluster, the simulation configuration module remotely accesses each host machine according to the connection information of each host machine included in the configuration information to check each host machine and confirm whether each host machine can be normally connected.
[0054] For any host machine, the connection information of the host machine may include: a host name, a host IP address, a user name, and a password. The host name is a unique identifier indicating the host machine; the host IP address is used to indicate the network address of the host machine; the user name is the user account name for logging in to the host machine; and the password is the password corresponding to the user name and is used for identity verification.
[0055] In a possible implementation, the network configuration information includes: network name, subnet mask, gateway, network type, IP address range; according to the network configuration information, a container network is created on each host included in the host list, including: for any host, a container network is created in the host according to the network name, subnet mask, gateway, network type, and IP address range.
[0056] The network configuration module traverses the host list and remotely accesses each host based on the connection information of each host. Furthermore, the network configuration module creates a container network in each host according to the network name, subnet mask, gateway, network type, and IP address range included in the network configuration information, so as to ensure that the simulation node containers created in each subsequent host can be connected to each other.
[0057] The network name is usually used to indicate the unique identifier of the container network. The container networks on each host use the same network name to ensure that different hosts can be connected to the network indicated by the same network name to achieve cross-host container communication.
[0058] The subnet mask is used to distinguish the network address and the host address, and it defines the size of the network and the range of assignable IP addresses. The container networks on each host use the same subnet mask to ensure that the IP addresses can be correctly assigned when creating simulation node containers in each host subsequently.
[0059] The gateway is the default exit point of the network and is used for the container to communicate with the external network. The container networks on each host use the same gateway.
[0060] The network type defines the topology and behavior of the network. The container networks on each host have a target network type, where the target network type can achieve cross-host container communication. For example, the target network type is an Overlay network.
[0061] The IP address range is used to indicate the range of IP addresses that can be assigned to containers in the network. This range should match the subnet mask and ensure that there is no conflict with other devices in the network.
[0062] For the specific process of creating a container network in a host according to the network name, subnet mask, gateway, network type, and IP address range, reference can be made to the container network creation process in related technologies, and this embodiment of the present disclosure does not make specific limitations thereon.
[0063] In a possible implementation, the host list includes: the number of target simulation nodes corresponding to each host; the simulation node configuration information includes: the container image address and the hardware configuration information; according to the container network and the simulation node configuration information, multiple simulation node containers are created on each host, and the simulation node information of each simulation node container is determined, including: for any one host, according to the IP address range, the container image address, and the hardware configuration information, the simulation node creation operation is executed multiple times on this host until the number of simulation node containers reaching the target number of simulation nodes is created on this host.
[0064] The host list also includes the number of target simulation nodes corresponding to each host. For any one host, the number of target simulation nodes corresponding to this host is used to indicate the upper limit of the number of simulation node containers that can be created in this host. For any one host, the specific value of the number of target simulation nodes corresponding to this host depends on the hardware resource configuration / performance of this host, and the present disclosure does not make specific limitations thereon.
[0065] In one example, for any one host, the higher the hardware resource configuration / performance of this host, the higher the value of the corresponding number of target simulation nodes; the lower the hardware resource configuration / performance of this host, the lower the value of the corresponding number of target simulation nodes.
[0066] The number of target simulation nodes corresponding to different hosts included in the host list may have the same value or different values, and the present disclosure does not make specific limitations thereon.
[0067] In one example, by flexibly setting the number of hosts included in the host list and the number of target simulation nodes corresponding to each host, it is possible to dynamically adjust the cluster scale according to different training tasks, so as to cope with the verification and testing of different scenarios.
[0068] The simulation node configuration information includes: the container image address and the hardware configuration information.
[0069] The node simulation module traverses the host list and remotely accesses each host based on the connection information of each host. Furthermore, for any one host, the node simulation module executes the simulation node creation operation multiple times on this host according to the IP address range configured by the container network and the container image address and the hardware configuration information included in the simulation node configuration information until the number of simulation node containers reaching the target number of simulation nodes is created on this host.
[0070] In a possible implementation, the simulation node information of each simulation node container includes: the node name and node IP address of the simulation node container; the simulation node creation operation includes: randomly generating a string, and determining the string as the node name of the target simulation node container, where the target simulation node container is any one of the multiple simulation node containers to be created in the host; selecting an idle IP address from the IP address range, and determining the idle IP address as the node IP address of the target simulation node container; obtaining the target container image from the container image address, and creating the target simulation node container by using the target container image; and performing corresponding hardware simulation in the target simulation node container according to the hardware configuration information.
[0071] For any host, the specific process of the node simulation module performing the simulation node creation operation on the host is as follows.
[0072] Randomly generate a string, and determine the string as the node name of the target simulation node container, where the target simulation node container is the simulation node container to be created in this simulation node creation operation. The method of randomly generating the string can refer to related technologies, and the present disclosure does not make specific limitations on this.
[0073] Select an idle IP address from the IP address range of the container network configuration, and determine the idle IP address as the node IP address of the target simulation node container. Among them, the idle IP address is the IP address that has not been assigned to any simulation node container.
[0074] Obtain the target container image from the container image address included in the simulation node configuration information, and create the target simulation node container by using the target container image. Among them, the target container image includes: simulation operating system files, hardware device simulation software and tools, etc.
[0075] Pass the hardware configuration information included in the simulation node configuration information to the target simulation node container through environment variables, so that in the target simulation node container, corresponding hardware simulation is performed by using the hardware device simulation software and tools.
[0076] In a possible implementation, the hardware configuration information includes: GPU model, number of GPUs, number of CPUs, and memory size.
[0077] According to actual application requirements, the hardware configuration information may include: GPU model, number of GPUs, number of CPUs, and memory size, so that in the constructed simulation node container, hardware simulation of GPUs, CPUs, memory, etc. is performed by using the hardware device simulation software and tools.
[0078] Among them, the specific parameter assignments of the GPU model, the number of GPUs, the number of CPUs, and the memory size can be flexibly set according to actual application requirements, and the present disclosure does not make specific limitations thereon.
[0079] By creating multiple simulation node containers on the host, it is possible to effectively utilize a small amount of physical hardware resources to implement large-scale AI cluster simulation, effectively saving physical hardware resources.
[0080] In addition, even in the GPU chip development stage when the physical GPU has not been completed or the resources are limited, the new GPU chip can be hardware simulated in the simulation node container in the above manner, thereby solving the problem that the GPU chip function cannot be verified in the real AI cluster environment and delaying the R & D progress.
[0081] In addition to the GPU model, the number of GPUs, the number of CPUs, and the memory size, the hardware configuration information can also include other hardware configuration information according to actual application requirements, and the present disclosure does not make specific limitations thereon.
[0082] In one example, for any one of the created simulation node containers, the node password corresponding to the node name of the simulation node container can be set for authentication, and the present disclosure does not make specific limitations thereon.
[0083] After performing the above simulation node creation operation and creating the target number of simulation node containers on each host, start all the simulation node containers, and send the simulation node information (the node name, node password, and node IP address of the simulation node container) of each simulation node container to the cluster component module to prepare for the subsequent construction of the target AI simulation cluster.
[0084] In one possible implementation, the method further includes: performing connectivity detection on each simulation node container according to the simulation node information of each simulation node container.
[0085] After receiving the simulation node information of all simulation node containers, the cluster component module can remotely connect to each simulation node container according to the simulation node information of each simulation node container for connectivity detection.
[0086] Among them, for any one of the simulation node containers, when the connectivity detection result of the simulation node container passes, it indicates that the simulation node container is successfully created and started and can be used to form the target AI simulation cluster.
[0087] In a possible implementation, the cluster configuration information includes: a cluster version number; creating a target AI simulation cluster according to the cluster configuration information and the simulation node information of each simulation node container, including: installing target software in each simulation node container according to the simulation node information of each simulation node container, and performing time zone synchronization processing on each simulation node container; determining a target cluster installation tool corresponding to the cluster version number according to the cluster version number; using the target cluster installation tool to create a target AI simulation cluster including each simulation node container, where the target AI simulation cluster corresponds to the cluster version indicated by the cluster version number.
[0088] The cluster component module installs target software in each simulation node container according to the simulation node information of each simulation node container, and performs time zone synchronization processing on each simulation node container.
[0089] The target software here can be some simple basic software that supports time zone synchronization and the internal operating system dependencies of the simulation node container. It is a system-level software. The specific type of the target software can be flexibly set according to the actual application scenario, and the present disclosure does not make specific limitations thereto.
[0090] Perform time zone synchronization processing on all simulation node containers to ensure that all simulation node containers use the same time zone, thereby avoiding communication problems and data inconsistencies caused by time zone differences.
[0091] The cluster configuration information includes: a cluster version number. The cluster version number is used to indicate the cluster version of the target AI simulation cluster to be built.
[0092] Different cluster versions may require different cluster installation tools. Therefore, use the target cluster installation tool corresponding to the cluster version number included in the cluster configuration information to create a target AI simulation cluster including all simulation node containers, thereby effectively forming the target AI simulation cluster of the version specified by the cluster version number.
[0093] In a possible implementation, the cluster configuration information includes: a network plugin, a component image address; using the target cluster installation tool to create a target AI simulation cluster including each simulation node container, including: using the target cluster installation tool to obtain a target component image from the component image address; using the target component image to install a target component in each simulation node container to add each simulation node container to the target AI simulation cluster; installing a network plugin in each simulation node container.
[0094] The cluster configuration information includes: a component image address. The target cluster installation tool can obtain the target component image from the component image address, and then use the target component image to install the target component in each simulation node container to add each simulation node container to the target AI simulation cluster.
[0095] The cluster configuration information includes: network plugins. Install network plugins in each simulation node container to provide network services for the loads in the target AI simulation cluster, thereby effectively enabling communication between processes in different simulation node containers in the target AI simulation cluster.
[0096] Since the target AI simulation cluster utilizes simulation node container components, the target AI simulation cluster can be a Kubernetes cluster, the cluster version number is the Kubernetes version number, and the target cluster installation tool is the Kubernetes installation tool corresponding to the Kubernetes version number. The target components and network plugins are both components and network plugins adapted to the Kubernetes cluster.
[0097] The process of creating the target AI simulation cluster based on the cluster component module can refer to the composition of building a Kubernetes cluster in related technologies, and the present disclosure does not make specific limitations on this.
[0098] In a possible implementation manner, the method further includes: determining the connection authentication information of the target AI simulation cluster; logging in to the target AI simulation cluster according to the connection authentication information of the target AI simulation cluster; and performing a functional test on the target AI simulation cluster by executing the target AI training task.
[0099] After the target AI simulation cluster is built, the connection authentication information of the target AI simulation cluster is generated. According to the connection authentication information of the target AI simulation cluster, the target AI simulation cluster can be logged in.
[0100] Furthermore, upper-layer components or loads for large-scale model training (AI training task) can be installed in the target AI simulation cluster for functional testing, thereby efficiently verifying the related components supporting large model training.
[0101] According to the embodiments of the present disclosure, determine the configuration information of the target AI simulation cluster, where the configuration information includes: a list of host machines, network configuration information, simulation node configuration information, and cluster configuration information; create a container network on each host machine included in the list of host machines according to the network configuration information; create multiple simulation node containers on each host machine according to the container network and the simulation node configuration information, and determine the simulation node information of each simulation node container; and create a target AI simulation cluster including the multiple simulation node containers created on each host machine according to the cluster configuration information and the simulation node information of each simulation node container.
[0102] Through container technology, it is possible to automatically create a specified number of simulation node containers on a small number of host machines according to the configuration, so as to effectively realize the AI cluster simulation by using a small amount of physical hardware resources, effectively save physical hardware resources, improve the efficiency of AI cluster simulation, and further improve the development and testing efficiency of large-scale training models. Moreover, through container technology for node simulation, it is possible to better achieve resource isolation and sharing, and physical resources can be allocated more dynamically and flexibly according to the simulation load situation to adjust the cluster scale.
[0103] The AI cluster simulation method based on containers can simulate the AI cluster fully automatically according to the configuration, reduce the difficulty of cluster simulation, and improve the efficiency of AI cluster simulation.
[0104] It can be understood that the above-mentioned various method embodiments mentioned in the present disclosure can be combined with each other to form a combined embodiment without violating the principle logic. Due to space limitations, the present disclosure will not elaborate further. Those skilled in the art can understand that in the above methods of the specific implementation manner, the specific execution order of each step should be determined according to its function and possible internal logic.
[0105] In addition, the present disclosure also provides an AI cluster simulation device, an electronic device, a computer-readable storage medium, and a program, all of which can be used to implement any one of the AI cluster simulation methods provided by the present disclosure. The corresponding technical solutions and descriptions are referred to the corresponding records in the method part and will not be elaborated further.
[0106] Figure 2 The block diagram of an AI cluster simulation device according to an embodiment of the present disclosure is shown. As Figure 2 shown, the device 20 includes:
[0107] A simulation configuration module 21, configured to determine the configuration information of the target AI simulation cluster, where the configuration information includes: a list of host machines, network configuration information, simulation node configuration information, and cluster configuration information;
[0108] A network configuration module 22, configured to create a container network on each host machine included in the list of host machines according to the network configuration information;
[0109] A node simulation module 23, configured to create a plurality of simulation node containers on each host machine according to the container network and the simulation node configuration information, and determine the simulation node information of each simulation node container;
[0110] A cluster formation module 24, configured to create a target AI simulation cluster according to the cluster configuration information and the simulation node information of each simulation node container, where the target AI simulation cluster includes a plurality of simulation node containers created on each host machine.
[0111] In a possible implementation, the host list includes: connection information of each host;
[0112] The apparatus 20 further includes:
[0113] A connection detection module, configured to check whether each host can be normally connected according to the connection information of each host.
[0114] In a possible implementation, the network configuration information includes: network name, subnet mask, gateway, network type, IP address range;
[0115] The network configuration module 22 is specifically configured to:
[0116] For any one host, create a container network in the host according to the network name, subnet mask, gateway, network type, and IP address range.
[0117] In a possible implementation, the host list includes: the number of target simulation nodes corresponding to each host; the simulation node configuration information includes: container image address, hardware configuration information;
[0118] The node simulation module 23 is specifically configured to:
[0119] For any one host, perform the simulation node creation operation multiple times on the host according to the IP address range, container image address, and hardware configuration information until the number of simulation node containers of the target simulation nodes is created on the host.
[0120] In a possible implementation, the simulation node information of each simulation node container includes: the node name of the simulation node container, the node IP address;
[0121] The node simulation module 23 is specifically configured to:
[0122] Randomly generate a string, and determine the string as the node name of the target simulation node container, where the target simulation node container is any one of the multiple simulation node containers that need to be created in the host;
[0123] Select an idle IP address from the IP address range, and determine the idle IP address as the node IP address of the target simulation node container;
[0124] Obtain the target container image from the container image address, and create the target simulation node container by using the target container image;
[0125] Perform corresponding hardware simulation in the target simulation node container according to the hardware configuration information.
[0126] In a possible implementation, the hardware configuration information includes: GPU model, number of GPUs, number of CPUs, and memory size.
[0127] In a possible implementation, the apparatus 20 further includes:
[0128] A connectivity detection module, configured to perform connectivity detection on each simulation node container according to the simulation node information of each simulation node container.
[0129] In a possible implementation, the cluster configuration information includes: cluster version number;
[0130] The cluster component module 24 is specifically configured to:
[0131] Install target software in each simulation node container according to the simulation node information of each simulation node container, and perform time zone synchronization processing on each simulation node container;
[0132] Determine a target cluster installation tool corresponding to the cluster version number according to the cluster version number;
[0133] Create a target AI simulation cluster including each simulation node container by using the target cluster installation tool, where the target AI simulation cluster corresponds to the cluster version indicated by the cluster version number.
[0134] In a possible implementation, the cluster configuration information includes: network plugin, component image address;
[0135] The cluster component module 24 is specifically configured to:
[0136] Obtain a target component image from the component image address by using the target cluster installation tool;
[0137] Install target components in each simulation node container by using the target component image to add each simulation node container to the target AI simulation cluster;
[0138] Install a network plugin in each simulation node container.
[0139] In a possible implementation, the apparatus 20 further includes:
[0140] A determination module, configured to determine connection authentication information of the target AI simulation cluster;
[0141] A login module, configured to log in to the target AI simulation cluster according to the connection authentication information of the target AI simulation cluster;
[0142] A functional testing module, configured to perform functional testing on the target AI simulation cluster for executing a target AI training task.
[0143] This method has a specific technical association with the internal structure of a computer system and can solve technical problems such as how to improve the operation efficiency or execution effect of hardware (including reducing the amount of data storage, reducing the amount of data transmission, and increasing the hardware processing speed), thereby obtaining a technical effect of improving the internal performance of the computer system that conforms to the laws of nature.
[0144] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the methods described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be repeated here.
[0145] The embodiments of the present disclosure also propose a computer-readable storage medium, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the above methods are implemented. The computer-readable storage medium can be a volatile or non-volatile computer-readable storage medium.
[0146] The embodiments of the present disclosure also propose an electronic device, including: a processor; a memory for storing instructions executable by the processor; wherein, the processor is configured to call the instructions stored in the memory to execute the above methods.
[0147] The embodiments of the present disclosure also provide a computer program product, including computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code. When the computer-readable code runs in the processor of the electronic device, the processor in the electronic device executes the above methods.
[0148] The electronic device can be provided as a terminal, a server, or other forms of devices.
[0149] Figure 3 A block diagram of an electronic device according to an embodiment of the present disclosure is shown. Referring to Figure 3 , the electronic device 1900 can be provided as a server or a terminal device. Referring to Figure 3 , the electronic device 1900 includes a processing component 1922, which further includes one or more processors, and memory resources represented by a memory 1932 for storing instructions executable by the processing component 1922, such as application programs. The application programs stored in the memory 1932 can include one or more modules each corresponding to a set of instructions. In addition, the processing component 1922 is configured to execute instructions to execute the above methods.
[0150] The electronic device 1900 may also include a power supply component 1926 configured to perform power management of the electronic device 1900, a wired or wireless network interface 1950 configured to connect the electronic device 1900 to a network, and an input / output interface 1958. The electronic device 1900 may operate based on an operating system stored in the memory 1932, such as the Microsoft server operating system (Windows Server TM ), the graphical user interface-based operating system launched by Apple Inc. (Mac OS X TM ), the multi-user and multi-process computer operating system (Unix TM ), the free and open-source Unix-like operating system (Linux TM ), the open-source Unix-like operating system (FreeBSD TM ), or the like.
[0151] In an exemplary embodiment, a non-volatile computer-readable storage medium is also provided, such as the memory 1932 including computer program instructions, and the above computer program instructions can be executed by the processing component 1922 of the electronic device 1900 to complete the above method.
[0152] The present disclosure may be a system, a method, and / or a computer program product. The computer program product may include a computer-readable storage medium having thereon computer-readable program instructions for causing a processor to implement various aspects of the present disclosure.
[0153] The computer-readable storage medium may be a tangible device that can retain and store instructions for use by an instruction execution device. The computer-readable storage medium may be, for example, (but is not limited to) an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanical encoding device, such as a punched card or raised structures in grooves having instructions stored thereon, and any suitable combination of the foregoing. The computer-readable storage medium used herein is not construed as an instantaneous signal itself, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagated through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.
[0154] The computer-readable program instructions described herein can be downloaded to various computing / processing devices from a computer-readable storage medium or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include copper transmission cables, optical fiber transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.
[0155] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state-setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet connection using an Internet service provider). In some embodiments, by using the state information of the computer-readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer-readable program instructions to implement various aspects of the present disclosure.
[0156] Aspects of the present disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0157] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that the instructions, when executed by the processor of the computer or other programmable data processing apparatus, create an apparatus that implements the functions / acts specified in one or more boxes of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer-readable medium storing the instructions comprises a manufacture including instructions that implement various aspects of the functions / acts specified in one or more boxes of the flowchart and / or block diagram.
[0158] The computer-readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device, such that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / acts specified in one or more boxes of the flowchart and / or block diagram.
[0159] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram may represent a module, a segment of a program, or a part of an instruction, and the module, segment of a program, or part of an instruction contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the boxes may occur in a different order than noted in the figures. For example, two consecutive boxes may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each box in the block diagrams and / or flowcharts, and combinations of boxes in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or acts, or by a combination of dedicated hardware and computer instructions.
[0160] The computer program product can be implemented specifically by hardware, software, or a combination thereof. In an alternative embodiment, the computer program product is embodied as a computer storage medium. In another alternative embodiment, the computer program product is embodied as a software product, such as a Software Development Kit (SDK), etc.
[0161] The above descriptions of the various embodiments tend to emphasize the differences between the various embodiments. The similarities or likenesses between them can be referred to each other. For the sake of brevity, they will not be elaborated herein.
[0162] Those skilled in the art will appreciate that, in the above method of specific implementation, the order in which the steps are written does not imply a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of the steps should be determined by their functions and possible internal logic.
[0163] If the technical solution of this application involves personal information, the product using the technical solution of this application has clearly informed the personal information processing rules and obtained the individual's voluntary consent before processing the personal information. If the technical solution of this application involves sensitive personal information, the product using the technical solution of this application has obtained the individual's separate consent before processing the sensitive personal information, and at the same time meets the "explicit consent" requirement. For example, on personal information collection devices such as cameras, clear and prominent signs are set to inform that the personal information collection scope has been entered and personal information will be collected. If the individual voluntarily enters the collection scope, it is deemed that he or she agrees to the collection of his or her personal information; or on the device that processes personal information, the personal information processing rules are notified by obvious signs / information, and the individual's authorization is obtained through pop-up information or by asking the individual to upload his or her personal information; among them, the personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the type of personal information processed.
[0164] The embodiments of the present disclosure have been described above, and the above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and changes will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The selection of terms used herein is intended to best explain the principles of the embodiments, practical applications, or improvements to the technology in the market, or to enable other persons of ordinary skill in the art to understand the embodiments disclosed herein.
Claims
1. An AI cluster simulation method, characterized in that, Including: Determine the configuration information of the target AI simulation cluster, where the configuration information includes: a list of host machines, network configuration information, simulation node configuration information, and cluster configuration information; Create a container network on each host machine included in the list of host machines according to the network configuration information; Create multiple simulation node containers on each host machine according to the container network and the simulation node configuration information, and determine the simulation node information of each simulation node container; Create the target AI simulation cluster according to the cluster configuration information and the simulation node information of each simulation node container, where the target AI simulation cluster includes multiple simulation node containers created on each host machine.
2. The method according to claim 1, characterized in that The list of host machines includes: connection information of each host machine; The method further includes: Check whether each host machine can be normally connected according to the connection information of each host machine.
3. The method according to claim 1, characterized in that, The network configuration information includes: network name, subnet mask, gateway, network type, IP address range; The step of creating a container network on each host machine included in the list of host machines according to the network configuration information includes: For any one host machine, create the container network in the host machine according to the network name, the subnet mask, the gateway, the network type, and the IP address range.
4. The method according to claim 3, characterized in that, The list of host machines includes: the number of target simulation nodes corresponding to each host machine; the simulation node configuration information includes: container image address, hardware configuration information; The step of creating multiple simulation node containers on each host machine according to the container network and the simulation node configuration information, and determining the simulation node information of each simulation node container includes: For any one host machine, perform the simulation node creation operation multiple times on the host machine according to the IP address range, the container image address, and the hardware configuration information until the number of simulation node containers of the target simulation nodes is created on the host machine.
5. The method according to claim 4, characterized in that, The simulation node information of each simulation node container includes: the node name of the simulation node container, the node IP address; The simulation node creation operation includes: Randomly generate a string, and determine the string as the node name of the target simulation node container, where the target simulation node container is any one of the multiple simulation node containers that need to be created in the host machine; Select an idle IP address from the IP address range, and determine the idle IP address as the node IP address of the target simulation node container; Obtain the target container image from the container image address, and create the target simulation node container by using the target container image; Perform corresponding hardware simulation in the target simulation node container according to the hardware configuration information.
6. The method according to claim 4 or 5, characterized in that, The hardware configuration information includes: GPU model, number of GPUs, number of CPUs, memory size.
7. The method according to claim 1, characterized in that, The method further includes: Perform connectivity detection on each simulation node container according to the simulation node information of each simulation node container.
8. The method according to claim 1, characterized in that, The cluster configuration information includes: cluster version number; Creating the target AI simulation cluster according to the cluster configuration information and the simulation node information of each simulation node container includes: Installing target software in each simulation node container according to the simulation node information of each simulation node container, and performing time zone synchronization processing on each simulation node container; Determining a target cluster installation tool corresponding to the cluster version number according to the cluster version number; Using the target cluster installation tool to create the target AI simulation cluster including each simulation node container, wherein the target AI simulation cluster corresponds to the cluster version indicated by the cluster version number.
9. The method according to claim 8, characterized in that, The cluster configuration information includes: network plugin, component image address; The using the target cluster installation tool to create the target AI simulation cluster including each simulation node container includes: Using the target cluster installation tool to obtain a target component image from the component image address; Installing target components in each simulation node container by using the target component image to add each simulation node container to the target AI simulation cluster; Installing the network plugin in each simulation node container.
10. The method according to claim 1, wherein The method further includes: Determining connection authentication information of the target AI simulation cluster; Logging in to the target AI simulation cluster according to the connection authentication information of the target AI simulation cluster; Performing a functional test on the target AI simulation cluster by executing a target AI training task.
11. An AI cluster simulation device, characterized in that, Including: A simulation configuration module, configured to determine configuration information of a target AI simulation cluster, wherein the configuration information includes: a list of host machines, network configuration information, simulation node configuration information, and cluster configuration information; A network configuration module, configured to create a container network on each host machine included in the list of host machines according to the network configuration information; A node simulation module, configured to create a plurality of simulation node containers on each host machine according to the container network and the simulation node configuration information, and determine simulation node information of each simulation node container; A cluster formation module, configured to create the target AI simulation cluster according to the cluster configuration information and the simulation node information of each simulation node container, wherein the target AI simulation cluster includes a plurality of simulation node containers created on each host machine.
12. An electronic device, characterized in that, Including: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to call the instructions stored in the memory to execute the method according to any one of claims 1 to 10.
13. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, The computer program instructions, when executed by the processor, implement the method according to any one of claims 1 to 10.
Citation Information
Cited By
Simulation method and platform of multi-machine multi-card AI computing cluster and storage medium
CN121841997A