High-availability cluster architecture artificial intelligence experiment cloud platform data processing method and system
Patent Information
- Application Number
- CN202211603530.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-13
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2042-12-13
AI Technical Summary
[0003]本申请实施例提供了一种高可用集群架构人工智能实验云平台数据处理方法及系统,旨在解决现有技术中进行人工智能相关实验使用的云平台集群中往往不能随时对集群增加或删减节点,这就导致人工智能相关实验面对的操作人员数量受限,只能开展少量人员参与的人工智能相关实验的问题
[0017]第四方面,本申请实施例还提供了一种计算机可读存储介质,其中计算机可读存储介质存储有计算机程序,计算机程序当被处理器执行时使处理器执行上述第一方面的高可用集群架构人工智能实验云平台数据处理方法。
Smart Images

Figure CN116192885B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of cloud platform technology, and in particular to a data processing method and system for a highly available cluster architecture artificial intelligence experimental cloud platform. Background Technology
[0002] Currently, some enterprises and universities are using experimental platform clusters to conduct AI-related experiments, placing experimental data on cloud platform clusters for cloud-based experiments. However, current cloud platform clusters often cannot easily add or remove nodes, limiting the number of operators available for AI experiments and making it impossible to handle large-scale cloud-based experiments. Furthermore, existing cloud platform clusters cannot automatically save experimental data in the event of power outages or other abnormal failures, posing significant data security risks. Summary of the Invention
[0003] This application provides a data processing method and system for a highly available cluster architecture artificial intelligence experimental cloud platform. It aims to solve the problem that in the existing technology, cloud platform clusters used for artificial intelligence-related experiments often cannot add or remove nodes from the cluster at any time. This results in a limited number of operators for artificial intelligence-related experiments, and only artificial intelligence-related experiments involving a small number of people can be carried out.
[0004] In a first aspect, embodiments of this application provide a data processing method for a highly available cluster architecture artificial intelligence experimental cloud platform, applied to an artificial intelligence experimental cloud platform, the artificial intelligence cloud platform including multiple master nodes and multiple slave nodes, all of which are communicatively connected; the method includes:
[0005] If a target slave node receives an experimental task deployment instruction sent by a target master node, it creates a target container according to the experimental task deployment instruction; wherein, the target slave node is any one of the plurality of slave nodes, and the target master node is the master node that is currently active among the plurality of master nodes;
[0006] If the target node receives an access request from the user terminal and passes the verification, it connects the target container instance corresponding to the target container to the user terminal.
[0007] The target receives target operation data from the user terminal from the node and stores the target operation data in the key-value database corresponding to the target container.
[0008] The target master node sends container operation commands to the target slave node;
[0009] If the target node receives the container operation instruction, it creates or deletes the container accordingly.
[0010] Secondly, this application provides a data processing system for a highly available cluster architecture artificial intelligence experimental cloud platform, which runs on the artificial intelligence experimental cloud platform and includes multiple master nodes and multiple slave nodes, all of which are communicatively connected; wherein, the target slave node is any one of the multiple slave nodes, and the target master node is the master node that is currently active among the multiple master nodes;
[0011] A target slave node is used to create a target container according to the experimental task deployment instruction sent by the target master node if it receives such an instruction. The target slave node is any one of the plurality of slave nodes, and the target master node is the master node that is currently active among the plurality of master nodes.
[0012] The target node is also used to connect the target container instance corresponding to the target container to the user terminal if it receives an access request from the user terminal and passes the verification.
[0013] The target slave node is also used to receive target operation data from the user terminal and store the target operation data in the key-value database corresponding to the target container;
[0014] The target master node is used to send container operation commands to the target slave node;
[0015] The target slave node is also used to create or delete a container according to the container operation instruction sent by the target master node if it receives the container operation instruction.
[0016] Thirdly, embodiments of this application provide a computer device including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the data processing method of the high-availability cluster architecture artificial intelligence experimental cloud platform described in the first aspect.
[0017] Fourthly, embodiments of this application also provide a computer-readable storage medium storing a computer program, which, when executed by a processor, causes the processor to perform the data processing method of the high-availability cluster architecture artificial intelligence experimental cloud platform described in the first aspect.
[0018] This application provides a data processing method and system for a highly available cluster architecture artificial intelligence experimental cloud platform. The artificial intelligence cloud platform includes multiple master nodes and multiple slave nodes, all of which are communicatively connected. The method includes: if a target slave node receives an experimental task deployment instruction sent by a target master node, it creates a target container according to the instruction; if a target slave node receives and verifies an access request from a user terminal, it connects the target container instance corresponding to the target container to the user terminal; the target slave node receives target operation data from the user terminal and stores the target operation data in a key-value database corresponding to the target container; the target master node sends container operation instructions to the target slave nodes; if a target slave node receives a container operation instruction, it creates or deletes a container according to the instruction. This enables the processing of artificial intelligence-related experimental tasks in the cloud within the artificial intelligence experimental cloud platform, and allows for the addition or removal of nodes from the cluster at any time, improving the high availability and load capacity of the cluster. Attached Figure Description
[0019] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0020] Figure 1 This is a schematic diagram illustrating an application scenario of the data processing method for a highly available cluster architecture artificial intelligence experimental cloud platform provided in this application embodiment;
[0021] Figure 2 A flowchart illustrating the data processing method for a highly available cluster architecture artificial intelligence experimental cloud platform provided in this application embodiment;
[0022] Figure 3 A schematic block diagram of a data processing system for a highly available cluster architecture artificial intelligence experimental cloud platform provided in this application embodiment;
[0023] Figure 4 A schematic block diagram of a computer device provided in an embodiment of this application. Detailed Implementation
[0024] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0025] It should be understood that, when used in this specification and the appended claims, the terms "comprising" and "including" indicate the presence of the described features, integrals, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.
[0026] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the scope of the application. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.
[0027] It should also be further understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0028] Please see Figure 1 and Figure 2 , Figure 1 This is a schematic diagram illustrating an application scenario of the data processing method for a highly available cluster architecture artificial intelligence experimental cloud platform provided in this application embodiment; Figure 2 This is a flowchart illustrating the data processing method for a high-availability cluster architecture artificial intelligence experimental cloud platform provided in this application embodiment. The data processing method for a high-availability cluster architecture artificial intelligence experimental cloud platform provided in this application embodiment is applied to an artificial intelligence experimental cloud platform, such as... Figure 1 As shown, the AI experimental cloud platform includes multiple master nodes and multiple slave nodes, all of which are communicatively connected. The AI experimental cloud platform can be viewed as a Kubernetes cluster comprising multiple master nodes and multiple slave nodes, and is a distributed system capable of orchestrating and scheduling the resources of a single container cluster.
[0029] Each of the multiple master nodes includes an APIServer module (which can be understood as an interface module), a Scheduler module (which can be understood as a scheduling module), a Controller-Manager module (which can be understood as a management and control module), and a key-value database (which can be represented as an Etcd database). The APIServer module is used to notify slave nodes to perform operations such as creating, deleting, and stopping cluster resources based on the decisions of the master nodes. The Scheduler module is used to schedule Pods (Pods are the smallest unit that can be created and managed in the Kubernetes system (i.e., the K8S system), and are the smallest resource object model created or deployed by the user) based on the resource consumption of each slave node in the cluster corresponding to the AI experimental cloud platform. The Controller-Manager module is used to detect the health status of each master node and each slave node in the cluster corresponding to the AI experimental cloud platform. The key-value database is used to store various important configuration information within the cluster corresponding to the AI experimental cloud platform, as well as to persist various data resources within the cluster. Of the multiple master nodes, only one is running and active at any given time, designated as the Leader-Master-Node (only the Leader-Master-Node can provide services externally), while the other master nodes are in an inactive standby state. If the working master node (i.e., the Leader-Master-Node) malfunctions, the corresponding cluster of the AI experimental cloud platform will automatically select a new master node from the standby master nodes to immediately replace the malfunctioning master node and continue the current operation.
[0030] Each of the multiple slave nodes can be considered a worker node in the cluster corresponding to the AI experimental cloud platform. It is the node that actually executes AI experimental tasks and is also the runtime container responsible for running actual business and resources. In addition to providing the runtime environment for Pods, each slave node also has infrastructure for management and communication. Specifically, each slave node interacts with each of the multiple master nodes through the Kubelet component (a proxy component on the slave node). The Kubelet component periodically receives work tasks from the API-Server module of the master node to handle matters related to the entire lifecycle of Pods on the master node; moreover, Kubelet periodically reports all work information to the master node through the API-Server module of the master node. Different slave nodes communicate via network proxy through the Kube-proxy component (a network proxy component on the slave nodes of the Kubernetes cluster).
[0031] like Figure 2As shown, the data processing method of the high-availability cluster architecture artificial intelligence experimental cloud platform includes steps S101 to S105.
[0032] S101. If the target slave node receives an experimental task deployment instruction sent by the target master node, it creates a target container according to the experimental task deployment instruction; wherein, the slave node is any one of the plurality of slave nodes, and the target master node is the master node that is currently active among the plurality of master nodes.
[0033] In this embodiment, when an AI experimental cloud platform cluster is composed of multiple master nodes and multiple slave nodes, the AI experimental cloud platform can serve as a cloud platform for conducting AI experiments. Specifically, when a user of the target master node (such as a platform administrator logging into the AI experimental cloud platform's master node with administrator privileges) operates on the user interface to deploy experimental tasks on slave nodes, an experimental task deployment instruction is triggered. The experimental task deployment instruction generated in the target master node is sent to each slave node, and a target container is created in the target slave node based on the experimental task deployment instruction. The target master node is the currently active master node among the multiple master nodes to ensure that only one master node in the cluster is currently working and performing various data processing tasks.
[0034] After the target container is created on the target slave node based on the experimental task deployment instruction, the target image corresponding to the experimental task deployment instruction needs to be added to each target container to obtain each target container instance. Once the target container instances are obtained, each target container instance can be associated with a corresponding user terminal and experimental participant, allowing each participant to connect to the corresponding target container instance using their user terminal to process artificial intelligence experimental tasks. Each target container instance already contains a container runtime environment and model code. The container engine of the target container instance can provide a container runtime environment and create images for different needs; the private image repository integrates framework images such as TensorFlow, Caffe, and PyTorch required for experiments in the field of artificial intelligence, and also supports deep neural networks (DNN), convolutional neural networks (CNN), and object detection-related YoLoV1~V5 models.
[0035] As can be seen, the target container is created on the slave node, not the master node. This ensures that the slave node is the actual cloud device running the container within the AI experimental cloud platform, while the master node serves as the cloud device for unified monitoring and management of the slave nodes. Even if a slave node fails, the AI experimental cloud platform employs highly available clusters and highly available application deployments to mitigate the damage caused by node failures, ensuring the platform's high reliability.
[0036] In one embodiment, the method further includes the following steps before step S101:
[0037] The interface module in the target master node establishes a communication connection with the Kubelet proxy component of the target slave node.
[0038] In this embodiment, when building a highly available cluster architecture AI experimental cloud platform, it is necessary to first establish communication connections between multiple master nodes and multiple slave nodes. Specifically, each slave node establishes a communication connection with the interface module in the target master node based on the Kubelet proxy component. Thus, the target slave node, as one of the slave nodes, also establishes a communication connection with the interface module in the target master node based on the Kubelet proxy component. The Kubelet proxy component can be figuratively understood as the link for data interaction between the target master node and each slave node. The Kubelet component periodically receives work tasks from the API-Server module of the master node to handle matters related to the entire lifecycle of Pods on the master node; moreover, the Kubelet also periodically reports all work information to the master node through the API-Server module of the master node. Furthermore, each slave node included in the AI experimental cloud platform can access the Internet based on the Kube-proxy component or communicate with user terminals via the Internet.
[0039] In one embodiment, the interface module in the target master node establishes a communication connection with the Kubelet proxy component of the target slave node, including:
[0040] The Keepalived component of the target master node automatically configures the virtual IP address of the artificial intelligence experimental cloud platform through the virtual routing redundancy protocol;
[0041] The interface module of the target master node establishes a communication connection with the Kubelet proxy component block of the target slave node based on the virtual IP address.
[0042] In this embodiment, each master node has a Keepalived component and a HAproxy component. The Keepalived component is used to automatically configure the virtual IP address of the AI experimental cloud platform through the Virtual Router Redundancy Protocol (VRRP) to ensure that the AI experimental cloud platform has a unified virtual IP for external access. The HAproxy component is used to provide load balancing services for the slave nodes.
[0043] Once the Keepalived component of the target master node obtains the virtual IP address of the AI experimental cloud platform, it automatically configures itself to have the same virtual IP address as the AI experimental cloud platform. Furthermore, besides the interface module of the target master node establishing a communication connection with the Kubelet proxy component of the target slave node based on the virtual IP address, the remaining master nodes also establish communication connections with the Kubelet proxy component of the target slave node based on the virtual IP address when switching from standby to active state. Therefore, this architecture ensures high availability and high load capacity of the system.
[0044] In one embodiment, step S101 includes:
[0045] If the experimental task deployment instruction is a unified experimental task deployment instruction, then the target slave node obtains the first target image resource, GPU resource and data storage volume path corresponding to the unified experimental task deployment instruction, and the target slave node creates a target container according to the first target image resource, GPU resource and data storage volume path corresponding to the unified experimental task deployment instruction.
[0046] If the experimental task deployment instruction is a personalized container deployment instruction, the target slave node obtains the second target image resource corresponding to the personalized container deployment instruction, and the target slave node creates a target container according to the second target image resource corresponding to the personalized container deployment instruction and the pre-stored data storage volume path.
[0047] In this embodiment, at least three types of user accounts with different permissions are pre-defined in the AI experimental cloud platform: administrator user accounts, first-permission user accounts (such as teacher user accounts), and second-permission user accounts (such as student user accounts). Administrator user accounts have the authority to manage all data within the entire AI experimental cloud platform; for example, creating multiple namespaces across multiple slave nodes of the AI experimental cloud platform is one such permission. First-permission user accounts have the authority to create multiple containers within their corresponding namespaces for second-permission user accounts to log in and use. Second-permission user accounts only have the authority to log in to the corresponding container within the corresponding namespace to process AI experimental tasks.
[0048] In this system, an administrator-privileged user account can receive a list of user accounts to be created from a user corresponding to a first-privilege user account. The administrator, after logging into the master or slave node of the AI Experiment Cloud Platform, can then combine the teacher's name and the student's class name from the user's bill list to create a combined name. Subsequently, the administrator-privileged user account creates a namespace on the slave node within the AI Experiment Cloud Platform using this combined name. This can be visualized as the administrator creating a class-specific namespace for the class taught by the teacher corresponding to that teacher's name within the AI Experiment Cloud Platform. Furthermore, the administrator-privileged user account can also match the number of second-privilege user accounts (i.e., the total number of student names in the user's bill list is the same as the total number of second-privilege user accounts) to the student name list (or student ID list) included in the user's bill list.
[0049] Of course, when creating multiple containers within each namespace according to actual needs, it could be that an administrator-privileged user account creates a corresponding number of containers based on the requirements of the AI experiment task and the student name list included in the user billing list; or it could be that a teacher with first-privilege user accounts creates a corresponding number of containers based on the requirements of the AI experiment task and the student name list included in the user billing list. In the AI experiment cloud platform, the relevant information for the namespaces is stored on the master node, and the containers corresponding to each namespace are deployed on the slave nodes of the AI experiment cloud platform.
[0050] After the namespace for a specific class, such as Class A, is created on the slave node of the AI experimental cloud platform, the teacher corresponding to the first-authority user account can generate a unified experimental task deployment instruction based on the requirements of the AI experimental task. This unified experimental task deployment instruction specifically sets the first target image resource, GPU resource, and data storage volume path. Subsequently, the resource container layer creates containers corresponding to the first target image resource, GPU resource, and data storage volume path specified in the unified experimental task deployment instruction.
[0051] The first target image resource can be selected from one or more of the following: images integrating frameworks such as TensorFlow, Caffe, and PyTorch required for experiments in the field of artificial intelligence; or images integrating artificial intelligence-related neural network models such as deep neural networks (DNN), convolutional neural networks (CNN), and object detection networks such as YoLoV1 to V5. More specifically, the first target image resource can be selected from the TensorFlow framework, and a convolutional neural network (CNN) can be deployed on the TensorFlow framework.
[0052] Once the teacher corresponding to the first-authority user account generates a unified experimental task deployment instruction based on the requirements of the artificial intelligence experimental task, and completes the creation of multiple containers corresponding to the namespace in the resource container layer, the initial environment setup for the artificial intelligence experimental task is completed.
[0053] Of course, after the namespace for a class, such as Class A, is created on the target slave node of the AI experiment cloud platform, the student corresponding to the second-authority user account can generate a personalized container deployment instruction based on their individual needs for conducting AI experiments. This personalized container deployment instruction is then sent by the slave node of the AI experiment cloud platform to the user terminal used by the teacher corresponding to the first-authority user account. Once the teacher approves the personalized container deployment instruction on the user terminal, the target slave node of the AI experiment cloud platform creates a container based on the second target image resource corresponding to the personalized container deployment instruction and the pre-stored data storage volume path. Similarly, the second target image resource can be selected from one or more of the following: images integrating frameworks such as TensorFlow, Caffe, and PyTorch required for AI-related experiments; or images integrating AI-related neural network models such as Deep Neural Networks (DNN), Convolutional Neural Networks (CNN), and object detection networks such as YoLoV1 to V5. More specifically, the second target image resource can select the TensorFlow framework and deploy the YoLoV5 object detection network on the TensorFlow framework.
[0054] Furthermore, when a student with the second-level user account generates a personalized container deployment instruction based on their individual needs for conducting AI experiments, GPU resources are not allocated by default. This means that containers created on the target slave node based on the personalized container deployment instruction are ordinary server containers, not GPU server containers. However, if the personalized container deployment instruction requires the use of a GPU server container, the instruction is sent to the user terminal used by the teacher with the first-level user account. Once the teacher approves the personalized container deployment instruction on their user terminal, the target slave node of the AI experiment cloud platform creates a container based on the second target image resources, GPU resources, and pre-stored data storage volume path corresponding to the personalized container deployment instruction.
[0055] In one embodiment, the target obtains the first target image resource, GPU resource, and data storage volume path corresponding to the unified experimental task deployment instruction from the node, including:
[0056] If the target slave node detects a unified experimental task deployment instruction, it obtains the teaching progress information and teacher teaching tag set corresponding to the unified experimental task deployment instruction, and generates the first target image resource, GPU resource and data storage volume path corresponding to the unified experimental task deployment instruction based on the teaching progress information, teacher teaching tag set and preset resource calling strategy.
[0057] In this embodiment, when the target slave node detects a unified experimental task deployment instruction, it can parse the instruction to determine whether it includes teaching progress information (e.g., learning the first chapter, fifth section of the AA Artificial Intelligence course) and a teacher's teaching tag set (e.g., tags such as face recognition and convolutional neural networks). If the unified experimental task deployment instruction parses to obtain the teaching progress information and the teacher's teaching tag set, then container information such as the first target image resource, GPU resource, and data storage volume path required for container creation is generated based on the teaching progress information, the teacher's teaching tag set, and a preset resource allocation strategy. The resource allocation strategy can be understood as a pre-set mapping table containing several teaching progress information entries, teacher's teaching tag sets, and target image resources, GPU resources, and data storage volume paths. Using the teaching progress information and teacher's teaching tag set as search conditions, the corresponding target image resources, GPU resources, and data storage volume paths can be retrieved.
[0058] S102. If the target slave node receives an access request from the user terminal and passes the verification, it connects the target container instance corresponding to the target container to the user terminal.
[0059] In this embodiment, after the target container and its instance are deployed on the target slave node, they can be accessed by users to perform AI-related experimental tasks. When a user needs to access the target container on the slave node, an access request containing the user's account information is first sent to the target slave node. Once the target slave node verifies the access request, a communication connection is established between the target container instance and the user terminal. This allows the user terminal to access the target slave node to perform AI-related experimental tasks.
[0060] S103. The target slave node receives the target operation data from the user terminal and stores the target operation data in the key-value database corresponding to the target container.
[0061] In this embodiment, after a user terminal connects to a target container instance in a target slave node, the target container instance can receive the target operation data from the user terminal. To improve the data security of the target operation data, the target operation data can be stored in a key-value database (such as an Etcd database, which is a type of key-value database) corresponding to the target container in the target master node. This way, even if the target slave node fails and stops running, the operation data of each container within it is still stored in the key-value database of the target master node. When the target slave node recovers after troubleshooting, all data from the target slave node is retrieved from the key-value database of the target master node for breakpoint recovery.
[0062] In one embodiment, step S103 is followed by:
[0063] If the target master node detects that the current working state is abnormal, it will randomly select one of the remaining master nodes as the target master node.
[0064] In this embodiment, typically only one master node is currently active and processing data in the cluster, while other master nodes are in an inactive standby state. If the active target master node (i.e., Leader-Master-Node) malfunctions, the corresponding cluster of the AI experimental cloud platform will automatically select a new master node from the standby master nodes to immediately replace the malfunctioning node and continue the current operation. Therefore, the failure of any single master node in the cluster will not affect the operation of the entire cluster. If the active master node malfunctions, the cluster will automatically select a new master node from the standby master nodes to immediately replace the malfunctioning node and continue the current operation.
[0065] In one embodiment, step S103 is followed by:
[0066] If the target slave node detects that the current working state is abnormal, it obtains the current node container data and stores the current node container data in the key-value database of the target master node corresponding to the target slave node.
[0067] In this embodiment, when the target slave node is currently in an abnormal state, before it restarts for troubleshooting, the current node container data of the target slave node can be retrieved again from the target master node and stored in the key-value database corresponding to the target slave node in the target master node for data backup. In this key-value database, a persistent volume declaration (PVC) and a persistent volume (PV) are created for each user. When a user stops and closes the operation on the target container instance in the target slave node, or exits the target container instance due to a fault, the current node container data generated by the user's operations on the target container instance before this closing or exiting operation is automatically stored in the key-value database corresponding to the target slave node in the target master node for data backup. When the user re-enters the target container instance, the AI experimental cloud platform automatically retrieves the current node container data from the key-value database to restore the target container instance to its state at the time of the last exit, so that the user can continue to operate on the target container instance upon re-entry.
[0068] Since all node container data for each container is persisted to a dynamically created persistent volume, the AI experimental cloud platform can automatically save historical node container data to the persistent volume, whether the user actively exits the container instance or is passively exited due to a fault. This allows the user to access the historical node container data and continue operating on the container instance after re-entering the container.
[0069] S104. The target master node sends container operation instructions to the target slave node.
[0070] In this embodiment, in addition to manipulating the AI-related experimental data within the containers of the target slave node, users can also access the target master node (e.g., a user account with administrator privileges logged into the target master node) and trigger container operation commands such as adding or deleting containers. After triggering the container operation command to add or delete a container, the target master node sends the container operation command to the target slave node.
[0071] S105. If the target slave node receives the container operation instruction, it creates or deletes the container according to the container operation instruction.
[0072] In this embodiment, the AI experimental cloud platform allows for the addition and deletion of cluster nodes at any time based on container operation commands, enabling convenient scaling of Pods through the cluster controller. More specifically, when resources are scarce on a target slave node, more slave nodes are added, and containers are created within those slave nodes to achieve scaling.
[0073] This method enables the processing of AI-related experimental tasks in the cloud within an AI experimental cloud platform, and allows for the addition or removal of nodes from the cluster at any time, thereby improving the cluster's high availability and load capacity.
[0074] This application also provides a data processing system for a highly available cluster architecture artificial intelligence experimental cloud platform. This system is used to execute any embodiment of the aforementioned data processing method for a highly available cluster architecture artificial intelligence experimental cloud platform. Specifically, please refer to... Figure 3 , Figure 3 This is a schematic block diagram of the data processing system 100 of the high-availability cluster architecture artificial intelligence experimental cloud platform provided in the embodiments of this application.
[0075] Among them, such as Figure 3 As shown, the high-availability cluster architecture artificial intelligence experimental cloud platform data processing system 100 includes multiple master nodes 101 and multiple slave nodes 102. The target slave node is any one of the multiple slave nodes 102, and the target master node is the currently active master node among the multiple master nodes 101.
[0076] The target slave node is used to create a target container according to the experimental task deployment instruction sent by the target master node if it receives such instruction. The slave node is any one of the plurality of slave nodes, and the target master node is the master node that is currently active among the plurality of master nodes.
[0077] In this embodiment, when an AI experimental cloud platform cluster is composed of multiple master nodes and multiple slave nodes, the AI experimental cloud platform can serve as a cloud platform for conducting AI experiments. Specifically, when a user of the target master node (such as a platform administrator logging into the AI experimental cloud platform's master node with administrator privileges) operates on the user interface to deploy experimental tasks on slave nodes, an experimental task deployment instruction is triggered. The experimental task deployment instruction generated in the target master node is sent to each slave node, and a target container is created in the target slave node based on the experimental task deployment instruction. The target master node is the currently active master node among the multiple master nodes to ensure that only one master node in the cluster is currently working and performing various data processing tasks.
[0078] After the target container is created on the target slave node based on the experimental task deployment instruction, the target image corresponding to the experimental task deployment instruction needs to be added to each target container to obtain each target container instance. Once the target container instances are obtained, each target container instance can be associated with a corresponding user terminal and experimental participant, allowing each participant to connect to the corresponding target container instance using their user terminal to process artificial intelligence experimental tasks. Each target container instance already contains a container runtime environment and model code. The container engine of the target container instance can provide a container runtime environment and create images for different needs; the private image repository integrates framework images such as TensorFlow, Caffe, and PyTorch required for experiments in the field of artificial intelligence, and also supports deep neural networks (DNN), convolutional neural networks (CNN), and object detection-related YoLoV1~V5 models.
[0079] As can be seen, the target container is created on the slave node, not the master node. This ensures that the slave node is the actual cloud device running the container within the AI experimental cloud platform, while the master node serves as the cloud device for unified monitoring and management of the slave nodes. Even if a slave node fails, the AI experimental cloud platform employs highly available clusters and highly available application deployments to mitigate the damage caused by node failures, ensuring the platform's high reliability.
[0080] In one embodiment, the target master node is further configured to establish a communication connection with the Kubelet proxy component of the target slave node through the interface module in the target master node.
[0081] In this embodiment, when building a highly available cluster architecture AI experimental cloud platform, it is necessary to first establish communication connections between multiple master nodes and multiple slave nodes. Specifically, each slave node establishes a communication connection with the interface module in the target master node based on the Kubelet proxy component. Thus, the target slave node, as one of the slave nodes, also establishes a communication connection with the interface module in the target master node based on the Kubelet proxy component. The Kubelet proxy component can be figuratively understood as the link for data interaction between the target master node and each slave node. The Kubelet component periodically receives work tasks from the API-Server module of the master node to handle matters related to the entire lifecycle of Pods on the master node; moreover, the Kubelet also periodically reports all work information to the master node through the API-Server module of the master node. Furthermore, each slave node included in the AI experimental cloud platform can access the Internet based on the Kube-proxy component or communicate with user terminals via the Internet.
[0082] In one embodiment, the interface module in the target master node establishes a communication connection with the Kubelet proxy component of the target slave node, including:
[0083] The Keepalived component of the target master node automatically configures the virtual IP address of the artificial intelligence experimental cloud platform through the virtual routing redundancy protocol;
[0084] The interface module of the target master node establishes a communication connection with the Kubelet proxy component block of the target slave node based on the virtual IP address.
[0085] In this embodiment, each master node has a Keepalived component and a HAproxy component. The Keepalived component is used to automatically configure the virtual IP address of the AI experimental cloud platform through the Virtual Router Redundancy Protocol (VRRP) to ensure that the AI experimental cloud platform has a unified virtual IP for external access. The HAproxy component is used to provide load balancing services for the slave nodes.
[0086] Once the Keepalived component of the target master node obtains the virtual IP address of the AI experimental cloud platform, it automatically configures itself to have the same virtual IP address as the AI experimental cloud platform. Furthermore, besides the interface module of the target master node establishing a communication connection with the Kubelet proxy component of the target slave node based on the virtual IP address, the remaining master nodes also establish communication connections with the Kubelet proxy component of the target slave node based on the virtual IP address when switching from standby to active state. Therefore, this architecture ensures high availability and high load capacity of the system.
[0087] In one embodiment, the target slave node is further used for:
[0088] If the experimental task deployment instruction is a unified experimental task deployment instruction, then the target slave node obtains the first target image resource, GPU resource and data storage volume path corresponding to the unified experimental task deployment instruction, and the target slave node creates a target container according to the first target image resource, GPU resource and data storage volume path corresponding to the unified experimental task deployment instruction.
[0089] If the experimental task deployment instruction is a personalized container deployment instruction, the target slave node obtains the second target image resource corresponding to the personalized container deployment instruction, and the target slave node creates a target container according to the second target image resource corresponding to the personalized container deployment instruction and the pre-stored data storage volume path.
[0090] In this embodiment, at least three types of user accounts with different permissions are pre-defined in the AI experimental cloud platform: administrator user accounts, first-permission user accounts (such as teacher user accounts), and second-permission user accounts (such as student user accounts). Administrator user accounts have the authority to manage all data within the entire AI experimental cloud platform; for example, creating multiple namespaces across multiple slave nodes of the AI experimental cloud platform is one such permission. First-permission user accounts have the authority to create multiple containers within their corresponding namespaces for second-permission user accounts to log in and use. Second-permission user accounts only have the authority to log in to the corresponding container within the corresponding namespace to process AI experimental tasks.
[0091] In this system, an administrator-privileged user account can receive a list of user accounts to be created from a user corresponding to a first-privilege user account. The administrator, after logging into the master or slave node of the AI Experiment Cloud Platform, can then combine the teacher's name and the student's class name from the user's bill list to create a combined name. Subsequently, the administrator-privileged user account creates a namespace on the slave node within the AI Experiment Cloud Platform using this combined name. This can be visualized as the administrator creating a class-specific namespace for the class taught by the teacher corresponding to that teacher's name within the AI Experiment Cloud Platform. Furthermore, the administrator-privileged user account can also match the number of second-privilege user accounts (i.e., the total number of student names in the user's bill list is the same as the total number of second-privilege user accounts) to the student name list (or student ID list) included in the user's bill list.
[0092] Of course, when creating multiple containers within each namespace according to actual needs, it could be that an administrator-privileged user account creates a corresponding number of containers based on the requirements of the AI experiment task and the student name list included in the user billing list; or it could be that a teacher with first-privilege user accounts creates a corresponding number of containers based on the requirements of the AI experiment task and the student name list included in the user billing list. In the AI experiment cloud platform, the relevant information for the namespaces is stored on the master node, and the containers corresponding to each namespace are deployed on the slave nodes of the AI experiment cloud platform.
[0093] After the namespace for a specific class, such as Class A, is created on the slave node of the AI experimental cloud platform, the teacher corresponding to the first-authority user account can generate a unified experimental task deployment instruction based on the requirements of the AI experimental task. This unified experimental task deployment instruction specifically sets the first target image resource, GPU resource, and data storage volume path. Subsequently, the resource container layer creates containers corresponding to the first target image resource, GPU resource, and data storage volume path specified in the unified experimental task deployment instruction.
[0094] The first target image resource can be selected from one or more of the following: images integrating frameworks such as TensorFlow, Caffe, and PyTorch required for experiments in the field of artificial intelligence; or images integrating artificial intelligence-related neural network models such as deep neural networks (DNN), convolutional neural networks (CNN), and object detection networks such as YoLoV1 to V5. More specifically, the first target image resource can be selected from the TensorFlow framework, and a convolutional neural network (CNN) can be deployed on the TensorFlow framework.
[0095] Once the teacher corresponding to the first-authority user account generates a unified experimental task deployment instruction based on the requirements of the artificial intelligence experimental task, and completes the creation of multiple containers corresponding to the namespace in the resource container layer, the initial environment setup for the artificial intelligence experimental task is completed.
[0096] Of course, after the namespace for a class, such as Class A, is created on the target slave node of the AI experiment cloud platform, the student corresponding to the second-authority user account can generate a personalized container deployment instruction based on their individual needs for conducting AI experiments. This personalized container deployment instruction is then sent by the slave node of the AI experiment cloud platform to the user terminal used by the teacher corresponding to the first-authority user account. Once the teacher approves the personalized container deployment instruction on the user terminal, the target slave node of the AI experiment cloud platform creates a container based on the second target image resource corresponding to the personalized container deployment instruction and the pre-stored data storage volume path. Similarly, the second target image resource can be selected from one or more of the following: images integrating frameworks such as TensorFlow, Caffe, and PyTorch required for AI-related experiments; or images integrating AI-related neural network models such as Deep Neural Networks (DNN), Convolutional Neural Networks (CNN), and object detection networks such as YoLoV1 to V5. More specifically, the second target image resource can select the TensorFlow framework and deploy the YoLoV5 object detection network on the TensorFlow framework.
[0097] Furthermore, when a student with the second-level user account generates a personalized container deployment instruction based on their individual needs for conducting AI experiments, GPU resources are not allocated by default. This means that containers created on the target slave node based on the personalized container deployment instruction are ordinary server containers, not GPU server containers. However, if the personalized container deployment instruction requires the use of a GPU server container, the instruction is sent to the user terminal used by the teacher with the first-level user account. Once the teacher approves the personalized container deployment instruction on their user terminal, the target slave node of the AI experiment cloud platform creates a container based on the second target image resources, GPU resources, and pre-stored data storage volume path corresponding to the personalized container deployment instruction.
[0098] In one embodiment, the target obtains the first target image resource, GPU resource, and data storage volume path corresponding to the unified experimental task deployment instruction from the node, including:
[0099] If the target slave node detects a unified experimental task deployment instruction, it obtains the teaching progress information and teacher teaching tag set corresponding to the unified experimental task deployment instruction, and generates the first target image resource, GPU resource and data storage volume path corresponding to the unified experimental task deployment instruction based on the teaching progress information, teacher teaching tag set and preset resource calling strategy.
[0100] In this embodiment, when the target slave node detects a unified experimental task deployment instruction, it can parse the instruction to determine whether it includes teaching progress information (e.g., learning the first chapter, fifth section of the AA Artificial Intelligence course) and a teacher's teaching tag set (e.g., tags such as face recognition and convolutional neural networks). If the unified experimental task deployment instruction parses to obtain the teaching progress information and the teacher's teaching tag set, then container information such as the first target image resource, GPU resource, and data storage volume path required for container creation is generated based on the teaching progress information, the teacher's teaching tag set, and a preset resource allocation strategy. The resource allocation strategy can be understood as a pre-set mapping table containing several teaching progress information entries, teacher's teaching tag sets, and target image resources, GPU resources, and data storage volume paths. Using the teaching progress information and teacher's teaching tag set as search conditions, the corresponding target image resources, GPU resources, and data storage volume paths can be retrieved.
[0101] The target node is also used to connect the target container instance corresponding to the target container to the user terminal if it receives an access request from the user terminal and passes the verification.
[0102] In this embodiment, after the target container and its instance are deployed on the target slave node, they can be accessed by users to perform AI-related experimental tasks. When a user needs to access the target container on the slave node, an access request containing the user's account information is first sent to the target slave node. Once the target slave node verifies the access request, a communication connection is established between the target container instance and the user terminal. This allows the user terminal to access the target slave node to perform AI-related experimental tasks.
[0103] The target slave node is also used to receive target operation data from the user terminal and store the target operation data in the key-value database corresponding to the target container.
[0104] In this embodiment, after a user terminal connects to a target container instance in a target slave node, the target container instance can receive the target operation data from the user terminal. To improve the data security of the target operation data, the target operation data can be stored in a key-value database (such as an Etcd database, which is a type of key-value database) corresponding to the target container in the target master node. This way, even if the target slave node fails and stops running, the operation data of each container within it is still stored in the key-value database of the target master node. When the target slave node recovers after troubleshooting, all data from the target slave node is retrieved from the key-value database of the target master node for breakpoint recovery.
[0105] In one embodiment, the target master node is further configured to randomly select one master node from the remaining multiple master nodes as the target master node if the current working state is detected to be an abnormal state.
[0106] In this embodiment, typically only one master node is currently active and processing data in the cluster, while other master nodes are in an inactive standby state. If the active target master node (i.e., Leader-Master-Node) malfunctions, the corresponding cluster of the AI experimental cloud platform will automatically select a new master node from the standby master nodes to immediately replace the malfunctioning node and continue the current operation. Therefore, the failure of any single master node in the cluster will not affect the operation of the entire cluster. If the active master node malfunctions, the cluster will automatically select a new master node from the standby master nodes to immediately replace the malfunctioning node and continue the current operation.
[0107] In one embodiment, the target slave node is further configured to, if an abnormal working state is detected, acquire the current node container data and store the current node container data in the key-value database of the target master node corresponding to the target slave node.
[0108] In this embodiment, when the target slave node is currently in an abnormal state, before it restarts for troubleshooting, the current node container data of the target slave node can be retrieved again from the target master node and stored in the key-value database corresponding to the target slave node in the target master node for data backup. In this key-value database, a persistent volume declaration (PVC) and a persistent volume (PV) are created for each user. When a user stops and closes the operation on the target container instance in the target slave node, or exits the target container instance due to a fault, the current node container data generated by the user's operations on the target container instance before this closing or exiting operation is automatically stored in the key-value database corresponding to the target slave node in the target master node for data backup. When the user re-enters the target container instance, the AI experimental cloud platform automatically retrieves the current node container data from the key-value database to restore the target container instance to its state at the time of the last exit, so that the user can continue to operate on the target container instance upon re-entry.
[0109] Since all node container data for each container is persisted to a dynamically created persistent volume, the AI experimental cloud platform can automatically save historical node container data to the persistent volume, whether the user actively exits the container instance or is passively exited due to a fault. This allows the user to access the historical node container data and continue operating on the container instance after re-entering the container.
[0110] The target master node is used to send container operation commands to the target slave node.
[0111] In this embodiment, in addition to manipulating the AI-related experimental data within the containers of the target slave node, users can also access the target master node (e.g., a user account with administrator privileges logged into the target master node) and trigger container operation commands such as adding or deleting containers. After triggering the container operation command to add or delete a container, the target master node sends the container operation command to the target slave node.
[0112] The target node is also configured to create or delete a container according to the container operation instruction if the container operation instruction is received.
[0113] In this embodiment, the AI experimental cloud platform allows for the addition and deletion of cluster nodes at any time based on container operation commands, enabling convenient scaling of Pods through the cluster controller. More specifically, when resources are scarce on a target slave node, more slave nodes are added, and containers are created within those slave nodes to achieve scaling.
[0114] This system enables the processing of AI-related experimental tasks in the cloud within an AI experimental cloud platform, and allows for the addition or removal of nodes from the cluster at any time, thereby improving the cluster's high availability and load capacity.
[0115] The aforementioned high-availability cluster architecture artificial intelligence experimental cloud platform data processing system can be implemented as a computer program, which can be used in, for example... Figure 4 It runs on the computer device shown.
[0116] Please see Figure 4 , Figure 4 This is a schematic block diagram of a computer device provided in an embodiment of this application. The computer device 500 is a server, or it can be a server cluster. The server can be a standalone server, or it can be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms.
[0117] See Figure 4 The computer device 500 includes a processor 502, a memory, and a network interface 505 connected via a device bus 501, wherein the memory may include a storage medium 503 and internal memory 504.
[0118] The storage medium 503 can store the operating system 5031 and the computer program 5032. When the computer program 5032 is executed, it enables the processor 502 to execute the data processing method of the high-availability cluster architecture artificial intelligence experimental cloud platform.
[0119] The processor 502 provides computing and control capabilities to support the operation of the entire computer device 500.
[0120] The internal memory 504 provides an environment for the computer program 5032 in the storage medium 503 to run. When the computer program 5032 is executed by the processor 502, the processor 502 can execute the data processing method of the high-availability cluster architecture artificial intelligence experimental cloud platform.
[0121] This network interface 505 is used for network communication, such as providing data transmission. Those skilled in the art will understand that... Figure 4The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device 500 to which the present application is applied. The specific computer device 500 may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0122] The processor 502 is used to run the computer program 5032 stored in the memory to implement the data processing method of the high-availability cluster architecture artificial intelligence experimental cloud platform disclosed in the embodiments of this application.
[0123] Those skilled in the art will understand that Figure 4 The embodiments of the computer device shown do not constitute a limitation on the specific configuration of the computer device. In other embodiments, the computer device may include more or fewer components than illustrated, or combine certain components, or have different component arrangements. For example, in some embodiments, the computer device may include only memory and a processor. In such embodiments, the structure and function of the memory and processor are different from those shown. Figure 4 The embodiments shown are consistent and will not be repeated here.
[0124] It should be understood that in the embodiments of this application, the processor 502 may be a central processing unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.
[0125] In another embodiment of this application, a computer-readable storage medium is provided. This computer-readable storage medium can be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. The computer-readable storage medium stores a computer program, wherein when executed by a processor, the computer program implements the data processing method for the high-availability cluster architecture artificial intelligence experimental cloud platform disclosed in the embodiments of this application.
[0126] Those skilled in the art will readily understand that, for the sake of convenience and brevity, the specific working processes of the devices, apparatuses, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in terms of function in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0127] In the several embodiments provided in this application, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. Units with the same function may be grouped into one unit. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces, devices, or units, or it may be an electrical, mechanical, or other form of connection.
[0128] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of the embodiments of this application, depending on actual needs.
[0129] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0130] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, a backend server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), magnetic disks, or optical disks.
[0131] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A data processing method for a highly available cluster architecture artificial intelligence experimental cloud platform, applied to an artificial intelligence experimental cloud platform, characterized in that, The artificial intelligence experimental cloud platform includes multiple master nodes and multiple slave nodes, all of which are communicatively connected; the method includes: If a target slave node receives an experimental task deployment instruction sent by a target master node, it creates a target container according to the experimental task deployment instruction; wherein, the target slave node is any one of the plurality of slave nodes, and the target master node is the master node that is currently active among the plurality of master nodes; If the target node receives an access request from the user terminal and passes the verification, it connects the target container instance corresponding to the target container to the user terminal. The target receives target operation data from the user terminal from the node and stores the target operation data in the key-value database corresponding to the target container. The target master node sends container operation commands to the target slave node; If the target node receives the container operation instruction, it creates or deletes the container accordingly based on the container operation instruction. The step of creating the target container according to the deployment instructions of the experimental task includes: If the experimental task deployment instruction is a unified experimental task deployment instruction, then the target slave node obtains the first target image resource, GPU resource and data storage volume path corresponding to the unified experimental task deployment instruction, and the target slave node creates a target container according to the first target image resource, GPU resource and data storage volume path corresponding to the unified experimental task deployment instruction. After receiving the target operation data from the user terminal from the target slave node and storing the target operation data in the key-value database corresponding to the target container, the method further includes: If the target slave node detects that the current working state is abnormal, it obtains the current node container data and stores the current node container data in the key-value database of the target master node corresponding to the target slave node. The target obtains the first target image resource, GPU resource, and data storage volume path corresponding to the unified experimental task deployment instruction from the node, including: If the target slave node detects a unified experimental task deployment instruction, it obtains the teaching progress information and teacher teaching tag set corresponding to the unified experimental task deployment instruction, and generates the first target image resource, GPU resource and data storage volume path corresponding to the unified experimental task deployment instruction based on the teaching progress information, teacher teaching tag set and preset resource calling strategy. The AI experiment cloud platform is pre-defined with administrator-privileged user accounts, first-privileged user accounts corresponding to teacher-privileged user accounts, and second-privileged user accounts corresponding to student-privileged user accounts. Administrator-privileged user accounts have the authority to manage all data on the entire AI experiment cloud platform. First-privileged user accounts have the authority to create multiple containers in the corresponding namespace for second-privileged user accounts to log in and use. Second-privileged user accounts have the authority to log in to the corresponding containers in the corresponding namespace to process AI experiment tasks. The unified experiment task deployment instruction is generated by the teacher corresponding to the first-privileged user account according to the needs of the AI experiment task.
2. The data processing method for a high-availability cluster architecture artificial intelligence experimental cloud platform according to claim 1, characterized in that, The step of creating the target container according to the deployment instructions of the experimental task also includes: If the experimental task deployment instruction is a personalized container deployment instruction, the target slave node obtains the second target image resource corresponding to the personalized container deployment instruction, and the target slave node creates a target container according to the second target image resource corresponding to the personalized container deployment instruction and the pre-stored data storage volume path.
3. The data processing method for a high-availability cluster architecture artificial intelligence experimental cloud platform according to claim 1, characterized in that, If the target slave node receives an experimental task deployment instruction sent by the target master node, before creating the target container according to the experimental task deployment instruction, the process further includes: The interface module in the target master node establishes a communication connection with the Kubelet proxy component of the target slave node.
4. The data processing method for a high-availability cluster architecture artificial intelligence experimental cloud platform according to claim 3, characterized in that, The interface module in the target master node establishes a communication connection with the Kubelet proxy component of the target slave node, including: The Keepalived component of the target master node automatically configures the virtual IP address of the artificial intelligence experimental cloud platform through the virtual routing redundancy protocol; The interface module of the target master node establishes a communication connection with the Kubelet proxy component block of the target slave node based on the virtual IP address.
5. The data processing method for a high-availability cluster architecture artificial intelligence experimental cloud platform according to claim 1, characterized in that, After receiving the target operation data from the user terminal from the target slave node and storing the target operation data in the key-value database corresponding to the target container, the method further includes: If the target master node detects that the current working state is abnormal, it will randomly select one of the remaining master nodes as the target master node.
6. A data processing system for a highly available cluster architecture artificial intelligence experimental cloud platform, running on an artificial intelligence experimental cloud platform, characterized in that, It includes multiple master nodes and multiple slave nodes, all of which are communicatively connected; A target slave node is used to create a target container according to the experimental task deployment instruction sent by the target master node if it receives such an instruction. The target slave node is any one of the plurality of slave nodes, and the target master node is the master node that is currently active among the plurality of master nodes. The target node is also used to connect the target container instance corresponding to the target container to the user terminal if it receives an access request from the user terminal and passes the verification. The target slave node is also used to receive target operation data from the user terminal and store the target operation data in the key-value database corresponding to the target container; The target master node is used to send container operation commands to the target slave node; The target slave node is also used to create or delete a container according to the container operation instruction sent by the target master node if it receives the container operation instruction. The target is created from the node according to the deployment instructions of the experimental task, including: If the experimental task deployment instruction is a unified experimental task deployment instruction, then the target slave node obtains the first target image resource, GPU resource and data storage volume path corresponding to the unified experimental task deployment instruction, and the target slave node creates a target container according to the first target image resource, GPU resource and data storage volume path corresponding to the unified experimental task deployment instruction. After receiving the target operation data from the user terminal and storing the target operation data in the key-value database corresponding to the target container, the method further includes: If the target slave node detects that the current working state is abnormal, it obtains the current node container data and stores the current node container data in the key-value database of the target master node corresponding to the target slave node. The target obtains the first target image resource, GPU resource, and data storage volume path corresponding to the unified experimental task deployment instruction from the node, including: If the target slave node detects a unified experimental task deployment instruction, it obtains the teaching progress information and teacher teaching tag set corresponding to the unified experimental task deployment instruction, and generates the first target image resource, GPU resource and data storage volume path corresponding to the unified experimental task deployment instruction based on the teaching progress information, teacher teaching tag set and preset resource calling strategy. The AI experiment cloud platform is pre-defined with administrator-privileged user accounts, first-privileged user accounts corresponding to teacher-privileged user accounts, and second-privileged user accounts corresponding to student-privileged user accounts. Administrator-privileged user accounts have the authority to manage all data on the entire AI experiment cloud platform. First-privileged user accounts have the authority to create multiple containers in the corresponding namespace for second-privileged user accounts to log in and use. Second-privileged user accounts have the authority to log in to the corresponding containers in the corresponding namespace to process AI experiment tasks. The unified experiment task deployment instruction is generated by the teacher corresponding to the first-privileged user account according to the needs of the AI experiment task.
7. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the data processing method for a high-availability cluster architecture artificial intelligence experimental cloud platform as described in any one of claims 1 to 5.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, causes the processor to perform the data processing method for the high-availability cluster architecture artificial intelligence experimental cloud platform as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Online experiment method and device
CN111176782A
Task processing method and device based on container cluster
CN115080207A