Artificial intelligence platform and method, device, and electronic device for expansion and update
By using update modules in the artificial intelligence platform to obtain and copy target processor files and start plug-in services, the problem of the existing technology being unable to expand and not supporting model GPU nodes is solved, and rapid adaptation and efficient expansion are achieved, improving the user experience.
Patent Information
- Application Number
- CN202510230901.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-02-28
AI Technical Summary
The existing capacity expansion technology cannot expand the capacity of nodes of GPUs that do not support models, and cannot be suitable for scenes where on-site node environment changes, resulting in poor user experience.
By introducing update modules into the artificial intelligence platform, obtaining the processor files and replacement files of the target processor, copying these files to a fixed directory, and starting the plug-in service of the target processor by executing the plug-in startup file, thereby achieving capacity expansion of GPU nodes that do not support models.
It realizes the rapid expansion process of GPU nodes that do not support models without appearance compatibility, ensuring that GPU nodes that do not support models can be expanded, improving the user experience.
Smart Images

Figure CN119718388B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and particularly to an artificial intelligence platform, an expansion and update method, device, and electronic device. Background Art
[0002] Currently, both the training and inference tasks of artificial intelligence require the participation of a Graphics Processing Unit (GPU). Therefore, in a computer cluster, a computer with a graphics processor is a necessary condition. When the GPU of the user's computer cluster changes, an expansion technology is required to expand the artificial intelligence platform to support the training and inference tasks of the computer cluster.
[0003] When the user's computer cluster adopts a brand-new GPU, it is possible that the new GPU is not in the support list when the artificial intelligence platform is installed. The existing expansion technology cannot expand the nodes with GPUs of unsupported models and is not applicable to the scenario of on-site node environment changes, resulting in a poor user experience. Summary of the Invention
[0004] The present application provides an artificial intelligence platform, an expansion and update method, device, and electronic device, so as to at least solve the problem in the related art that the existing expansion technology cannot expand the nodes with GPUs of unsupported models and is not applicable to the scenario of on-site node environment changes, resulting in a poor user experience.
[0005] The present application provides an artificial intelligence platform, including: a management node,
[0006] The management node includes an update module, a software package module, and a database;
[0007] The update module is configured to obtain the processor file of the target processor from the software package module and obtain a plurality of replacement files corresponding to the target processor from the database;
[0008] The update module is further configured to copy the plurality of replacement files and the processor file to their respective fixed directories corresponding to the plurality of replacement files and the processor file; wherein, the management node further includes a plurality of fixed directories;
[0009] The update module is further configured to, when there are the plurality of replacement files and the processor file corresponding to each in the plurality of fixed directories, start the plugin service of the target processor by executing the plugin startup file; wherein, the processor file at least includes the plugin startup file;
[0010] The management node is configured to expand the target processor to the artificial intelligence platform based on the started plugin service.
[0011] The present application provides an expansion and update method, which is applied to the above-mentioned artificial intelligence platform and includes:
[0012] Obtain the processor file of the target processor and multiple replacement files corresponding to the target processor; wherein, the multiple replacement files are files obtained in advance for expanding the target processor;
[0013] Copy the multiple replacement files corresponding to the target processor and the processor file to their respective fixed directories corresponding to the multiple replacement files and the processor file respectively; wherein, the fixed directory is a directory set in advance for storing files;
[0014] When there are multiple replacement files and processor files corresponding to them in multiple fixed directories, start the plugin service of the target processor by executing the plugin startup file; wherein, the processor file at least includes the plugin startup file;
[0015] Based on the started plugin service, expand the target processor to the artificial intelligence platform.
[0016] The present application also provides an expansion and update device, including:
[0017] An acquisition unit, configured to obtain the processor file of the target processor and multiple replacement files corresponding to the target processor; wherein, the multiple replacement files are files obtained in advance for expanding the target processor;
[0018] A copying unit, configured to copy the multiple replacement files corresponding to the target processor and the processor file to their respective fixed directories corresponding to the multiple replacement files and the processor file respectively; wherein, the fixed directory is a directory set in advance for storing files;
[0019] A startup unit, configured to start the plugin service of the target processor by executing the plugin startup file when there are multiple replacement files and processor files corresponding to them in multiple fixed directories; wherein, the processor file at least includes the plugin startup file;
[0020] An expansion unit, configured to expand the target processor to the artificial intelligence platform based on the started plugin service.
[0021] The present application also provides an electronic device, including: a memory for storing a computer program; a processor for implementing the steps of any of the above-mentioned expansion and update methods when executing the computer program.
[0022] The present application also provides a computer-readable storage medium, in which a computer program is stored, and wherein the computer program implements the steps of any of the above-mentioned expansion and update methods when executed by a processor.
[0023] The present application also provides a computer program product, including a computer program, which implements the steps of any of the above expansion and update methods when executed by a processor.
[0024] Through an artificial intelligence platform, an expansion and update method, a device, and an electronic device according to the present application, in the process of updating and expanding the artificial intelligence platform, various expansion files of the target processor, namely, a plurality of replacement files and the processor files related to the target processor, are copied to a fixed directory through an update module. After that, various expansion files of the target processor exist in the fixed directory, and the plug-in service corresponding to the target processor can be started to ensure that after the target processor is expanded in the artificial intelligence platform, the functions corresponding to the target processor can run normally, and the expansion of the target processor is completed. Therefore, the technical problem that the existing expansion technology cannot expand the nodes of GPUs of unsupported models and cannot be applied to the scenario of on-site node environment change, resulting in a poor user experience, can be solved, and the technical effect of quickly updating the expansion process of the nodes of GPUs of unsupported models without factory compatibility and ensuring that the nodes of GPUs of unsupported models can be expanded can be achieved. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] To more clearly illustrate the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0026] Figure 1 It is a schematic structural diagram of an artificial intelligence platform provided by an embodiment of the present application;
[0027] Figure 2 It is a schematic diagram of modules in a management node provided by an embodiment of the present application;
[0028] Figure 3 It is another schematic structural diagram of an artificial intelligence platform provided by an embodiment of the present application;
[0029] Figure 4 It is a schematic flowchart of an expansion and update method provided by an embodiment of the present application;
[0030] Figure 5 It is a schematic structural diagram of an expansion and update device provided by an embodiment of the present application;
[0031] Figure 6 It is another schematic structural diagram of an expansion and update device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0032] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present application.
[0033] It should be noted that in the description of the present application, the terms "including", "comprising" or any other variant thereof are intended to cover a non-exclusive inclusion, such that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects and are not used to describe a specific order or sequence.
[0034] In order to enable those skilled in the art of this technology to better understand the solution of the present application, the present application will be further described in detail below in conjunction with the accompanying drawings and specific embodiments.
[0035] Combined with the specific application environment architecture or specific hardware architecture on which the execution of the expansion and update method depends, the specific application environment architecture or specific hardware architecture will be described herein.
[0036] Figure 1 It is a schematic structural diagram of an artificial intelligence platform provided by the present application, and the execution of the expansion and update method depends on the artificial intelligence platform.
[0037] As Figure 1 shown, the artificial intelligence platform includes: a management node 11,
[0038] The management node 11 includes an update module 111, a software package module 112, and a database 113;
[0039] The update module 111 is used to obtain the processor file of the target processor from the software package module 112, and obtain a plurality of replacement files corresponding to the target processor from the database 113;
[0040] The update module 111 is further used to copy the plurality of replacement files and the processor file to their respective fixed directories corresponding to the plurality of replacement files and the processor file; wherein, the management node 11 further includes a plurality of fixed directories;
[0041] The update module 111 is further used to start the plugin service of the target processor by executing the plugin startup file in the case where there are the plurality of replacement files and the processor file corresponding to each of the plurality of fixed directories; wherein, the processor file includes the plugin startup file;
[0042] The management node 11 is used to expand the target processor to the artificial intelligence platform based on the started plug-in service.
[0043] Among them, the artificial intelligence platform is the artificial intelligence training platform, which is used to implement the training and inference tasks of artificial intelligence of the computer cluster. The artificial intelligence platform is deployed in a computer cluster composed of many computers (referred to as nodes), and generally includes one or an odd number (3 or 5) of management nodes 11, and the rest are computing nodes 12. The management node 11 deploys resource management services, databases, cluster monitoring, user management services, and image repositories (harbor) (preset image file databases), etc. The computing node 12 deploys node services, etc. The development environment is usually run on the computing node 12, and the development environment is the infrastructure provided for users to conduct artificial intelligence training. The main functions of the node service include but are not limited to: collecting node information, reporting cluster monitoring, resource management and other services. Users can perform artificial intelligence training tasks on the artificial intelligence platform. Specifically, for the relationship between the update module 111 and the software package module 112, reference can be made to Figure 2 , Figure 2 which is a schematic diagram of the modules in a management node provided by this application. Among them, the update module 111 and the software package module 112 are directly stored in the management node 11.
[0044] The artificial intelligence platform is a service platform that provides tools, resources, and environments to build, train, and deploy artificial intelligence models. The platform usually includes the following key functions and features: Data cleaning: helps remove errors and duplicate information in the data; Data transformation: converts the data into a format suitable for model training; Data augmentation: increases sample diversity by transforming the data and improves the generalization ability of the model.
[0045] Model Building: Provide various machine learning and deep learning frameworks, such as commonly used tools in the field of deep learning like TensorFlow, PyTorch (PythonTorch), Keras, etc.; Programming Environment: Provide an online code editor, Jupyter Notebook (an interactive computing platform that allows users to mix elements such as code, text, mathematical formulas, images, and videos in the same document), or other Integrated Development Environments (IDEs). Version Control: Support version management and collaboration of models; Model Training: Computational Resources: Provide powerful GPU or Tensor Processing Unit (TPU) computational resources to accelerate model training; Automatic Hyperparameter Tuning: Use automated tools (such as Hyperparameter tuning) to optimize model parameters; Training Monitoring: Monitor the training process in real-time, including metrics such as loss function and accuracy.
[0046] Evaluation Metrics: Provide various evaluation metrics, for example, accuracy, recall, F1 Score, etc.; Cross-Validation: Support different cross-validation methods to evaluate the generalization ability of the model; Model Deployment: Support deploying the model as an Application Programming Interface (API) service or integrating it into existing applications; Management and Monitoring: Provide management and monitoring tools during the runtime of the model to ensure stable operation of the model; Team Collaboration: Support multi-person collaboration and sharing of data and models; Community Sharing: Allow users to share models and experiences to promote knowledge exchange.
[0047] It should be noted that the fixed directory is a directory preset in the management node for storing files. The update module 111 copies the pre-obtained replacement files, driver data, and plugin files to the preset fixed directory. The replacement files are used to expand or upgrade the target processor. The setting of the fixed directory helps standardize the file storage location, facilitating management and access.
[0048] The mirror plugin is the Kubernetes device plugin mirror (deviceplugin, dp) of the target processor, which can be directly obtained from the manufacturer of the target processor. The mirror plugin is an extension of Kubernetes for the target processor, helping Kubernetes manage and schedule special hardware resources in the cluster, namely the target processor. The mirror plugin ensures that resources can be used by Pods (the most basic running unit in Kubernetes), and provides a mechanism for resource isolation and allocation. Among them, Kubernetes (k8s) is an open-source system used to manage containerized applications on multiple hosts in a cloud platform. The goal of Kubernetes is to make the deployment of containerized applications simple and efficient. Kubernetes provides a mechanism for application deployment, planning, updating, and maintenance.
[0049] Through an artificial intelligence platform of the present application, during the update and expansion process in the artificial intelligence platform, the update module copies various expansion files for the target processor, namely multiple replacement files, and the processor files related to the target processor, to a fixed directory. After that, various expansion files for the target processor exist in the fixed directory, and the plugin service corresponding to the target processor can be started to ensure that after expanding the target processor in the artificial intelligence platform, the functions corresponding to the target processor can run normally, completing the expansion of the target processor. Therefore, it can solve the technical problem that the existing expansion technology cannot expand nodes with GPUs of unsupported models and cannot be applied to scenarios where the on-site node environment changes, resulting in a poor user experience, and achieve the technical effect of quickly updating the expansion process for nodes with GPUs of unsupported models without factory compatibility, ensuring that nodes with GPUs of unsupported models can be expanded.
[0050] In an implementable embodiment of the present application, the update module 111 is further configured to:
[0051] Modify the tag of the plugin mirror in the processor file to obtain the modified plugin mirror, and push the modified plugin mirror to a preset mirror file database; wherein, the management node further includes a preset mirror file database, and the processor file at least includes a plugin mirror, driver data, and a plugin file.
[0052] In the embodiment of the present application, the plugin mirror includes but is not limited to metadata. Among them, a tag is a string used to identify the plugin mirror, including but not limited to information such as version number, build date, and environment identifier. The tag modification process refers to editing the metadata of the plugin mirror to update or change the identification information of the mirror.
[0053] The preset mirror file database is a database specifically used to store and manage plugin mirrors, such as: a mirror repository (harbor).
[0054] In an implementable embodiment of the present application, the update module 111 is further configured to:
[0055] Copy the plugin deployment serialization file, the plugin startup file, and the plugin stop file to the corresponding fixed directory of the plugin file according to a preset control instruction; wherein, the plugin file at least includes the plugin deployment serialization file, the plugin startup file, and the plugin stop file;
[0056] Copy the driver data to the corresponding fixed directory of the driver data according to a preset control instruction; wherein, the driver data is used to add the driver of the target processor in the preset server after the artificial intelligence platform is installed in the preset server.
[0057] Among them, the main function of the update module 111 is to execute the update of the entire expansion process. The update module 111 stores preset control instructions and various replacement files. The preset control instructions are some instructions written customarily, for example: command statements written according to the syntax of bash (Bourne Again Shell, whose syntax can be called bash script syntax or bash command syntax) and ansible (an automation tool used for configuration management and application deployment). Specifically, regarding the writing of the preset control instructions, the present application does not impose any restrictions.
[0058] The preset control instructions store multiple instructions, such as: the instruction to copy the plugin deployment serialization file to the corresponding fixed directory, the instruction to copy the driver data to the corresponding fixed directory, etc. Specifically, regarding the content in the preset control instructions, the present application does not impose any restrictions.
[0059] The plugin deployment serialization file is the configuration file of the plugin, including but not limited to: the parameters and setting information required by the plugin. Among them, the plugin deployment serialization file is saved in a serialized format. Serialization refers to the process of converting a data structure or object state into a storable or transmittable format. Regarding the plugin deployment serialization file, for example: the plugin deployment yaml file (a yaml file usually refers to a file written in the yaml (YAML Ain’t Markup Language) data serialization format and used for configuring and managing applications), etc. Specifically, regarding the plugin deployment serialization file, the present application does not impose any restrictions.
[0060] The plug-in startup file contains instructions or scripts for starting the plug-in, which are used to run the plug-in of the target processor when the artificial intelligence platform starts or under specific conditions. The plug-in stop file contains instructions or scripts for stopping the plug-in, which are used to stop the plug-in of the target processor when the artificial intelligence platform shuts down or under specific conditions. The driver data refers to the driver program files required for installing and configuring the target processor (such as GPU, Field-Programmable Gate Array (FPGA), etc.) in the preset server after the artificial intelligence platform is installed in the preset server.
[0061] Specifically, for the implementation process of the embodiments of the present application, it can be achieved by, but not limited to, the following method: The update module 111 executes the command statements in the preset control instruction: copy the plug-in deployment yaml file and the startup and stop files of this plug-in to the fixed directory where such files are stored on the management node of the platform.
[0062] Copy the key files of the plug-in deployment serialization file and the driver data to the specific and predefined directory location in the artificial intelligence platform, which can ensure that when the artificial intelligence platform starts and runs, it can automatically load and use the plug-in and driver of the target processor without manual intervention, helping to improve the stability and manageability of the artificial intelligence platform.
[0063] In an implementable embodiment of the present application, the update module 111 is further used for:
[0064] Copy the expansion and replacement file of the target processor to the corresponding fixed directory according to the preset control instruction; among them, multiple replacement files include the expansion and replacement file, and the expansion and replacement file is used to add the first expansion information of the target processor in the artificial intelligence platform;
[0065] Copy the module start / stop replacement file of all modules in the artificial intelligence platform to the corresponding fixed directory according to the preset control instruction; among them, multiple replacement files include the module start / stop replacement file, and the module start / stop replacement file is used to control the start and stop of all modules during the process of expanding the target processor in the artificial intelligence platform;
[0066] Copy the expansion template replacement file of the target processor to the corresponding fixed directory according to the preset control instruction; among them, multiple replacement files include the expansion template replacement file, and the expansion template replacement file is used to add the second expansion information of the target processor in the artificial intelligence platform, and the first expansion information is different from the second expansion information.
[0067] Among them, the preset control instructions at least further include instructions for copying the expansion replacement file to the corresponding fixed directory, instructions for copying the module start / stop replacement file to the corresponding fixed directory, instructions for copying the expansion template replacement file to the corresponding fixed directory, etc.
[0068] The deployment of the expansion replacement file for the target processor is a file used to add the first expansion information of the target processor in the artificial intelligence platform. The first expansion information includes, but is not limited to: the hardware configuration of the target processor, the resource allocation strategy, or parameter settings related to expansion. The preset control instructions will guide the copying of the expansion replacement file to the fixed directory corresponding to the expansion replacement file, and the fixed directory corresponding to the expansion replacement file is specifically used to store the expansion replacement file to ensure that the platform can access and use this file when needed.
[0069] The deployment of the module start / stop replacement files for all modules contains instructions or configurations for controlling the start and stop of all modules in the artificial intelligence platform. Since some modules need to be reconfigured or restarted in the artificial intelligence platform during the process of expanding the target processor to adapt to the new hardware resources, it is necessary to copy the module start / stop replacement files of all modules in the artificial intelligence platform to the corresponding fixed directory to ensure that the platform can access and use this file when needed.
[0070] The expansion template replacement file for the target processor contains the second expansion information used to add the target processor in the artificial intelligence platform. The second expansion information is different from the first expansion information and includes configuration parameters different from the first expansion information or settings for different expansion stages. Copying the expansion template replacement file to the corresponding fixed directory provides more detailed expansion control for the platform, allowing the platform to perform different degrees of hardware resource expansion according to different needs.
[0071] Specifically, for the implementation process of the embodiments of the present application, it can be achieved by, but not limited to, the following methods: The update module 111 executes the command statements in the preset control instructions: Copy the expansion replacement file of the target processor to the fixed directory where the platform's management node stores such files. Copy the replacement file for starting and stopping all modules managed by the platform to the fixed directory where the platform's management node stores such files. Copy the template replacement file for expansion to the fixed directory where the platform's management node stores such files.
[0072] Copy specific replacement files to the corresponding fixed directories in the artificial intelligence platform so that when the platform expands the target processor, new configurations and start / stop instructions can be automatically applied. By using the preset control instructions, the copying of files can be automated, reducing the complexity and error probability of manual operations. The expansion replacement file, the module start / stop replacement file, and the expansion template replacement file together ensure the stable operation of the platform and the effective management of resources during the expansion process.
[0073] In an implementable embodiment of the present application, the update module 111 is further configured to:
[0074] Copy the label addition and replacement file of the target processor to the fixed directory corresponding to the label addition and replacement file according to a preset control instruction; wherein, among the multiple replacement files, there is a label addition and replacement file, and the label addition and replacement file is used to add the label information of the target processor in the artificial intelligence platform;
[0075] Copy the node addition and replacement file of the target processor to the fixed directory corresponding to the node addition and replacement file according to a preset control instruction; wherein, among the multiple replacement files, there is a node addition and replacement file, and the node addition and replacement file is used to add the node information of the target processor in the artificial intelligence platform.
[0076] Wherein, the label addition and replacement file is a file containing the label information of the target processor, and the label information is usually used to identify and classify processor resources for resource management and scheduling in the artificial intelligence platform. The preset control instruction copies the label addition and replacement file to the corresponding fixed directory to ensure that the platform can allocate and manage resources according to the label information of the label addition and replacement file.
[0077] The node addition and replacement file is a file containing the node information of the target processor. The node information usually refers to the location and role information of the processor in the network or computing cluster, which is crucial for resource scheduling and task allocation of the platform. The preset control instruction copies the node addition and replacement file to the corresponding fixed directory, and the platform can optimize resource allocation according to the node information to ensure that tasks can be executed on the correct processor.
[0078] Furthermore, by adding label information through the label addition and replacement file, the platform can quickly identify and process specific processor resources through the label, which is of great significance for multi-dimensional management of resources. By adding node information through the node addition and replacement file, the platform can understand the specific location and role of the target processor in the network or cluster through the node information, which is crucial for task distribution and execution.
[0079] Specifically, regarding the implementation process of the embodiment of the present application, it can be achieved through but not limited to the following methods: The update module 111 executes the command statements in the preset control instruction: Copy the label addition and replacement file for expansion to the fixed directory where the management node of the platform stores such files. Copy the node addition and replacement file for expansion to the fixed directory where the management node of the platform stores such files.
[0080] By automatically deploying replacement files for tags and replacement files for node addition, the artificial intelligence platform can more efficiently manage and utilize the processor resources of the target processor, improving the overall computing efficiency and platform scalability.
[0081] In an implementable embodiment of the present application, the update module 111 is further configured to:
[0082] Copy the cleaning node replacement file of the target processor to the corresponding fixed directory according to a preset control instruction; wherein, among the multiple replacement files, there is a cleaning node replacement file, and the cleaning node replacement file is used to clean the data of the target processor in the preset server after the artificial intelligence platform is uninstalled from the preset server;
[0083] Copy the detection template replacement file of the target processor to the corresponding fixed directory according to a preset control instruction; wherein, among the multiple replacement files, there is a detection template replacement file, and the detection template replacement file is used to monitor the running state of the target processor through a first monitoring method after the artificial intelligence platform is installed on the preset server;
[0084] Copy the detection replacement file of the target processor to the corresponding fixed directory according to a preset control instruction; wherein, among the multiple replacement files, there is a detection replacement file, and the detection replacement file is used to monitor the running state of the target processor through a second monitoring method after the artificial intelligence platform is installed on the preset server.
[0085] Among them, the cleaning node replacement file of the target processor is a file used to clean the data related to the target processor in the preset server after the artificial intelligence platform is uninstalled from the preset server, including but not limited to: deleting configuration files, log files, temporary files, etc., to ensure that there is no residual data left in the preset server after the platform is uninstalled and keep the preset server environment clean. The preset control instruction copies the cleaning node replacement file to the corresponding fixed directory, and the data cleaning operation can be performed when needed.
[0086] The detection template replacement file of the target processor is a file used to monitor the running state of the target processor through a first monitoring method after the artificial intelligence platform is installed on the preset server. The first monitoring method includes but not limited to: a specified monitoring policy or template for collecting and processing the performance data, error logs, etc. of the target processor. The preset control instruction copies the detection template replacement file to the corresponding fixed directory, enabling the platform to immediately monitor the state of the target processor after installation and ensure its normal operation.
[0087] The detection and replacement file of the target processor is a file used to monitor the running status of the target processor through the second monitoring method after the artificial intelligence platform is installed on the preset server. The second monitoring method is different from the first monitoring method and provides another monitoring perspective or strategy to enhance the comprehensiveness and reliability of monitoring. The preset control instruction copies the detection and replacement file to the corresponding fixed directory, enabling the artificial intelligence platform to use multiple monitoring means to ensure the stability and performance of the target processor.
[0088] Furthermore, data cleaning can be performed through the cleaning node replacement file to ensure that the preset server environment is cleaned up after the artificial intelligence platform is uninstalled, avoiding data residue and security risks. Through the detection template replacement file and the monitoring of the running status of the detection and replacement file, the running status of the processor can be monitored in real time through different monitoring templates and strategies, so as to discover and solve problems in a timely manner and ensure the stable operation of the artificial intelligence platform.
[0089] Specifically, regarding the implementation process of the embodiments of the present application, it can be achieved through but not limited to the following methods: The update module 111 executes the command statements in the preset control instruction: Copy the cleaning node replacement file for capacity reduction to the fixed directory where such files are stored on the management node of the platform. Copy the detection template replacement file for monitoring the GPU to the fixed directory where such files are stored on the management node of the platform. Copy the detection and replacement file for monitoring the GPU to the fixed directory where such files are stored on the existing nodes of the platform.
[0090] Through the automated deployment of the cleaning node replacement file, the detection template replacement file, and the detection and replacement file, the artificial intelligence platform can manage and monitor resources more efficiently, improving the maintainability and reliability of the system.
[0091] In an implementable embodiment of the present application, the artificial intelligence platform further includes: a computing node 12,
[0092] The computing node 12 is used to deploy node services; among them, the node services at least include collecting node information, reporting cluster monitoring, and resource management services;
[0093] The computing node 12 is also used to run the development environment corresponding to artificial intelligence training to provide the infrastructure for artificial intelligence training.
[0094] Specifically, regarding the specific structures and functions of the management node 11 and the computing node 12 in the artificial intelligence platform, reference can be made to Figure 3 , Figure 3Schematic diagram of another artificial intelligence platform provided by this application. Among them, the computing node 12 is a key component in the artificial intelligence platform. Usually, it is an entity with computing capabilities, which can be a physical server, virtual machine, or container, etc. The computing node 12 is responsible for executing the computing tasks in the platform, especially during the artificial intelligence training process.
[0095] One of the main functions of the computing node 12 is to deploy node services. Node services include but are not limited to the following: Collect node information: The computing node 12 will collect its own configuration information, resource usage (such as CPU, memory, storage, and network usage), operating status, etc., so that the artificial intelligence platform can understand the specific situation of each node. Report cluster monitoring: The computing node 12 regularly reports the collected information to the cluster monitoring system, and the cluster manager can monitor the health status and performance of the entire cluster. Resource management service: The computing node 12 provides resource management services, including but not limited to resource allocation, scheduling, and optimization, ensuring that resources are effectively utilized and meet the needs of artificial intelligence training tasks.
[0096] The computing node 12 can also run the development environment required for artificial intelligence training. The development environment includes but is not limited to: operating system, programming language interpreter or compiler, artificial intelligence framework (such as TensorFlow, PyTorch, etc.), related libraries and dependencies, data processing and model training tools, etc.
[0097] The computing node 12 provides the necessary environment and resources for artificial intelligence training tasks, and model development, training, and testing can be carried out on the artificial intelligence platform. The computing node 12 in the artificial intelligence platform is not only responsible for deploying services to manage and monitor its own status, but also provides a complete development environment required for running artificial intelligence training tasks. Users of the artificial intelligence platform can focus on model design and optimization without worrying about the configuration and maintenance of the underlying infrastructure.
[0098] In an implementable embodiment of this application, the management node 11 at least further includes: cluster monitoring 114;
[0099] The management node 11 is also used to deploy resource management services and user management services.
[0100] Among them, the management node 11 is another key component in the artificial intelligence platform. It is responsible for managing and coordinating the work of the entire platform. The management node 11 usually does not directly participate in computing tasks, but focuses on providing support services to ensure the efficient and stable operation of the platform.
[0101] Please continue to refer to Figure 3, The components of the management node include but are not limited to the database 113. The management node 11 contains a database for storing various data of the platform, including but not limited to user information, training task records, resource allocation status, monitoring data, etc. The database 113 is a key component for the persistent storage of platform data.
[0102] The cluster monitoring 114 provides cluster monitoring services for real-time monitoring of the status of the entire cluster, including but not limited to the health status of computing nodes 12, resource usage, task execution progress, etc. This helps to promptly detect and solve problems and ensure the stable operation of the cluster.
[0103] Regarding the service deployment of the management node 11: The management node deploys a resource management service. The resource management service is responsible for resource allocation and scheduling of the entire artificial intelligence platform. Its specific functions include but are not limited to: receiving resource usage reports from computing nodes, dynamically allocating and recycling computing resources according to the requirements of training tasks, scheduling training tasks to appropriate computing nodes for execution, and optimizing resource usage to improve resource utilization.
[0104] The management node 11 also deploys a user management service, which is used to handle user-related operations, including but not limited to: user authentication and authorization; creation, update, and deletion of user accounts; management of user resource quotas; recording and analysis of user operation logs.
[0105] In an implementable embodiment of the present application, the management node 11 is pre-deployed in the artificial intelligence platform;
[0106] The computing nodes 12 are added to the artificial intelligence platform by way of expansion.
[0107] Among them, all computing nodes 12 are equipped with GPU devices, and the management node 11 may or may not be. In the face of unsupported GPU models, i.e., target processors, the adopted methods include but are not limited to: when starting to deploy the platform, only deploy the management node 11 (select a node without a GPU as the management node 11), and then add the computing nodes 12 to the cluster by way of expansion.
[0108] The management node 11 is a core component of the artificial intelligence platform and has been pre-deployed in the initialization stage of the artificial intelligence platform, which means that at the beginning of building the artificial intelligence platform, all services and functions of the management node 11 have been configured properly and are ready to start managing the entire platform. The pre-deployed management node 11 ensures that the training platform has a stable control center, which can be responsible for operations such as starting, configuring, monitoring, and maintaining the training platform.
[0109] The computing node 12 is not deployed all at once during the initialization of the artificial intelligence platform, but is dynamically added to the artificial intelligence platform through capacity expansion as needed. Capacity expansion refers to increasing more computing resources according to the changes in the processing requirements of the artificial intelligence platform to enhance the overall computing power of the platform.
[0110] In the artificial intelligence platform, the management node serves as the permanent control center of the platform, is pre-deployed and always running to ensure that the basic management and monitoring functions of the platform are always available. The computing nodes can be dynamically added to the platform according to actual needs, enabling the platform to flexibly handle different computing loads, thereby improving resource utilization and the scalability of the platform.
[0111] Embodiments of the present application provide a capacity expansion and update method. In combination with the execution process of the capacity expansion and update method, the method is described in detail.
[0112] As Figure 4 shown, Figure 4 is a schematic flowchart of a capacity expansion and update method provided by the present application. The capacity expansion and update method is applied to Figures 1 to 3 the described artificial intelligence platform, including:
[0113] Step 401, obtain the processor file of the target processor and multiple replacement files corresponding to the target processor; among them, the multiple replacement files are files pre-obtained for expanding the target processor.
[0114] In the embodiments of the present application, the processor file is a file related to the target processor, at least including a plugin image, driver data, and plugin files. The multiple replacement files are pre-obtained files specifically used for expanding the target processor.
[0115] Among them, the driver data refers to software components used to establish effective communication between the operating system and the target processor, usually including hardware drivers, enabling the operating system to recognize and correctly interact with the hardware components of the target processor (such as: Central Processing Unit (CPU), GPU, network interface, etc.). The purpose of obtaining the driver data is to ensure that the artificial intelligence platform can correctly install and configure the target processor on the preset server so that it can efficiently execute computing tasks.
[0116] The plugin image usually refers to a software package containing specific functions or services, which can expand the functions of the artificial intelligence platform and can be quickly deployed and used in the platform. The purpose of obtaining the plugin image is to integrate additional functions in the artificial intelligence platform, that is, the functions corresponding to the target processor.
[0117] The plugin file refers to the configuration file, script, or other auxiliary files related to the plugin of the target processor, which works together with the plugin image to ensure that the plugin can run correctly in the artificial intelligence platform. The plugin file may include, but is not limited to, startup scripts, stop scripts, configuration files, serialized data, etc., which are used to control the behavior of the plugin and configure its runtime environment.
[0118] Step 402: Copy the multiple replacement files corresponding to the target processor and the processor files to their respective fixed directories; the fixed directory is a pre-set directory for storing files.
[0119] In the embodiments of the present application, the multiple replacement files are pre-obtained files specifically used for expanding the target processor. Expansion includes, but is not limited to: increasing processing power, optimizing configuration, updating system parameters, etc. The types of the multiple replacement files include, but are not limited to: scripts, configuration files, or binary programs, etc.
[0120] Among them, the multiple replacement files include, but are not limited to: the expansion replacement file of the target processor, the module start / stop replacement file of all modules in the artificial intelligence platform, the expansion template replacement file of the target processor, the label addition replacement file of the target processor, the node addition replacement file of the target processor, the cleaning node replacement file of the target processor, the detection template replacement file of the target processor, and the detection replacement file of the target processor, etc. Specifically, the present application does not limit the multiple replacement files.
[0121] The fixed directory is pre-set and is a directory for storing specific types of files. Each type of file has its corresponding fixed directory, which helps the management and maintenance of the system.
[0122] Step 403: When there are corresponding multiple replacement files and processor files in the multiple fixed directories, start the plugin service of the target processor by executing the plugin startup file; among them, the processor file at least includes the plugin startup file.
[0123] In the embodiments of the present application, in the artificial intelligence platform, it is necessary to copy the drive data and plugin files in the multiple replacement files and processor files to their respective fixed directories, and all necessary files need to be placed in the corresponding positions to prepare for starting and processing the plugin service of the target processor.
[0124] The plugin startup file is a part of the plugin file in the processor file and contains instructions and configuration information for starting the plugin service. Regarding the type of the plugin startup file, for example: script or executable program, etc., which is used to initialize the plugin environment, load necessary resources, and start the plugin to run.
[0125] By executing the plug-in startup file, the artificial intelligence platform will start the plug-in service of the target processor. The process of executing the plug-in startup file includes but is not limited to the following steps: Load the plug-in's dependent libraries and resources. Configure the environment variables and parameters required for the plug-in to run. Initialize the state and objects inside the plug-in. Start the main process or service of the plug-in to start executing the predetermined function.
[0126] After ensuring that all necessary replacement files, driver data in the processor files, and plug-in files are located in the correct fixed directory, the AI platform starts the plug-in service of the target processor by executing the plug-in startup file. This enables the AI platform to utilize the additional functions and services provided by the plug-in to enhance the performance and functionality of the target processor. The AI platform can dynamically expand its capabilities according to the needs of the training task, thereby providing a more flexible and efficient AI training environment.
[0127] Step 404, based on the activated plug-in service, expand the target processor to the artificial intelligence platform.
[0128] In an embodiment of the present application, after the plug-in service is successfully started, the artificial intelligence platform can provide additional functions and support, at which time the target processor can be expanded in the artificial intelligence platform.
[0129] Regarding the process of expanding the target processor to the artificial intelligence platform, it can be done through but not limited to the following methods: Identification and registration: The artificial intelligence platform identifies the newly added target processor and registers it with the platform's resource management system. Resource allocation: The platform allocates resources to the newly added processor (target processor), such as: CPU core, memory, storage space, and network bandwidth. Service synchronization: Ensure that the plug-in service of the target processor is synchronized with other parts of the platform, including: monitoring, logging, and task scheduling. Status monitoring: Start status monitoring of the target processor to ensure that it runs stably and can respond to the platform's scheduling. Specifically, regarding the process of expanding the target processor to the artificial intelligence platform, it can be determined based on actual conditions, and this application does not impose any restrictions.
[0130] By obtaining the necessary files and data and ultimately expanding the target processor to the AI platform, the stability and scalability of the platform can be ensured, allowing the AI platform to flexibly adjust resources to adapt to changing workloads and performance requirements.
[0131] In an achievable embodiment of the present application, when copying multiple replacement files, driver data and plug-in files corresponding to the target processor to a fixed directory, it is necessary to perform it through pre-set control instructions. Regarding the setting of control instructions, it can be implemented by but not limited to the following methods: obtain the control instructions generated by the preset command statements to obtain the preset control instructions.
[0132] In an embodiment of the present application, the preset command statement is a custom - selected command statement. For example: bash statement, ansible statement, etc. Specifically, the present application does not limit the preset command statement.
[0133] The preset control instruction stores multiple instructions, including but not limited to: control instructions corresponding to multiple replacement files, driver data, and plugin files of the target processor, etc., to copy files to a fixed directory.
[0134] Obtaining the control instruction generated by the preset command statement ensures that the artificial intelligence platform can automatically execute tasks according to a predefined operation process, reduces manual intervention, improves the accuracy and efficiency of operations. By using the preset command statement to generate control instructions, the artificial intelligence platform can maintain a high degree of flexibility and maintainability, thus better adapting to different operating environments and user requirements.
[0135] In an implementable embodiment of the present application, when copying multiple replacement files, driver data, and plugin files corresponding to the target processor to a fixed directory, the following methods can be adopted but are not limited to: performing tag modification processing on the plugin image to obtain a modified plugin image, and pushing the modified plugin image to a preset image file database; where the preset image file database is a database pre - set for storing plugin images; by executing the preset control instruction, copying multiple replacement files, driver data, and plugin files corresponding to the target processor to their respective fixed directories.
[0136] In an embodiment of the present application, the plugin image includes but is not limited to metadata. Among them, a tag is a string used to identify the plugin image, including but not limited to information such as version number, build date, environment identifier, etc. The tag modification processing refers to editing the metadata of the plugin image to update or change the identification information of the image.
[0137] The preset image file database is a database specifically used for storing and managing plugin images. The preset image file database can be private or public, depending on the security requirements and deployment strategies of the artificial intelligence platform. Pushing the modified plugin image to the preset image file database is to save the modified plugin image so that it can be used by other components or services in the artificial intelligence platform.
[0138] Regarding the copying of files, it can be achieved by but not limited to the following method: executing command statement 1: obtaining the plugin image of the target processor from the software package module, loading the image, modifying the image tag, and finally pushing the image to the harbor repository (preset image file database).
[0139] Execute command statement 2: Copy the plugin deployment YAML file (plugin deployment serialization file) and the plugin's startup and stop files (plugin startup file and plugin stop file) to the fixed directory on the management node of the artificial intelligence platform where such files are stored.
[0140] Execute command statement 3: Copy the expansion replacement file for unsupported GPUs (expansion replacement file for the target processor) to the fixed directory on the management node of the artificial intelligence platform where such files are stored.
[0141] Execute command statement 4: Copy the replacement file for starting and stopping all modules of the platform management (module start / stop replacement file) to the fixed directory on the management node of the artificial intelligence platform where such files are stored.
[0142] Execute command statement 5: Copy the template replacement file for expansion (expansion template replacement file for the target processor) to the fixed directory on the management node of the artificial intelligence platform where such files are stored.
[0143] Execute command statement 6: Copy the add label replacement file for expansion (label addition replacement file for the target processor) to the fixed directory on the management node of the artificial intelligence platform where such files are stored.
[0144] Execute command statement 7: Copy the add node replacement file for expansion (node addition replacement file for the target processor) to the fixed directory on the management node of the artificial intelligence platform where such files are stored.
[0145] Execute command statement 8: Copy the clean node replacement file for scaling down (clean node replacement file for the target processor) to the fixed directory on the management node of the artificial intelligence platform where such files are stored.
[0146] Execute command statement 9: Copy the GPU detection template replacement file for monitoring (detection template replacement file for the target processor) to the fixed directory on the management node of the artificial intelligence platform where such files are stored.
[0147] Execute command statement 10: Copy the GPU detection replacement file for monitoring (detection replacement file for the target processor) to the fixed directory on the existing nodes of the artificial intelligence platform where such files are stored.
[0148] Execute command statement 11: Copy the driver provided by the GPU manufacturer stored in the software package module (driver data) to the fixed directory on the management node of the artificial intelligence platform where such files are stored.
[0149] Execute command statement 12: Execute the plugin's startup file on the management node to start the plugin.
[0150] Among them, command statements 1-12 are the instructions included in the preset control instructions.
[0151] Executing the preset control instructions to copy the replacement files, driver data, and plugin files of the target processor to their respective corresponding fixed directories ensures that all necessary files are correctly placed in the positions expected by the system, facilitating subsequent service startup and platform expansion operations.
[0152] In summary, the present application can achieve the following technical effects:
[0153] In the process of updating and expanding the artificial intelligence platform of the present application, various expansion files of the target processor, namely multiple replacement files, driver data, and plugin files, are copied to the fixed directory through the update module. After that, various expansion files of the target processor exist in the fixed directory, and the plugin service corresponding to the target processor can be started to ensure that after expanding the target processor in the artificial intelligence platform, the functions corresponding to the target processor can run normally, completing the expansion of the target processor. Therefore, it can solve the technical problem that the existing expansion technology cannot expand the nodes of GPUs of unsupported models and is not applicable to the scenario of on-site node environment changes, resulting in a poor user experience, and achieve the technical effect of quickly updating the expansion process of the nodes of GPUs of unsupported models without factory compatibility and ensuring that the nodes of GPUs of unsupported models can be expanded.
[0154] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method.
[0155] The embodiments of the present application also provide an expansion and update device. Figure 5 As shown in the structural schematic diagram of an expansion and update device provided by the present application, as Figure 5 shown, it includes:
[0156] An acquisition unit 51, configured to acquire the processor file of the target processor and multiple replacement files corresponding to the target processor; among them, the multiple replacement files are files acquired in advance for expanding the target processor.
[0157] A copying unit 52, configured to copy the multiple replacement files corresponding to the target processor and the processor file to their respective corresponding fixed directories; among them, the fixed directory is a directory preset for storing files.
[0158] A startup unit 53, which is configured to start the plugin service of a target processor by executing a plugin startup file when there are multiple corresponding replacement files and processor files in multiple fixed directories; wherein, the processor file includes at least the plugin startup file.
[0159] An expansion unit 54, which is configured to expand the target processor to an artificial intelligence platform based on the started plugin service.
[0160] In an embodiment of the present application, as Figure 6 shown, the expansion and update device further includes:
[0161] A generation unit 55, which is configured to obtain a control instruction generated by a preset command statement to obtain a preset control instruction.
[0162] In an embodiment of the present application, the replication unit 52 is further configured to:
[0163] Modify the tag of the plugin image to obtain a modified plugin image, and push the modified plugin image to a preset image file database; wherein, the preset image file database is a database preset for storing plugin images.
[0164] By executing the preset control instruction, copy the multiple replacement files, driver data, and plugin files corresponding to the target processor to their respective corresponding fixed directories for the multiple replacement files, driver data, and plugin files.
[0165] For the description of the features in the embodiments corresponding to the expansion and update device, reference can be made to the relevant descriptions in the embodiments corresponding to the expansion and update method, which will not be elaborated here one by one.
[0166] An embodiment of the present application further provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above embodiments of the expansion and update method.
[0167] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps in any one of the above embodiments of the expansion and update method when running.
[0168] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: USB flash drives, read-only memories (ROM for short), random access memories (RAM for short), mobile hard disks, magnetic disks, or optical discs and other various media that can store computer programs.
[0169] Embodiments of the present application also provide a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above-described embodiments of the expansion and update method.
[0170] Embodiments of the present application also provide another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above-described embodiments of the expansion and update method.
[0171] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0172] The above has introduced in detail an artificial intelligence platform and an expansion and update method, device, and electronic device provided by the present application. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application. It should be noted that for those of ordinary skill in the art in the technical field, without departing from the principle of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.
Claims
1. An artificial intelligence platform, characterized in that: include: A management node, the management node including an update module, a software package module and a database; The update module is used to obtain a processor file of a target processor from the software package module, and to obtain a plurality of replacement files corresponding to the target processor from the database; The update module is also used to copy the multiple replacement files and the processor file to the fixed directories corresponding to the multiple replacement files and the processor files respectively; wherein the management node also includes a plurality of the fixed directories, the processor file includes at least a plug-in image, driver data and a plug-in file, and the plug-in image is a k8s device plug-in image of the target processor; The update module is also used to start the plug-in service of the target processor by executing the plug-in startup file when the corresponding multiple replacement files and the processor files exist in the multiple fixed directories; wherein the processor file includes the plug-in startup file, and the plug-in startup file at least contains an instruction or script to start the plug-in, which is used to run the plug-in of the target processor when the artificial intelligence platform is started or under specific conditions; The management node is used to expand the target processor to the artificial intelligence platform based on the plug-in service after startup, wherein the artificial intelligence platform is used to implement artificial intelligence training and reasoning tasks of the computer cluster.
2. The artificial intelligence platform according to claim 1, characterized in that: The update module is also used for: The plug-in image in the processor file is subjected to label modification processing to obtain a modified plug-in image, and the modified plug-in image is pushed to a preset image file database; wherein the management node also includes the preset image file database.
3. The artificial intelligence platform according to claim 2, characterized in that: The update module is also used for: Copy the plug-in deployment serialization file, the plug-in startup file and the plug-in stop file to the fixed directory corresponding to the plug-in file according to the preset control instruction; wherein the plug-in file at least includes the plug-in deployment serialization file, the plug-in startup file and the plug-in stop file; The driving data is copied to the fixed directory corresponding to the driving data according to the preset control instruction; wherein the driving data is used to add the driver of the target processor in the preset server after the artificial intelligence platform is installed in the preset server.
4. The artificial intelligence platform according to claim 3, characterized in that: The update module is also used for: According to the preset control instruction, copy the expansion replacement file of the target processor to the fixed directory corresponding to the expansion replacement file; wherein the multiple replacement files include the expansion replacement file, and the expansion replacement file is used to add the first expansion information of the target processor in the artificial intelligence platform; According to the preset control instruction, the module start-stop replacement files of all modules in the artificial intelligence platform are copied to the fixed directory corresponding to the module start-stop replacement files; wherein the multiple replacement files include the module start-stop replacement file, and the module start-stop replacement file is used to control the start and stop of all modules in the process of expanding the target processor of the artificial intelligence platform; According to the preset control instruction, the expansion template replacement file of the target processor is copied to the fixed directory corresponding to the expansion template replacement file; wherein the multiple replacement files include the expansion template replacement file, and the expansion template replacement file is used to add the second expansion information of the target processor in the artificial intelligence platform, and the first expansion information is different from the second expansion information.
5. The artificial intelligence platform according to claim 4, characterized in that: The update module is also used for: According to the preset control instruction, copy the label addition replacement file of the target processor to the fixed directory corresponding to the label addition replacement file; wherein the multiple replacement files include the label addition replacement file, and the label addition replacement file is used to add the label information of the target processor in the artificial intelligence platform; According to the preset control instruction, the node addition and replacement file of the target processor is copied to the fixed directory corresponding to the node addition and replacement file; wherein the multiple replacement files include the node addition and replacement file, and the node addition and replacement file is used to add the node information of the target processor in the artificial intelligence platform.
6. The artificial intelligence platform according to claim 5, characterized in that: The update module is also used for: According to the preset control instruction, the cleanup node replacement file of the target processor is copied to the fixed directory corresponding to the cleanup node replacement file; wherein the multiple replacement files include the cleanup node replacement file, and the cleanup node replacement file is used to clean up the data of the target processor in the preset server after the artificial intelligence platform is uninstalled from the preset server; Copying the detection template replacement file of the target processor to the fixed directory corresponding to the detection template replacement file according to the preset control instruction; wherein the multiple replacement files include the detection template replacement file, and the detection template replacement file is used to monitor the running state of the target processor through a first monitoring method after the artificial intelligence platform is installed on the preset server; According to the preset control instruction, the detection replacement file of the target processor is copied to the fixed directory corresponding to the detection replacement file; wherein the multiple replacement files include the detection replacement file, and the detection replacement file is used to monitor the operating status of the target processor through a second monitoring method after the artificial intelligence platform is installed on the preset server.
7. The artificial intelligence platform according to claim 1, characterized in that: The artificial intelligence platform also includes: computing nodes, The computing node is used to deploy node services; wherein the node services at least include collecting node information, reporting cluster monitoring, and resource management services; The computing node is also used to run the development environment corresponding to the artificial intelligence training to provide an infrastructure for artificial intelligence training.
8. The artificial intelligence platform according to claim 1, characterized in that: The management node also includes at least: cluster monitoring; The management node is also used to deploy resource management services and user management services.
9. The artificial intelligence platform according to claim 7, characterized in that: The management node is pre-deployed in the artificial intelligence platform; The computing nodes are added to the artificial intelligence platform by means of capacity expansion.
10. A capacity expansion and updating method, characterized in that: The capacity expansion and updating method is applied to the artificial intelligence platform according to any one of claims 1 to 9, comprising: Acquire a processor file of a target processor and a plurality of replacement files corresponding to the target processor; wherein the plurality of replacement files are pre-acquired files for expanding the capacity of the target processor; Copy the multiple replacement files and the processor file corresponding to the target processor to the fixed directories corresponding to the multiple replacement files and the processor files respectively; wherein the fixed directory is a pre-set directory for storing files, the processor file includes at least a plug-in image, driver data and a plug-in file, and the plug-in image is a k8s device plug-in image of the target processor; In the case where the plurality of replacement files and the processor files corresponding to each other exist in the plurality of fixed directories, the plug-in service of the target processor is started by executing the plug-in startup file; wherein the processor file at least includes the plug-in startup file, and the plug-in startup file at least includes an instruction or script for starting the plug-in, which is used to run the plug-in of the target processor when the artificial intelligence platform is started or under specific conditions; Based on the started plug-in service, the target processor is expanded to the artificial intelligence platform, wherein the artificial intelligence platform is used to implement artificial intelligence training and reasoning tasks of a computer cluster.
11. The capacity expansion and updating method according to claim 10, characterized in that: Before obtaining the driver data, plug-in image and plug-in file of the target processor, the method further includes: Acquire the control instruction generated by the preset command statement to obtain the preset control instruction.
12. The capacity expansion and updating method according to claim 11, characterized in that: The step of copying the plurality of replacement files corresponding to the target processor and the processor file to fixed directories corresponding to the plurality of replacement files and the processor file respectively comprises: Performing tag modification processing on the plug-in image in the processor file to obtain a modified plug-in image, and pushing the modified plug-in image to a preset image file database; wherein the preset image file database is a pre-set database for storing the plug-in image; By executing the preset control instruction, the multiple replacement files and the processor file corresponding to the target processor are respectively copied to the fixed directories corresponding to the multiple replacement files and the processor file.
13. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the capacity expansion and updating method as claimed in any one of claims 10 to 12 when executing the computer program.
14. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the capacity expansion and update method according to any one of claims 10 to 12.
15. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the capacity expansion and updating method as claimed in any one of claims 10 to 12 are implemented.
Citation Information
Patent Citations
Fault domain capacity expansion method and device, computer equipment and storage medium
CN117631996A