Artificial intelligence model deployment method, computer system, computer readable storage medium and computer program product

By flexibly allocating storage locations based on parameter quantity and storage resource information, the problem of low resource utilization caused by parameter storage of artificial intelligence models of different scales is solved, and more efficient model training and inference speed is achieved.

WO2025148772A1PCT designated stage expired Publication Date: 2025-07-17HUAWEI TECH CO LTD

Patent Information

Application Number
PCT/CN2025/070156
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-12
Filing Date
2025-01-02
Publication Date
2025-07-17

AI Technical Summary

Technical Problem

In the prior art, the parameters of storing artificial intelligence models of different scales based on a storage solution lead to low resource utilization and the system performance cannot be fully utilized.

Method used

According to the parameter quantity of the artificial intelligence model and the storage resource information of the computing system, the storage location of the parameters is flexibly allocated, and the parameters are stored in storage media that meet the needs, including storage media associated with a dedicated processor, storage media associated with a general processor and remote storage devices, making full use of the resources of the computing system.

Benefits of technology

It improves resource utilization, ensures that special processors can quickly obtain parameters, improves the training or inference speed of artificial intelligence models, and gives full play to the system performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025070156_17072025_PF_FP_ABST
    Figure CN2025070156_17072025_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed are an artificial intelligence model deployment method, a computer system, a computer readable storage medium and a computer program product, relating to the field of artificial intelligence. The method comprises: acquiring parameter quantity of parameters used during inference or training of an artificial intelligence model, and information of available storage resources of a computing system deployed by the artificial intelligence model; on the basis of the parameter quantity and the information of the available storage resources of the computing system, determining storage resources used for parameter storage; and storing parameters to the determined storage resources when the artificial intelligence model is deployed. The information of the storage resources comprises multiple types of storage resources, and each type of storage resources among the multiple types of storage resources has different access performance. Therefore, the storage position of parameters is flexibly allocated, the resources of the computing system are fully utilized, the resource utilization rate is effectively improved, parameters are obtained from the determined storage resources as soon as possible, computing is accelerated, and the system performance is fully exerted.
Need to check novelty before this filing date? Find Prior Art

Description

Artificial intelligence model deployment method, computer system, computer-readable storage medium, and computer program product

[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office on January 12, 2024, with application number 202410054568.5 and application name “Artificial Intelligence Model Deployment Method, Computer System, Computer-Readable Storage Medium and Computer Program Product”, all contents of which are incorporated by reference into this application. Technical Field

[0002] The present application relates to the field of artificial intelligence, and in particular to an artificial intelligence model deployment method, a computer system, a computer-readable storage medium, and a computer program product. Background Art

[0003] At present, general-purpose processors and dedicated processors perform the calculations for training and reasoning of artificial intelligence models to improve computing efficiency. For artificial intelligence models such as recommendation systems and natural language processing, only some parameters in the artificial intelligence model (such as sparse parameters) are required to participate in the calculation in one calculation. For example, the parameters are vectorized (embedding), and dedicated processors process the vector data. Generally, since the vector data of artificial intelligence models of different scales have different data volumes, storing the vector data of artificial intelligence models of different scales based on one storage solution results in poor storage flexibility, low resource utilization, and inability to fully utilize system performance. Summary of the Invention

[0004] The present application provides an artificial intelligence model deployment method, a computer system, a computer-readable storage medium, and a computer program product, thereby effectively improving resource utilization and system performance.

[0005] In a first aspect, a method for deploying an artificial intelligence model is provided, the method comprising obtaining parameter quantities of parameters used when the artificial intelligence model performs inference or training, and information about available storage resources of a computing system in which the artificial intelligence model is deployed; determining storage resources for storing the parameters based on the parameter quantities and the information about the available storage resources of the computing system; and, when deploying the artificial intelligence model, storing the parameters in the determined storage resources. The storage resource information includes multiple types of storage resources, each type of storage resource having different access performance.

[0006] Compared with storing the parameters of artificial intelligence models of different scales based on a single storage solution, for example, storing the parameters of a smaller-scale artificial intelligence model in a remote storage device, or storing the parameters of an ultra-large-scale artificial intelligence model in a dedicated processor, which results in low resource utilization and inability to fully utilize system performance. The artificial intelligence model deployment method provided in this application is that when deploying an artificial intelligence model, for artificial intelligence models of different scales, the parameter quantities used by the artificial intelligence model during inference or training are different, and the parameters are stored in at least one storage medium that meets the parameter quantity based on the information of the available storage resources of the computing system where the artificial intelligence model is deployed. Thus, the storage location of the parameters is flexibly allocated based on multi-dimensional information such as the parameter quantity and storage resource information, fully utilizing the resources of the computing system, effectively improving resource utilization, and enabling the dedicated processor to obtain parameters from the determined storage resources as quickly as possible when training or inferring the artificial intelligence model, and accelerate the execution of model training calculations or model inference calculations based on the parameters, fully utilizing system performance. Avoid using only one storage medium to store the parameters of the artificial intelligence model, limiting the storage location of the parameters, resulting in the dedicated processor being unable to obtain the parameters in a timely manner, and reducing the training or inference speed of the artificial intelligence model.

[0007] The parameters described herein may be data obtained by embedding sparse parameters. For example, sparse parameters may be converted into dense data. Consequently, since the number of parameters is small, storage and computing resources are reduced, computing and storage resource utilization is improved, and computational speed is increased.

[0008] In a possible implementation, the multiple types of storage resources include: a first storage medium associated with a dedicated processor, a second storage medium associated with a general-purpose processor, and a remote storage device.

[0009] In another possible implementation, the information on the available storage resources of the computing system includes the remaining storage capacity; and the storage resources for storing the parameters are determined based on the parameter quantity and the information on the available storage resources of the computing system, including: when the remaining storage capacity of the first storage medium is greater than the parameter quantity of the parameter, determining that the determined storage resources include the first storage medium.

[0010] In another possible implementation, storing the parameter in the determined storage resource includes: when a remaining storage capacity of the first storage medium is greater than a parameter amount of the parameter, storing the parameter in the first storage medium.

[0011] When the remaining storage capacity of the first storage medium associated with the dedicated processor meets the parameter quantity of the parameters, that is, the first storage medium has sufficient storage space to store the parameters related to the training calculation or inference calculation of the artificial intelligence model, the parameters are stored in the first storage medium, so that when the dedicated processor trains or infers the artificial intelligence model, it can obtain the parameters from the determined storage resources as quickly as possible, and accelerate the execution of model training calculation or model inference calculation based on the parameters, so as to give full play to the system performance.

[0012] In another possible implementation, the storage resources for storing the parameters are determined based on the parameter quantity and information about the available storage resources of the computing system, including: when the remaining storage capacity of the first storage medium is less than the parameter quantity of the parameter, determining that the determined storage resources include the second storage medium.

[0013] In another possible implementation, storing the parameter in the determined storage resource includes: when the remaining storage capacity of the first storage medium is less than the parameter amount of the parameter, storing the parameter in a second storage medium.

[0014] If the remaining storage capacity of the first storage medium associated with the dedicated processor does not meet the parameter quantity, that is, the first storage medium does not have sufficient storage space to store the parameters related to the training calculation or inference calculation of the artificial intelligence model, the parameters will be stored in the second storage medium associated with the general-purpose processor. This allows the storage location of the parameters to be flexibly allocated based on storage requirements and storage resource characteristics, fully utilizing the computing system resources and effectively improving resource utilization.

[0015] In another possible implementation, determining that the determined storage resource includes the second storage medium includes: when a remaining storage capacity of the second storage medium is greater than a parameter amount of the parameter, determining that the determined storage resource includes the second storage medium.

[0016] In another possible implementation, storing the parameter in the determined storage resource includes: when a remaining storage capacity of the second storage medium is greater than a parameter amount of the parameter, storing the parameter in the second storage medium.

[0017] When the remaining storage capacity of the second storage medium associated with the general-purpose processor meets the parameter quantity of the parameter, that is, the second storage medium has sufficient storage space to store parameters related to the training calculation or inference calculation of the artificial intelligence model, the parameters are stored in the second storage medium.

[0018] In another possible implementation, determining that the determined storage resources include the second storage medium includes: when the remaining storage capacity of the second storage medium is less than a parameter amount of the parameter, determining that the determined storage resources include the second storage medium and the remote storage device.

[0019] In another possible implementation, storing the parameters to the determined storage resources includes: when the remaining storage capacity of the second storage medium is less than the parameter amount of the parameters, storing a first number of the parameters to the second storage medium, and storing a second number of the parameters to a remote storage device.

[0020] If the remaining storage capacity of the second storage medium associated with the general-purpose processor does not meet the parameter quantity, that is, the second storage medium does not have sufficient storage space to store parameters related to the training calculation or inference calculation of the artificial intelligence model, the parameters are divided into two parts, with a first number of parameters stored in the second storage medium and a second number of parameters stored in a remote storage device. In this way, the storage location of the parameters can be flexibly allocated based on storage requirements and storage resource characteristics, fully utilizing the resources of the computing system and effectively improving resource utilization.

[0021] In another possible implementation, the storage resources for storing the parameters are determined based on the parameter quantity and information about the available storage resources of the computing system, including: when the remaining storage capacity of the first storage medium is less than the parameter quantity of the parameter, and the remaining storage capacity of the second storage medium is less than the parameter quantity of the parameter, determining that the determined storage resources include a remote storage device.

[0022] In another possible implementation, storing the parameter to the determined storage resource includes: when the remaining storage capacity of the first storage medium is less than the parameter amount of the parameter, and the remaining storage capacity of the second storage medium is less than the parameter amount of the parameter, storing the parameter to a remote storage device.

[0023] If the remaining storage capacity of the first storage medium associated with the dedicated processor does not meet the parameter quantity, and the remaining storage capacity of the second storage medium associated with the general-purpose processor does not meet the parameter quantity, the parameters are stored in a remote storage device. This allows flexible allocation of parameter storage locations based on storage requirements and storage resource characteristics, fully utilizing computing system resources and effectively improving resource utilization.

[0024] In another possible implementation, the parameters include multiple feature tables related to the artificial intelligence model; the storage resources for storing the parameters are determined based on the parameter quantity and information about the available storage resources of the computing system, including: when the remaining storage capacity of the first storage medium is less than the parameter quantity of the parameter, determining that the determined storage resources include the first storage medium and the second storage medium; storing the parameters in the determined storage resources, including: when the remaining storage capacity of the first storage medium is less than the parameter quantity of the parameter, storing the first number of feature tables in the first storage medium, and storing the second number of feature tables in the second storage medium.

[0025] In another possible implementation, the parameters include multiple feature tables related to the artificial intelligence model; determining the storage resources for storing the parameters based on the parameter quantity and information about the available storage resources of the computing system, including: when the remaining storage capacity of the first storage medium is less than the parameter quantity of the parameter, determining that the determined storage resources include the first storage medium and the remote storage device; storing the parameters to the determined storage resources, including: when the remaining storage capacity of the first storage medium is less than the parameter quantity of the parameter, storing a first number of feature tables to the first storage medium, and storing a second number of feature tables to the remote storage device.

[0026] Therefore, multiple feature tables are divided into two parts, one part is stored in a first storage medium associated with a dedicated processor, and the other part is stored in a second storage medium or a remote storage medium associated with a general-purpose processor. That is, the first storage medium is used as a cache layer to cache some feature tables in multiple feature tables, so that the dedicated processor can obtain parameters as quickly as possible, accelerate the execution of model training calculations or model inference calculations based on the parameters, improve the training or inference speed of the artificial intelligence model, and give full play to the system performance.

[0027] In another possible implementation, the method also includes: obtaining the user's performance requirements for artificial intelligence model reasoning or training; determining the storage resources for storing parameters based on the parameter quantity and information about the available storage resources of the computing system includes: determining the storage resources for storing parameters based on the parameter quantity, performance requirements and information about the available storage resources of the computing system.

[0028] In another possible implementation, parameters are obtained from the determined storage resource storage, and a dedicated processor performs training calculations or inference calculations on the artificial intelligence model based on the obtained parameters.

[0029] In a second aspect, an artificial intelligence model deployment apparatus is provided. The artificial intelligence model deployment apparatus includes modules for executing the artificial intelligence model deployment method of the first aspect or any possible design of the first aspect. For example, the artificial intelligence model deployment apparatus is configured to implement the functions of a general-purpose processor and includes a communication module and a processing module.

[0030] A communication module is used to obtain the parameter quantities of the parameters used by the artificial intelligence model to be deployed for reasoning or training, and information about the available storage resources of the computing system where the artificial intelligence model is deployed, wherein the information about the storage resources includes multiple types of storage resources, and each type of storage resource has different access performance.

[0031] A processing module is used to determine the storage resources for storing the parameters based on the parameter quantity and information about the available storage resources of the computing system; and, when deploying the artificial intelligence model, store the parameters in the determined storage resources.

[0032] In a possible implementation, the multiple types of storage resources include: a first storage medium associated with a dedicated processor, a second storage medium associated with a general-purpose processor, and a remote storage device.

[0033] In another possible implementation, the information on the available storage resources of the computing system includes the remaining storage capacity; when the processing module determines the storage resources for storing the parameters based on the parameter quantity and the information on the available storage resources of the computing system, it is specifically used to: when the remaining storage capacity of the first storage medium is greater than the parameter quantity of the parameter, determine that the determined storage resources include the first storage medium.

[0034] In another possible implementation, when the processing module stores the parameter in the determined storage resource, it is specifically configured to: store the parameter in the first storage medium when the remaining storage capacity of the first storage medium is greater than the parameter amount of the parameter.

[0035] In another possible implementation, when the processing module determines the storage resources for storing parameters based on the parameter quantity and information about the available storage resources of the computing system, it is specifically used to: when the remaining storage capacity of the first storage medium is less than the parameter quantity of the parameter, determine that the determined storage resources include the second storage medium.

[0036] In another possible implementation, when the processing module stores the parameter in the determined storage resource, it is specifically configured to: when the remaining storage capacity of the first storage medium is less than the parameter amount of the parameter, store the parameter in the second storage medium.

[0037] In another possible implementation, when the processing module determines that the determined storage resources include the second storage medium, it is specifically configured to: when the remaining storage capacity of the second storage medium is greater than the parameter amount of the parameter, determine that the determined storage resources include the second storage medium.

[0038] In another possible implementation, when the processing module stores the parameter in the determined storage resource, it is specifically configured to: store the parameter in the second storage medium when the remaining storage capacity of the second storage medium is greater than the parameter amount of the parameter.

[0039] In another possible implementation, when the processing module determines that the determined storage resources include a second storage medium, it is specifically used to: when the remaining storage capacity of the second storage medium is less than the parameter amount of the parameter, determine that the determined storage resources include the second storage medium and a remote storage device.

[0040] In another possible implementation, when the processing module stores the parameters to the determined storage resources, it is specifically used to: when the remaining storage capacity of the second storage medium is less than the parameter amount of the parameters, store the first number of parameters in the parameters to the second storage medium, and store the second number of parameters in the parameters to the remote storage device.

[0041] In another possible implementation, when the processing module determines the storage resources for storing parameters based on the parameter quantity and information about the available storage resources of the computing system, it is specifically used to: when the remaining storage capacity of the first storage medium is less than the parameter quantity of the parameter, and the remaining storage capacity of the second storage medium is less than the parameter quantity of the parameter, determine that the determined storage resources include a remote storage device.

[0042] In another possible implementation, when the processing module stores the parameters to the determined storage resource, it is specifically used to: when the remaining storage capacity of the first storage medium is less than the parameter amount of the parameter, and the remaining storage capacity of the second storage medium is less than the parameter amount of the parameter, store the parameters to the remote storage device.

[0043] In another possible implementation, the parameters include multiple feature tables related to the artificial intelligence model; when the processing module determines the storage resources for storing the parameters based on the parameter quantity and information about the available storage resources of the computing system, it is specifically used to: when the remaining storage capacity of the first storage medium is less than the parameter quantity of the parameter, determine that the determined storage resources include the first storage medium and the second storage medium; when the processing module stores the parameters to the determined storage resources, it is specifically used to: when the remaining storage capacity of the first storage medium is less than the parameter quantity of the parameter, store the first number of feature tables to the first storage medium, and store the second number of feature tables to the second storage medium.

[0044] In another possible implementation, the parameters include multiple feature tables related to the artificial intelligence model; when the processing module determines the storage resources for storing the parameters based on the parameter quantity and information about the available storage resources of the computing system, it is specifically used to: when the remaining storage capacity of the first storage medium is less than the parameter quantity of the parameter, determine that the determined storage resources include the first storage medium and the remote storage device; when storing the parameters to the determined storage resources, it is specifically used to: when the remaining storage capacity of the first storage medium is less than the parameter quantity of the parameter, store the first number of feature tables to the first storage medium, and store the second number of feature tables to the remote storage device.

[0045] In another possible implementation, the communication module is also used to obtain the user's performance requirements for artificial intelligence model inference or training; when the processing module determines the storage resources for storing parameters based on the parameter quantity and information about the available storage resources of the computing system, it is specifically used to: determine the storage resources for storing parameters based on the parameter quantity, performance requirements and information about the available storage resources of the computing system.

[0046] In a third aspect, a computing system is provided, which includes a general-purpose processor and multiple special-purpose processors, and the general-purpose processor and the multiple special-purpose processors jointly execute the operating steps of the method in the first aspect or any possible implementation of the first aspect.

[0047] In a fourth aspect, a computer system is provided, which includes a storage node and multiple computing nodes, the storage node is used to store parameters indicated by the computing node, and the computing nodes jointly execute the operating steps of the method in the first aspect or any possible implementation of the first aspect.

[0048] In a fifth aspect, a computer system is provided, which includes a memory and multiple processors, the memory being used to store a set of computer instructions; when the processor executes the set of computer instructions, the multiple processors jointly execute the operating steps of the method in the first aspect or any possible implementation of the first aspect.

[0049] In a sixth aspect, a computer-readable storage medium is provided, comprising: computer software instructions; when the computer software instructions are executed in a processor, the processor executes the operating steps of the method described in the first aspect or any possible implementation of the first aspect.

[0050] In a seventh aspect, a computer program product is provided. When the computer program product is run on a computer, it enables the computer to perform the operating steps of the method described in the first aspect or any possible implementation of the first aspect.

[0051] The technical effects brought about by any design method in the second to seventh aspects can be referred to the technical effects brought about by the first aspect or different design methods in the first aspect, and will not be repeated here.

[0052] Based on the implementation methods provided in the above aspects, this application can also be further combined to provide more implementation methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] FIG1 is a schematic diagram of the structure of a neural network provided by this application;

[0054] FIG2 is a schematic diagram of a sparse parameter provided by this application;

[0055] FIG3 is a schematic diagram of one-hot encoding and vectorization provided by this application;

[0056] FIG4 is a schematic diagram of the architecture of a parameter server provided by this application;

[0057] FIG5 is a schematic diagram of the architecture of a computer system provided by the present application;

[0058] FIG6 is a schematic diagram of the architecture of a computing system provided by the present application;

[0059] FIG7 is a schematic diagram of the architecture of a computer system provided by the present application;

[0060] FIG8 is a schematic diagram of the architecture of a computing system provided by the present application;

[0061] FIG9 is a schematic diagram of parameter storage of a medium model provided by this application;

[0062] FIG10 is a schematic diagram of parameter storage of a large model provided by this application;

[0063] FIG11 is a schematic diagram of parameter storage of a super-large model provided by this application;

[0064] FIG12 is a flow chart of an artificial intelligence model deployment method provided by this application;

[0065] FIG13 is a flow chart of an artificial intelligence model deployment method provided by this application;

[0066] FIG14 is a schematic diagram of the structure of an artificial intelligence model deployment device provided by this application;

[0067] FIG15 is a schematic structural diagram of a computer device provided in this application. DETAILED DESCRIPTION

[0068] To facilitate understanding, the main terms involved in this application are first explained.

[0069] Artificial Neural Network (ANN): also known as Neural Network (NN) or Neural Network-like. In the fields of machine learning and cognitive science, it is a mathematical or computational model that mimics the structure and function of biological neural networks (the central nervous system of animals, particularly the brain). A neural network is a network formed by connecting multiple individual neurons together, meaning that the output of one neuron can be the input of another. The input of each neuron can be connected to the local receptive field of the previous layer to extract features from that local receptive field, which can be an area consisting of several neurons.

[0070] Each node represents a specific output function, called an activation function. Each connection between two nodes represents a weighted value for the signal passing through that connection, called a weight, which acts as the memory of the artificial neural network. The output of the neural network varies depending on the network's connection structure, weight values, and activation function. Neural networks themselves are often approximations of natural algorithms or functions, or they may express a logical strategy.

[0071] As shown in Figure 1, it is a schematic diagram of the structure of a neural network provided in the present application. The neural network 100 includes N processing layers, where N is an integer greater than or equal to 3. The first layer of the neural network 100 is the input layer 110, which is responsible for receiving input signals, and the last layer of the neural network 100 is the output layer 130, which is responsible for outputting the processing results of the neural network. The other layers excluding the first and last layers are intermediate layers 140, and these intermediate layers 140 together constitute the hidden layer 120. Each intermediate layer 140 in the hidden layer 120 can both receive input signals and output signals. The hidden layer 120 is responsible for the processing of the input signal. Each layer represents a logical level of signal processing. Through multiple layers, the data signal can be processed by multiple levels of logic.

[0072] In some feasible embodiments, the input signal of the neural network can be a video signal, a voice signal, a text signal, an image signal, a temperature signal, an engineering signal that can be processed by a computer, or other signals in various forms.

[0073] Model training refers to the use of a training set to train an AI model, enabling it to predict or classify unknown data. During AI model training, learning is performed based on the features and target values ​​in the training set. Upon completion, an AI model is generated that can be used to predict or classify unknown data. AI model training is one of the most critical steps in machine learning, impacting the accuracy and reliability of the AI ​​model.

[0074] Parameters: These are variables or weights that need to be learned or adjusted in an AI model. Parameters can influence the AI ​​model's predictive power and performance. During AI model training, different parameter combinations are tried to optimize the model's performance. Common parameters include weights, biases, learning rates, and regularization coefficients. These parameters are used to calculate the output when the AI ​​model makes predictions.

[0075] Sparse parameters: Parameters that are only partially involved in the calculation during a training process. For example, parameters that participate in both forward calculations and backward updates.

[0076] For example, Figure 2 is a schematic diagram of a sparse parameter provided by this application. As shown in Figure 2, the weight part involved in the calculation in one training step only includes half of the total weight.

[0077] Typically, sparse parameters are very large. For example, in a production-level recommendation system, the size of sparse parameters can reach 10TB to 30TB.

[0078] Recommendation system: A type of application that provides users with personalized decision support and information services based on massive data mining, and determines the items or services that users currently need or are interested in based on information such as their historical behavior, social relationships, points of interest, and the context in which they are located.

[0079] In AI models like recommendation systems and natural language processing, input data often contains discrete features. Embedding table parameters are used to convert the input data into continuous vector data before processing the vector data. During a single training session, only a subset of the embedding table parameters are used for calculation and training updates.

[0080] Embedding technology: A dense vector representation. It converts sparse parameters into dense vector data. Embedding can represent object features, such as height, gender, name, or item.

[0081] For example, as shown in Figure 3, a schematic diagram of one-hot encoding and vectorization provided by this application is shown. Embedding is equivalent to smoothing one-hot encoding, and one-hot encoding is equivalent to performing max pooling on embedding.

[0082] During the training of an AI model, parameters (e.g., sparse parameters) are continuously adjusted so that the predicted values ​​after the input data and parameter calculations are close to the actual values. The parameters can be loaded into the storage medium of a dedicated processor, which then calculates and updates the parameters. For example, dedicated processors include, but are not limited to, graphics processing units (GPUs), data processing units (DPUs), neural processing units (NPUs), and embedded neural-network processing units (NPUs). The storage medium of a dedicated processor includes high-bandwidth memory (HBM).

[0083] For example, as shown in (a) in Figure 4, it is a schematic diagram of the architecture of a parameter server provided by this application. The Embedding table is entirely stored in the storage medium of the GPU (such as HBM). The storage medium of the GPU is divided into two parts, one part is used to store data during the training process of the artificial intelligence model; the other part is used to store the Embedding table. During the training process of the artificial intelligence model, the required Embedding vectors are pulled to the specified GPU through the All2All method to achieve the purpose of sharing the Embedding table. The Embedding table is referred to as the Embedding table. The Embedding table contains vector data of sparse parameters.

[0084] However, this solution places a high demand on GPUs. For example, approximately 300 GPU cards are required to store a 10TB Emb table, making AI model training extremely expensive. Furthermore, the large number of GPUs required leads to poor system linearity, insufficient overall system performance, and poor computational reliability. Since all GPUs participate in training using an All2All algorithm, a single GPU failure can easily lead to loss of Emb table data stored on the failed GPU.

[0085] The storage capacity of a dedicated processor is only 16GB to 32GB, which is far too little for tens of TB of sparse parameters. Sparse parameters can also be stored in storage media such as the server's volatile memory (such as RAM) or non-volatile memory (such as a solid state drive (SSD)).

[0086] In some embodiments, the sparse parameters are stored in a server including a general-purpose processor, which converts the sparse parameters into vector data, which is then transferred to a dedicated processor for computation.

[0087] For example, as shown in (b) in Figure 4, it is a schematic diagram of the architecture of a parameter server provided by this application. The Emb table is all stored in the CPU cluster; the GPU trains the artificial intelligence model. The CPU provides the functions of storing, querying and updating the Emb table, which is called a parameter server. The storage capacity of the GPU is expanded by the parameter server, and the CPU and GPU collaborate to perform forward calculations and reverse updates of the artificial intelligence model; during forward calculation, the required Emb table is pulled from the parameter server through the pull interface; the GPU trains the artificial intelligence model according to the Emb table to obtain the gradient; the gradient is pushed back to the parameter server through the push interface, and the Emb table is updated on the parameter server.

[0088] Since each time an artificial intelligence model is trained, a large number of Emb tables are transmitted through the network, that is, pull operations and push operations are performed, the computing efficiency is low.

[0089] In other embodiments, if the server cannot store the complete Emb table, the Emb table is stored in the server's storage medium and a remote storage device.

[0090] For example, as shown in (c) of Figure 4, a schematic diagram of the architecture of a parameter server provided by this application is provided. When the storage capacity of the memory is insufficient, the SSD is also used to store the Embed table. The CPU can pull the required Embed table from the SSD and transfer the Embed table to the GPU.

[0091] In order to solve the problem of low resource utilization and inability to fully utilize system performance when storing parameters of artificial intelligence models of different sizes based on a single storage solution, the present application provides an artificial intelligence model deployment method, namely, obtaining the parameter quantity of the parameters used by the artificial intelligence model for inference or training, and information about the available storage resources of the computing system where the artificial intelligence model is deployed, determining the storage resources for storing the parameters based on the parameter quantity and the information about the available storage resources of the computing system; and storing the parameters in the determined storage resources when deploying the artificial intelligence model. The storage resource information includes multiple types of storage resources, and each type of storage resource has different access performance.

[0092] When deploying an artificial intelligence model, for artificial intelligence models of different sizes, the parameter quantities used when the artificial intelligence model is reasoned or trained are different. In this case, the parameters are stored in at least one storage medium that meets the parameter quantity based on the information of the available storage resources of the computing system where the artificial intelligence model is deployed. Thus, the storage location of the parameters is flexibly allocated based on multi-dimensional information such as the parameter quantity and storage resource information, and the resources of the computing system are fully utilized, which effectively improves the resource utilization rate. In addition, when the dedicated processor trains or reasoned the artificial intelligence model, it can obtain the parameters from the determined storage resources as quickly as possible, and accelerate the execution of model training calculations or model reasoning calculations based on the parameters, giving full play to the system performance. Avoid using only one storage medium to store the parameters of the artificial intelligence model, which limits the storage location of the parameters, resulting in the dedicated processor being unable to obtain the parameters in a timely manner, and reducing the training or reasoning speed of the artificial intelligence model.

[0093] The implementation of the artificial intelligence model deployment method provided in this application is described in detail below with reference to the accompanying drawings.

[0094] FIG5 is a schematic diagram of the architecture of a computer system provided by the present application. As shown in FIG5 , the computer system 500 includes a client 510 , a computing cluster 520 , and a storage cluster 530 .

[0095] The computing cluster 520 includes multiple computing nodes 521. The multiple computing nodes 521 can be connected through network devices (such as switches, network cards, etc.) based on high-speed interconnection technology, so that the multiple computing nodes 521 can communicate with each other.

[0096] In some embodiments, the computing node 521 may include computing units with computing capabilities such as a graphics processing unit (GPU), a data processing unit (DPU), a neural processing unit (NPU), and an embedded neural-network processing unit (NPU) to provide high-performance computing.

[0097] The computing cluster 520 further includes a control node 522. The control node 522 is used to manage and allocate tasks, and multiple computing nodes execute multiple tasks in parallel to increase the data processing rate.

[0098] Computing nodes with computing capabilities (such as GPUs, DPUs, and NPUs) can be referred to as devices, training cards, or accelerator cards. Control nodes can be referred to as hosts. Alternatively, computing nodes can be referred to as dedicated processors. Control nodes can be referred to as general-purpose processors.

[0099] In the present application, the control node 522 is also used to perform preprocessing operations on the parameters required for model training or model reasoning, and cooperate with multiple computing nodes 521 to manage parameters. For example, the control node 522 vectorizes the sparse parameters and converts them into dense vector data (such as: Emb table). Furthermore, when the control node 522 obtains the model training task or the model reasoning task, the storage resources for storing the parameters are determined based on the parameter quantity and the information of the available storage resources of the computing system. The determined storage resources include at least one of the storage medium associated with the computing node, the storage medium associated with the control node, or the storage medium associated with the storage node. That is, for artificial intelligence models of different scales, since the parameter quantities of the parameters that need to be loaded are different, the control node 522 performs hierarchical storage on the parameters with different storage requirements, that is, the parameters related to model training or model reasoning are stored in the storage medium of at least one of the computing nodes, control nodes, or storage nodes.

[0100] When computing node 521 performs training calculations or inference calculations on an artificial intelligence model based on parameters, if the parameters are stored in a storage medium of computing node 521 , computing node 521 obtains parameters related to model training or model inference from the storage medium of computing node 521 .

[0101] If the parameters are stored in the storage medium of the control node 522 or the storage medium of the storage node 531, the control node 522 obtains the parameters from the storage medium of the control node 522 or the storage medium of the storage node 531, loads the parameters to the computing node 521, and the computing node 521 performs training calculations or inference calculations on the artificial intelligence model based on the parameters.

[0102] Optionally, if the parameters are stored in the storage medium of the control node 522 or the storage medium of the storage node 531 , the computing node 521 may also obtain parameters related to model training or model inference from the storage medium of the control node 522 or the storage medium of the storage node 531 .

[0103] The storage cluster 530 includes multiple storage nodes 531. A storage node 531 includes one or more controllers, a network card, and multiple hard disks. The hard disks are used to store data. The hard disks can be magnetic disks or other types of storage media, such as solid-state drives or shingled magnetic recording hard disks. The network cards are used to communicate with the computing nodes 521 included in the computing cluster 520. The controllers are used to write data to or read data from the hard disks based on read / write data requests sent by the computing nodes 521. During the data reading and writing process, the controllers need to convert the addresses carried in the read / write data requests into addresses that the hard disks can recognize.

[0104] The storage cluster 530 includes multiple storage nodes 531 and the computing cluster 520 includes multiple computing nodes 521 which are connected through network devices (such as switches, network cards, etc.) based on high-speed interconnection technology, so that the multiple computing nodes 521 and the multiple storage nodes 531 can communicate with each other.

[0105] In this application, the storage nodes 531 included in the storage cluster 530 can serve as the remote storage devices described herein. For example, the storage node can be a server, which includes a CPU and storage media. If the remaining storage capacity in the computing cluster is insufficient, the storage cluster 530 can provide storage expansion capabilities, namely, storing parameters related to model training or model inference, such as the Embed table required for model training or model inference.

[0106] This application does not limit the deployment form of the computing nodes 521 and the control nodes 522. For example, the computing nodes 521 and the control nodes 522 can be deployed on the same server or multiple servers.

[0107] In other embodiments, the computing node 521 and the control node 522 can each be an independent server, etc. The computing node 521 can be a server that provides a heterogeneous computing architecture to provide high-performance computing. The computing node 521 includes a dedicated processor and a general-purpose processor. For example, the general-purpose processor can include a central processing unit (CPU). The dedicated processor includes a computing node with computing capabilities such as a GPU, a DPU, and an NPU to provide high-performance computing. The control node 522 can be an ordinary application server, which includes a CPU and a storage medium.

[0108] The control node 522 is used to manage and allocate tasks, instructing multiple computing nodes to execute multiple tasks in parallel to improve the training or reasoning speed of the artificial intelligence model.

[0109] The general-purpose processor in the computing node 521 is used to determine the storage resources for storing the parameters based on the parameter quantity and the available storage resources of the computing system. The dedicated processor in the computing node 521 is used to perform training calculations or inference calculations on the artificial intelligence model based on the parameters.

[0110] Optionally, the server used to store the parameters required for model training or model reasoning can be called a parameter server. The server used to perform model training calculations or model reasoning calculations can be called a training server. The training server can include multiple accelerator cards. For example, as shown in Figure 6, an architectural diagram of a computing system provided in this application is provided. The computing system includes 3 parameter servers and 3 training servers, each of which includes 2 NPUs. The parameter server is used to store the parameters required for model training or model reasoning (such as: Emb table). The training server is used to perform training calculations or reasoning calculations on the artificial intelligence model based on the parameters.

[0111] Client 510 communicates with computing cluster 520 and storage cluster 530 via network 540. For example, client 510 sends a request to computing cluster 520 via network 540, requesting that computing cluster 520 perform model training or model inference. Network 540 can be an internal enterprise network (e.g., a local area network (LAN)) or the Internet. Client 510 can be a computer connected to network 540, also known as a workstation. Different clients can share network resources (e.g., computing resources and storage resources).

[0112] In some embodiments, client 510 is installed with a client program 511. Client 510 runs client program 511 to display a user interface (UI). User 550 operates the UI to submit a request. For example, user 550 operates the UI to submit a model training request or a model inference request. After receiving the request, control node 522 can obtain the parameters required for model training or model inference from storage cluster 530 and determine the storage resources used to store the parameters based on the parameter quantity and the available storage resources of the computing system.

[0113] Optionally, the system administrator 560 can call the application platform interface (API) 512 or the command-line interface (CLI) interface 513 through the client 510 to configure system information, for example, the system information includes the storage strategy of the parameters required for model training or model inference.

[0114] FIG5 is merely a schematic diagram, and the embodiments of the present application do not limit the device connection method, device quantity, and device form in the computer system.

[0115] For example, Figure 7 is a schematic diagram of the architecture of a computer system provided in this application. Dedicated servers and general-purpose servers are interconnected through network devices based on high-speed interconnection technology. The computer system may include dedicated servers and general-purpose servers. The dedicated server and the general-purpose server can each be an independent server. The dedicated server includes a storage medium, a general-purpose processor, and multiple dedicated processors. The dedicated server can be an artificial intelligence server. The artificial intelligence server includes a general-purpose processor, a storage medium, and multiple dedicated processors. Multiple artificial intelligence servers are interconnected via a network, and multiple artificial intelligence servers form an AI cluster to implement model training calculations and model inference calculations. The general-purpose server includes a general-purpose processor.

[0116] Among them, general servers are used to manage and allocate tasks, and multiple dedicated servers execute multiple tasks in parallel to improve the training or inference speed of artificial intelligence models.

[0117] The general-purpose processor in the dedicated server is used to determine storage resources for storing parameters based on the parameter quantity and information about available storage resources of the computing system. The determined storage resources include at least one of the storage medium of the dedicated processor, the storage medium of the dedicated server, and the storage medium of the general-purpose server. The storage medium is used to store parameters related to training calculations or inference calculations of the artificial intelligence model, such as the storage medium used to store an Emb table or a portion of the data in the Emb table.

[0118] Dedicated processors in dedicated servers are used to perform training calculations or inference calculations on artificial intelligence models based on parameters.

[0119] It should be noted that the specific form of the storage medium described in this application is not limited. The storage medium includes a volatile memory pool or a non-volatile memory pool, or may include both volatile and non-volatile memory. For example, the storage medium of a dedicated processor includes at least one of HBM or SSD.

[0120] The following describes the hierarchical storage of parameters required for model training calculations or model inference calculations based on a multi-level parameter server provided by this application.

[0121] Figure 8 is a schematic diagram of the architecture of a computing system provided by the present application. The computing system provides a unified architecture of multi-level parameter servers. As shown in Figure 8, the computing system 800 includes one or more local servers 810 and one or more storage servers 820. The local server 810 includes a general-purpose processor and multiple dedicated processors. The general-purpose processor and multiple dedicated processors can be deployed together on the same server. The storage medium in the dedicated processor serves as the primary storage layer, and the storage medium in the general-purpose processor serves as the secondary storage layer, that is, the storage medium in the general-purpose processor serves as the local parameter server (Local PS Server), and the storage server 820 serves as the tertiary storage layer, that is, the storage server 820 serves as the remote parameter server (Remote PS Server).

[0122] The storage medium in the dedicated processor is used to store parameters in the model training or model inference process, as well as all or part of the parameters of the Emb table.

[0123] The storage medium in the general-purpose processor is used to store all or part of the parameters of the Emb table, gradient accumulation, optimizer data, cache data, etc.

[0124] The storage server is used to store all or part of the parameters of the Emb table, gradient accumulation, optimizer data, cache data, etc.

[0125] The general-purpose processor is used to perform preprocessing operations on the parameters required for model training or model inference. For example, sparse parameters are vectorized and converted into dense vector data to obtain an Embed table. The general-purpose processor is also used to store the Embed table in a determined storage resource based on the storage requirements of the Embed table and the storage resource characteristics of the system. The determined storage resource includes a storage medium in at least one of the aforementioned primary storage layer, secondary storage layer, and tertiary storage layer.

[0126] A general-purpose processor can run multiple worker processes, and each worker process can control a dedicated processor. For example, a worker can instruct a dedicated processor to perform model training or model inference calculations, pull the Emb table and transmit it to the dedicated processor, and then update the Emb table based on the gradient pushed back from the dedicated processor.

[0127] Optionally, the model described in this application may refer to a recommendation model.

[0128] Based on the above-mentioned unified multi-level parameter server architecture, unified hierarchical storage solutions are proposed for medium-sized models, large-scale models, and ultra-large-scale models.

[0129] In the first possible implementation, the Emb table corresponding to the medium-sized model has a small number of parameters and requires a small storage capacity. The Emb table can be stored in the first-level storage layer, that is, the Emb table is stored in the storage medium in the dedicated processor.

[0130] In some embodiments, the Emb table can be divided into multiple data blocks according to the number of dedicated processors, and the multiple data blocks of the Emb table can be stored in the storage media of multiple dedicated processors respectively. The way of dividing the Emb table is not limited in this application. The number of data blocks of the Emb table can also be less than the number of dedicated processors, and the data blocks of the Emb table are stored in some of the multiple dedicated processors. The number of data blocks of the Emb table can also be greater than or equal to the number of dedicated processors, and the data blocks of the Emb table are stored in the storage media of multiple dedicated processors. One dedicated processor stores one data block, or one dedicated processor can also store multiple data blocks.

[0131] For example, as shown in Figure 9, the Emb table of a medium-sized model is stored in the storage medium of a dedicated processor (e.g., HBM), which is referred to as Form 1. Each time the dedicated processor iteratively trains or infers the model, it retrieves the Emb vector in the Emb table from the dedicated processor's storage medium. After training is complete, the Emb vector in the HBM is updated based on the gradient. Alternatively, after training is complete, a general-purpose processor can update the Emb table based on the gradient.

[0132] In other embodiments, dedicated processors are expanded, that is, dedicated processors are added to the local server, so that the Emb table is stored based on the storage media in multiple dedicated processors.

[0133] In the second possible implementation, the Embed table corresponding to a large model has a large number of parameters and requires a large storage capacity. For example, the Embed table can have up to 10TB of parameters. If the Embed table is stored on the storage medium of a dedicated processor, assuming each dedicated processor has a storage capacity of 32GB, 300 dedicated processors would be required to start training, which is too many dedicated processors. Alternatively, the Embed table can be stored on the secondary storage tier, that is, on the storage medium of a general-purpose processor, or on a local parameter server.

[0134] For example, as shown in FIG10 , the general-purpose processor may store the large-scale model Emb table to the storage medium (eg, DDR) of the general-purpose processor, that is, store the Emb table to the local parameter server, which is referred to as form 2.

[0135] After the general processor receives a model training task request or a model inference task request, the worker can obtain the parameter set identifier from the request. The identifier can indicate the Emb vector in the Emb table. The general processor obtains the Emb vector corresponding to the identifier and transmits it to the dedicated processor.

[0136] Optionally, the HBM in the dedicated processor is used to cache a portion of the Emb vectors in the full Emb table, and the storage medium in the general processor is used to store another portion of the Emb vectors in the full Emb table.

[0137] When a dedicated processor trains an AI model or infers an AI model, it reads the required Emb vector from the HBM. After training is complete, it updates the Emb vector in the HBM based on the gradient. Alternatively, after training is complete, a general-purpose processor can update the Emb vector in the HBM based on the gradient.

[0138] In some embodiments, the dedicated processor periodically writes the Emb vector in the HBM to the local parameter server according to a cycle / event trigger.

[0139] Optionally, the Emb table of a large-scale model may also be stored in a remote parameter server, and the general processor may obtain the Emb vector corresponding to the identifier from the remote parameter server.

[0140] In the third possible implementation, the Embed table corresponding to an ultra-large-scale model has a very large number of parameters, for example, up to 100TB. Storing the Embed table on dedicated processors would require thousands of dedicated processors to store the Embed table, and using thousands of dedicated processors to train an ultra-large-scale model would be prohibitively expensive. Alternatively, the Embed table can be stored on a tertiary storage tier, storing it on the storage media of a storage server, or on a remote parameter server.

[0141] For example, as shown in FIG11 , a general-purpose processor may store an ultra-large-scale model Emb table to a storage server, that is, store the Emb table to a remote parameter server, which is referred to as form three.

[0142] The general processor reads the Emb vector corresponding to the identifier from the remote parameter server and transmits it to the dedicated processor.

[0143] The HBM in the dedicated processor is used to cache a portion of the Emb vectors in the full Emb table.

[0144] Optionally, when there are many workers, the workers can be grouped, with workers in a group handing off requests to a general-purpose processor. The local parameter server can provide deduplication and caching services for the workers; ultimately, the Emb vectors are transferred to the storage medium in the dedicated processor.

[0145] The unified architecture of multi-level parameter servers provides a hierarchical storage path, expanding the storage methods of Embed tables. This supports hierarchical storage of Embed tables for models of varying sizes and various hardware deployment forms, enabling a naturally scalable Embed table storage solution. A single architecture supports three recommendation models: medium-scale, large-scale, and extra-large Embed, eliminating the need to develop and deploy specific software for each scenario.

[0146] For example, for a medium-sized model, by expanding dedicated processors, the Embedding table is divided into blocks and stored in the storage media (Embedding Storage) of multiple dedicated processors, and multiple Embedding Storages are synthesized into a complete Embedding table.

[0147] For example, for large-scale models, the Emb table is stored in the storage medium of the general-purpose processor, that is, the local parameter server. The storage medium in the dedicated processor serves as a cache layer, and the Emb table is read from the general-purpose processor on demand for model training or model inference.

[0148] For example, for extremely large models, the Emb table is stored in a remote storage device, namely a remote parameter server. The storage medium in the dedicated processor serves as a cache layer, and the Emb table is read from the remote storage device on demand for model training or model inference.

[0149] In some embodiments, the unified architecture of the multi-level parameter server provides a unified programming interface to the outside world. The interface definition is as follows:

[0150] The SparseEmbeddingManager is used to manage the Embed table, such as creating an Embed table and setting parameter sets.

[0151] SparseEmbedding corresponds to an Embedding table and provides corresponding search and update methods based on the storage type (storage_type). External interfaces include: Embedding vector search (embedding_lookup); Embedding vector update (embedding_update); Embedding table storage (save). Key attributes include the storage type (storage_type) and the Embedding table storage method.

[0152] The unified architecture of the multi-level parameter server presents a unified programming interface to the outside world. For models of different sizes, user code does not need to be modified.

[0153] Next, the artificial intelligence model deployment process is explained in detail with reference to the accompanying drawings.

[0154] FIG12 is a flow chart of an artificial intelligence model deployment method provided by this application. Here, the storage of sparse parameters related to the training calculation or inference calculation of the artificial intelligence model is mainly described. As shown in FIG12, the method includes the following steps 1210 to 1230.

[0155] Step 1210: The general processor obtains the parameter quantities of the parameters used by the artificial intelligence model for reasoning or training, and information about the available storage resources of the computing system where the artificial intelligence model is deployed.

[0156] This application does not limit the original storage location of the parameters to be stored. A general-purpose processor can obtain the parameters to be stored from a storage device such as a storage server or storage array. The storage resource information includes multiple types of storage resources, each of which has different access performance.

[0157] Step 1220: The general purpose processor determines storage resources for storing the parameters based on the parameter quantity and information about available storage resources of the computing system.

[0158] Step 1230: When the general processor deploys the artificial intelligence model, the parameters are stored in the determined storage resources.

[0159] The parameters to be stored include sparse parameters related to the training calculation or inference calculation of the artificial intelligence model. For example, the parameters include one or more Emb tables.

[0160] The information about the storage resources of the computing system includes the remaining storage capacity of the storage resources. For example, the information about the storage resources of the computing system includes at least one of the remaining storage capacity of the storage resources in the primary storage tier, the remaining storage capacity of the storage resources in the secondary storage tier, or the remaining storage capacity of the storage resources in the secondary storage tier as shown in FIG8 . The remaining storage capacity may refer to the storage capacity of the storage space where no data is stored or the storage capacity of the storage space that can be used.

[0161] In a first possible implementation, a general-purpose processor determines a storage resource for storing the parameter based on the parameter quantity and information about available storage resources of the computing system. The determined storage resource includes at least one of a first storage medium associated with the dedicated processor, a second storage medium associated with the general-purpose processor, or a remote storage device.

[0162] The first storage medium associated with the dedicated processor includes a storage medium in the dedicated processor. The first storage medium includes volatile memory, non-volatile memory, or both. The volatile memory includes at least one of HBM or DDR. The non-volatile memory may include an SSD.

[0163] The second storage medium associated with the general-purpose processor includes a storage medium in the general-purpose processor and a storage medium located on the same server as the general-purpose processor. The second storage medium includes volatile memory, non-volatile memory, or both. For example, the second storage medium associated with the general-purpose processor is a local parameter server.

[0164] The remote storage device includes a storage medium that belongs to a server different from the general-purpose processor. The remote storage device is primarily used to store parameters. For example, the remote storage device includes a storage array. For example, the remote storage device can be a remote parameter server.

[0165] The general-purpose processor determines whether the remaining storage capacity of the storage resource meets the storage capacity required to store the parameter. If so, the general-purpose processor determines the storage resource to be determined. It is understood that the general-purpose processor determines a storage medium from the system that meets the parameter capacity, and the storage medium that meets the parameter capacity is capable of storing the parameter. Storage media that meet the parameter capacity include at least one storage medium in the system. For example, the general-purpose processor determines whether the remaining storage capacity of the storage resource is greater than the parameter capacity of the parameter. If so, the general-purpose processor determines the storage resource to be determined.

[0166] For example, the general-purpose processor determines whether the remaining storage capacity of the first storage medium associated with the dedicated processor is greater than the parameter quantity of the parameter. If the remaining storage capacity of the first storage medium is greater than the parameter quantity of the parameter, the first storage medium associated with the dedicated processor is determined as the determined storage resource, and the parameter is stored in the first storage medium associated with the dedicated processor.

[0167] Alternatively, the general-purpose processor may divide the Emb table into multiple data blocks and store the multiple data blocks on multiple storage media in the dedicated processor. For example, one storage medium may store one data block, and multiple data blocks stored on multiple storage media may form a complete Emb table. In another example, one storage medium may store multiple data blocks.

[0168] For another example, if the remaining storage capacity of the first storage medium is less than the parameter amount of the parameter, it is determined that the determined storage resource includes at least one of the second storage medium or the remote storage device. That is, at least one of the second storage medium or the remote storage device associated with the general-purpose processor is used to assist the dedicated processor in storing the Emb table.

[0169] For example, the general processor may divide the Emb table into a first parameter set and a second parameter set, store the first parameter set in a storage medium in the dedicated processor, and store the second parameter set in at least one of a second storage medium associated with the general processor or a remote storage device.

[0170] As another example, the general-purpose processor can divide multiple Emb tables into a first number of Emb tables and a second number of Emb tables, store the first number of Emb tables to the storage medium in the dedicated processor, and store the second number of Emb tables to at least one of the second storage medium or remote storage device associated with the general-purpose processor.

[0171] Optionally, if the remaining storage capacity of the first storage resource associated with the dedicated processor is equal to the parameter amount of the parameter, the parameter may be stored in the first storage medium associated with the dedicated processor.

[0172] Optionally, if the remaining storage capacity of the second storage resource associated with the general-purpose processor is equal to the parameter quantity of the parameter, the parameter may be stored in the second storage medium associated with the general-purpose processor.

[0173] For details about the storage process of sparse parameters related to the training calculation or inference calculation of the artificial intelligence model, please see Figure 13.

[0174] FIG13 is a schematic diagram of the storage process of sparse parameters provided in this application.

[0175] Determine whether the remaining storage capacity of the first storage medium (e.g., HBM) is greater than the parameter quantity of the parameter, for example, whether the HBM can store the Embed table (step 1310). When the remaining storage capacity of the first storage medium is greater than the parameter quantity of the parameter, store the parameter in the first storage medium, that is, store the Embed table in the card (step 1320). When the remaining storage capacity of the first storage medium is less than the parameter quantity of the parameter, determine whether the number of Embed tables is 1 (step 1330).

[0176] If an Embed table exists, a determination is made as to whether the remaining storage capacity of the second storage medium (e.g., the local parameter server) is greater than the parameter quantity of the parameters, for example, whether the local parameter server can store the Embed table (step 1340). If the remaining storage capacity of the second storage medium is greater than the parameter quantity of the parameters, the parameters are stored in the second storage medium, i.e., the Embed table is stored in the local parameter server (step 1350). If the remaining storage capacity of the second storage medium is less than the parameter quantity of the parameters, a determination is made as to whether the number of workers is small (step 1360). If the number of workers is large, the parameters are stored in the remote storage device and the second storage medium, i.e., the Embed table is stored in the remote parameter server and the local parameter server (step 1370). For example, a first number of parameters is stored in the second storage medium, and a second number of parameters is stored in the remote storage device. If the number of workers is small, the parameters are stored in the remote storage device, i.e., the Embed table is stored in the remote parameter server (step 1380). For example, if the remaining storage capacity of the second storage medium is less than the parameter quantity of the parameters, the parameters are stored in the remote storage device.

[0177] If there are multiple Emb tables, when the remaining storage capacity of the first storage medium is less than the parameter quantity of the parameter, the first number of Emb tables are stored in the first storage medium, that is, the first number of Emb tables are stored in the card (step 1390), and it is determined whether the remaining storage capacity of the second storage medium (such as the local parameter server) is greater than the remaining parameter quantity, that is, whether the local parameter server can store the remaining Emb tables (step 13100). The local parameter server can store the remaining Emb tables, and the second number of feature tables are stored in the second storage medium, that is, the second number of Emb tables are stored in the local parameter server (step 13110).

[0178] When the remaining storage capacity of the first storage medium is less than the parameter amount of the parameter, the first number of Emb tables is stored in the first storage medium (step 1390), the local parameter server cannot store the remaining Emb tables, and the second number of feature tables is stored in the remote storage device, that is, the second number of Emb tables is stored in the remote parameter server (step 13120).

[0179] In a second possible implementation, the general-purpose processor determines the determined storage resource based on the parameter quantity of the parameter and the remaining storage capacity and storage type of the storage resource. Based on the remaining storage capacity and storage type of the storage resource, the general-purpose processor determines a storage medium that satisfies the parameter quantity of the parameter, determines the storage medium that satisfies the parameter quantity of the parameter as the determined storage resource, and stores the parameter in the determined storage resource. The storage type may refer to the type of storage medium. Storage medium types include both volatile memory and non-volatile memory.

[0180] In some embodiments, the general-purpose processor may first select a volatile memory from the system, determine whether the remaining storage capacity of the volatile memory satisfies the parameter quantity of the parameter, and if so, determine the volatile memory as the determined storage resource. If the remaining storage capacity of the volatile memory does not satisfy the parameter quantity of the parameter, determine the non-volatile memory as the determined storage resource.

[0181] In other embodiments, the general-purpose processor can also obtain the user's performance requirements for artificial intelligence model reasoning or training; and determine the storage resources used to store parameters based on the parameter quantity, performance requirements, and information about the available storage resources of the computing system.

[0182] Different businesses have different performance requirements for model training or model inference. For example, performance requirements include data processing speed and data access speed.

[0183] For example, for businesses with high requirements for data processing speed and data access speed, a storage medium that meets the performance requirements is selected from the system. For example, if the access speed of the storage medium meets the data access speed requirement and the dedicated processor to which the storage medium belongs meets the data processing speed requirement, a determination is then made as to whether the remaining storage capacity of the storage medium that meets the performance requirements is greater than the parameter quantity of the parameter. If the remaining storage capacity of the storage medium that meets the performance requirements is greater than the parameter quantity of the parameter, the determined storage resource is determined to be the storage medium that meets the performance requirements, and the parameter is stored in the storage medium that meets the performance requirements. If the remaining storage capacity of the storage medium that meets the performance requirements is less than the parameter quantity of the parameter, then other storage media assist the dedicated processor in storing the parameter.

[0184] Optionally, the general-purpose processor can also jointly determine the storage resources based on multi-dimensional characteristics such as the performance and cost ratio of the dedicated processor or the remaining storage capacity. The most cost-effective storage medium is determined to store the Emb table based on factors such as performance and storage requirements. Thus, by combining the model's required hardware resources, available hardware resources, and the user's expected processing performance, appropriate hardware resource storage parameters are automatically selected, eliminating the need for complex manual selection and testing. This improves resource utilization, the training or inference speed of the AI ​​model, and fully maximizes system performance.

[0185] Furthermore, after the parameters are stored in the determined storage resource, when it is necessary to perform training calculations or inference calculations on the artificial intelligence model, the parameters are obtained from the determined storage resource. This application may also include step 1240.

[0186] Step 1240: Obtain parameters from the determined storage resources, and the dedicated processor performs training calculations or inference calculations on the artificial intelligence model based on the parameters.

[0187] If the Emb table is stored in a storage medium in a dedicated processor, the Emb table is obtained from the storage medium in the dedicated processor, and the dedicated processor then performs model training calculations or model inference calculations based on the data in the Emb table.

[0188] If the Emb table is stored in the storage medium of the general-purpose processor, the Emb table is obtained from the storage medium of the general-purpose processor, and the dedicated processor then performs model training calculations or model inference calculations based on the data in the Emb table.

[0189] If the Emb table is stored in a remote storage device, the Emb table is obtained from the storage medium in the remote storage device, and the dedicated processor then performs model training calculations or model inference calculations based on the data in the Emb table.

[0190] Optionally, a general-purpose processor or a dedicated processor can obtain the data in the Emb table from the determined storage resources, load the data in the Emb table into the dedicated processor, and the dedicated processor performs model training calculations or model inference calculations based on the data in the Emb table.

[0191] Therefore, the Emb table storage solution provided by this application, that is, when storing the Emb table required for model training or model inference, uses a dedicated processor, a general-purpose processor, and a storage medium provided by a storage device to jointly store the Emb table. If the storage medium of the dedicated processor is capable of storing the Emb table, the complete Emb table is divided into data blocks and stored in the dedicated processor. During model training or model inference, each dedicated processor pulls the Emb vector it needs on demand for calculation.

[0192] When the dedicated processor's storage medium cannot store the Emb table, the Emb table is stored in the general-purpose processor's storage medium. A portion of the Emb table is cached on demand in the dedicated processor's storage medium. During model training or inference, the required Emb vector is pulled from the cache for calculation.

[0193] When the storage medium of a general-purpose processor cannot store the Emb table, a remote storage device is used to store the Emb table. Part of the Emb table is cached on demand in the storage medium of a dedicated processor. During model training or model inference, the required Emb vector is pulled from the cache for calculation.

[0194] A unified multi-level parameter server architecture is used to support the above three scenarios, provide a unified external interface, and provide an automatic deployment component that automatically selects the optimal hardware facilities for deployment based on the data volume of the Embed table, hardware resources, and the user's pursuit of extreme performance or optimal cost-effectiveness.

[0195] It is understood that in order to implement the functions in the above embodiments, the computer device includes hardware structures and / or software modules corresponding to the execution of each function. It should be readily apparent to those skilled in the art that, in combination with the units and method steps of each example described in the embodiments disclosed in this application, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in hardware or in a manner driven by computer software depends on the specific application scenario and design constraints of the technical solution.

[0196] The above description is in detail based on the artificial intelligence model deployment method provided by the present application in conjunction with Figures 1 to 13. The following description will be based on the device provided by the present application in conjunction with Figure 14. These devices can be used to implement the functions of the dedicated processor or general-purpose processor in the above method embodiment, and thus can also achieve the beneficial effects of the above method embodiment. In this embodiment, the device can be a dedicated processor or a general-purpose processor as shown in Figure 12, or it can be a module (such as a chip) applied to a computer device.

[0197] As shown in Figure 14, the artificial intelligence model deployment device 1400 includes a communication module 1401, a processing module 1402 and a storage module 1403.

[0198] The artificial intelligence model deployment device 1400 is used to implement the functions of the general processor in the method embodiment shown in Figure 12 or Figure 13 above.

[0199] The communication module 1401 is used to obtain a request, where the request is used to instruct the execution of training calculations or inference calculations of an artificial intelligence model.

[0200] The communication module 1401 is also used to obtain the parameter quantities of the parameters used by the artificial intelligence model for reasoning or training, and information about the available storage resources of the computing system where the artificial intelligence model is deployed. For example, the communication module 1401 is used to execute step 1210 in Figure 12.

[0201] Processing module 1402 is configured to determine a storage resource for storing the parameters based on the parameter quantity and information about the available storage resources of the computing system, store the parameters in the determined storage resource, and store the parameters in the determined storage resource when deploying the artificial intelligence model. The determined storage resource includes at least one of a first storage medium associated with a dedicated processor, a second storage medium associated with a general-purpose processor, or a remote storage device. For example, processing module 1402 is configured to execute steps 1220 and 1230 in Figure 12. For example, processing module 1402 is configured to execute steps 1310 to 13120 in Figure 13.

[0202] The artificial intelligence model deployment device 1400 is used to implement the functions of the dedicated processor in the method embodiment shown in Figure 12 or Figure 13 above.

[0203] Optionally, the communication module 1401 is used to obtain a request, where the request is used to instruct execution of training calculations or inference calculations of the artificial intelligence model.

[0204] Optionally, the processing module 1402 is configured to obtain parameters from the determined storage resources and perform training calculations or inference calculations on the artificial intelligence model based on the parameters. For example, the processing module 1402 is configured to execute step 1240 in FIG. 12 .

[0205] The storage module 1403 is used to store the Emb table, etc., to facilitate model training or model inference.

[0206] It should be understood that the artificial intelligence model deployment device 1400 of the embodiment of the present application can be implemented by an application-specific integrated circuit (ASIC) or a programmable logic device (PLD), and the above-mentioned PLD can be a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL) or any combination thereof. When the method shown in Figure 12 or Figure 13 is implemented by software, and each module thereof can also be a software module, the artificial intelligence model deployment device 1400 and each module thereof can also be a software module.

[0207] According to the embodiment of the present application, the artificial intelligence model deployment device 1400 can correspond to executing the method described in the embodiment of the present application, and the above-mentioned and other operations and / or functions of each unit in the artificial intelligence model deployment device 1400 are respectively for implementing the corresponding processes of each method in 12 or Figure 13. For the sake of brevity, they will not be repeated here.

[0208] FIG15 is a schematic diagram of the structure of a computer device 1500 provided in this application. As shown in FIG15 , computer device 1500 includes a processor 1510, a bus 1520, a memory 1530, a communication interface 1540, a memory 1550 (also referred to as a main memory unit), and a processor 1560. Processor 1510, processor 1560, memory 1530, memory 1550, and communication interface 1540 are connected via bus 1520.

[0209] It should be understood that in this embodiment, the processor 1510 may be a CPU, but may also be other general-purpose processors, digital signal processors (DSP), ASICs, FPGAs or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.

[0210] Computer device 1500 may also include a graphics processing unit (GPU), a neural network processing unit (NPU), a microprocessor, an ASIC, or one or more integrated circuits for controlling the execution of the program of the present application. For example, processor 1560 may be a GPU or an NPU.

[0211] The communication interface 1540 is used to implement communication between the computer device 1500 and external devices or components.

[0212] In the present application, when the computer device 1500 is used to implement the functions of the general-purpose processor shown in Figure 12 or Figure 13, the communication interface 1540 is used to obtain requests, etc., so that the processor 1560 can determine the storage resources for storing parameters based on the parameter amount and the information about the available storage resources of the computing system, and store the parameters in the determined storage resources when deploying the artificial intelligence model.

[0213] When the computer device 1500 is used to implement the functions of the dedicated processor shown in FIG. 12 or FIG. 13 , the communication interface 1540 is used to transmit data, etc., so that the processor 1510 or the processor 1560 is used to instruct the acquisition of parameters from the determined storage resource, and the processor 1560 performs training calculations or inference calculations on the artificial intelligence model according to the parameters.

[0214] Bus 1520 may include a path for transmitting information between the aforementioned components (e.g., processor 1510, memory 1550, and storage 1530). In addition to a data bus, bus 1520 may also include a power bus, a control bus, and a status signal bus. However, for the sake of clarity, various buses are labeled as bus 1520 in the figure. Bus 1520 may be a Peripheral Component Interconnect Express (PCIe) bus, an extended industry standard architecture (EISA) bus, a unified bus (Ubus or UB), a compute express link (CXL), a cache coherent interconnect for accelerators (CCIX), etc. Bus 1520 may be divided into an address bus, a data bus, a control bus, etc.

[0215] As an example, computer device 1500 may include multiple processors. The processor may be a multi-core (multi-CPU) processor. A processor herein may refer to one or more devices, circuits, and / or computing units for processing data (e.g., computer program instructions).

[0216] It is worth noting that FIG15 only uses a computer device 1500 including one processor 1510 and one memory 1530 as an example. Here, the processor 1510 and the memory 1530 are respectively used to indicate a type of device or equipment. In specific embodiments, the number of each type of device or equipment can be determined based on business requirements. For example, the computer device 1500 may include multiple GPUs or multiple NPUs.

[0217] Memory 1550 may be a volatile memory pool or a nonvolatile memory pool, or may include both volatile and nonvolatile memory. Nonvolatile memory may be read-only memory (ROM), programmable ROM (PROM), erasable programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may be random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic random access memory (DRAM), synchronous DRAM (SDRAM), double data rate synchronous DRAM (DDR SDRAM), enhanced synchronous DRAM (ESDRAM), synchronous link DRAM (SLDRAM), and direct RAM bus RAM (DR RAM). Memory 1550 is used to store the Emb table, etc.

[0218] The memory 1530 may correspond to the storage medium used to store information such as the Emb table in the above method embodiment, for example, a disk such as a mechanical hard disk or a solid-state drive.

[0219] The computer device 1500 may be a general-purpose device or a dedicated device. For example, the computer device 1500 may be a server or other device with computing capabilities.

[0220] It should be understood that the computer device 1500 according to this embodiment may correspond to the artificial intelligence model deployment device 1400 in this embodiment, and may correspond to executing the corresponding subject in any method in Figure 12 or Figure 13, and the above-mentioned and other operations and / or functions of each module in the artificial intelligence model deployment device 1400 are respectively for implementing the corresponding processes of each method in Figure 12 or Figure 13. For the sake of brevity, they will not be repeated here.

[0221] The method steps in this embodiment can be implemented by hardware or by a processor executing software instructions. The software instructions can be composed of corresponding software modules, which can be stored in random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, mobile hard disks, CD-ROMs, or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be an integral part of the processor. The processor and storage medium can be located in an ASIC. In addition, the ASIC can be located in a computing device. Of course, the processor and storage medium can also exist as discrete components in a computing device.

[0222] In the above embodiments, all or part of the embodiments can be implemented by software, hardware, firmware or any combination thereof. When implemented using software, all or part of the embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer programs or instructions. When the computer program or instructions are loaded and executed on a computer, the process or function described in the embodiments of the present application is performed in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, a network device, a user device or other programmable device. The computer program or instruction can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer program or instruction can be transmitted from one website, computer, server or data center to another website, computer, server or data center via wired or wireless means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium, such as a floppy disk, a hard disk, or a tape; it can also be an optical medium, such as a digital video disc (DVD); it can also be a semiconductor medium, such as a solid state drive (SSD). The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present application, and such modifications or substitutions should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.

Claims

1. An artificial intelligence model deployment method, characterized in that, Including: Obtaining the number of parameters used when the artificial intelligence model to be deployed performs inference or training, and information on the available storage resources of the computing system on which the artificial intelligence model is deployed, where the information on the storage resources includes multiple types of storage resources, and each type of storage resource in the multiple types of storage resources has different access performances; Determining the storage resource for storing the parameters according to the number of parameters and the information on the available storage resources of the computing system; And When deploying the artificial intelligence model, storing the parameters in the determined storage resource.

2. The method according to claim 1, wherein The multiple types of storage resources include: a first storage medium associated with a dedicated processor in the computing system, a second storage medium associated with a general-purpose processor in the computing system, and a remote storage device connected to the server where the dedicated processor and the general-purpose processor are located through a network.

3. The method according to claim 2, wherein The information on the available storage resources of the computing system includes the remaining storage capacity; Determining the storage resource for storing the parameters according to the number of parameters and the information on the available storage resources of the computing system includes: When the remaining storage capacity of the first storage medium is greater than the number of parameters, determining that the determined storage resource includes the first storage medium.

4. The method according to claim 2 or 3, characterized in that, Determining the storage resource for storing the parameters according to the number of parameters and the information on the available storage resources of the computing system includes: When the remaining storage capacity of the first storage medium is less than the number of parameters, determining that the determined storage resource includes the second storage medium.

5. The method according to claim 4, characterized in that, Determining that the determined storage resource includes the second storage medium includes: When the remaining storage capacity of the second storage medium is greater than the number of parameters, determining that the determined storage resource includes the second storage medium.

6. The method according to claim 4 or 5, characterized in that, Determining that the determined storage resource includes the second storage medium includes: When the remaining storage capacity of the second storage medium is less than the number of parameters, determining that the determined storage resource includes the second storage medium and the remote storage device.

7. The method according to claim 2 or 3, characterized in that, Determining the storage resource for storing the parameters according to the number of parameters and the information on the available storage resources of the computing system includes: When the remaining storage capacity of the first storage medium is less than the number of parameters and the remaining storage capacity of the second storage medium is less than the number of parameters, determining that the determined storage resource includes the remote storage device.

8. The method according to claim 2 or 3, characterized in that, The parameters include multiple feature tables related to the artificial intelligence model; Determining the storage resource for storing the parameters according to the number of parameters and the information on the available storage resources of the computing system includes: When the remaining storage capacity of the first storage medium is less than the number of parameters, determining that the determined storage resource includes the first storage medium and the second storage medium; Storing the parameters in the determined storage resource includes: Storing a first number of feature tables in the first storage medium and storing a second number of feature tables in the second storage medium.

9. The method according to claim 2 or 3, characterized in that, The parameters include multiple feature tables related to the artificial intelligence model; Determine the storage resources for storing the parameters based on the information of the parameter quantity and the available storage resources of the computing system, including: When the remaining storage capacity of the first storage medium is less than the parameter quantity, determine that the determined storage resources include the first storage medium and the remote storage device; Store the parameters in the determined storage resources, including: Store a first quantity of feature tables in the first storage medium and a second quantity of feature tables in the remote storage device.

10. The method according to claim 1, characterized in that, The method further includes: Obtain the performance requirements of the user during the inference or training of the artificial intelligence model; The determining the storage resources for storing the parameters based on the information of the parameter quantity and the available storage resources of the computing system includes: Determine the storage resources for storing the parameters based on the information of the parameter quantity, the performance requirements, and the available storage resources of the computing system.

11. A computer system, characterized in that, The computer system includes a memory and a processor. The memory is used to store a set of computer instructions; when the processor executes the set of computer instructions, the processor is caused to execute the operation steps of the method described in any one of claims 1-10 above.

12. A computer-readable storage medium, characterized in that, Including: Computer software instructions; when the computer software instructions run in the processor, the processor is caused to execute the operation steps of the method described in any one of claims 1-10 above.

13. A computer program product, characterized in that, Including: When the computer program product runs on a computer, the computer is caused to execute the operation steps of the method described in any one of claims 1-10 above.

Citation Information

Patent Citations

  • Method and equipment for deploying machine learning model and computer program product

    CN111507476A

  • Neural network parameter deployment method, AI integrated chip and related device thereof

    CN115705301A

  • Dynamically determining tracks to prestage from storage to cache by training a machine learning module

    US20190391920A1

  • Utilizing machine learning to streamline telemetry processing of storage media

    US20210334253A1

  • Information infrastructure management method, management server of information infrastructure, and non-transitory computer-readable recording medium for information infrastructure management program

    US20230185620A1

Cited By

  • Inference method and device, inference cluster, storage medium and program product

    CN121480725A