Neural network model loading method, device and equipment
By grading the parameters of the neural network model and selecting appropriate levels based on client device information for transmission and loading, the problems of large model transmission delay and insufficient computing resources are solved, and normal operation under various equipment and network conditions is achieved.
Patent Information
- Application Number
- CN202410381682.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-12-19
- Filing Date
- 2024-03-30
- Publication Date
- 2025-06-20
AI Technical Summary
In the prior art, the number of parameters of the neural network model is huge, resulting in the model transmission delay and cannot operate normally in the event of poor edge networks and insufficient computing resources.
By grading the full parameters of the neural network model into multi-level parameters, and selecting appropriate parameter levels based on the client's device information for transmission and loading, we ensure that the transmission delay is small and conforming to the client's computing resource capabilities.
It effectively reduces the transmission delay and computing resource requirements of the neural network model on the client side, ensuring that the model can operate normally under various devices and network conditions.
Smart Images

Figure CN120181189A_ABST
Abstract
Description
[0001] This application claims the priority of a Chinese patent application with the application number 202311758255.X and the invention title "Method, apparatus, and device for loading neural network models" filed on December 19, 2023, the entire content of which is incorporated herein by reference. Technical Field
[0002] This application relates to the field of cloud technologies, and particularly to a method, apparatus, and device for loading neural network models. Background Art
[0003] With the development of Artificial Intelligence (AI) technology, AI has been widely applied in various industries such as automobiles, robots, and healthcare. In AI technology, various neural network models play a key role. Neural network models can achieve functions such as object recognition, object segmentation, and scene understanding.
[0004] However, currently, the number of parameters of neural network models is increasing, resulting in an increasing amount of data for neural network models. In the case of poor edge networks, the latency of transmitting neural network models from the cloud to the edge will be relatively large, and in the case of insufficient computing resources at the edge, neural network models with a large amount of data cannot run properly. Summary of the Invention
[0005] This application provides a method, apparatus, and device for loading neural network models, which can better meet the client's requirements for transmitting and loading neural network models. The corresponding technical solutions are as follows:
[0006] In a first aspect, a method for loading a neural network model is provided. The method includes:
[0007] The cloud platform classifies the full-scale parameters of the first neural network model into N levels of parameters and stores the N levels of parameters. The client sends a first acquisition request for the first neural network model to the cloud platform, where the first acquisition request carries the first device information of the client, and the first device information includes at least one of the first network information and the first computing resource information. The cloud platform selects M levels of parameters from the N levels of parameters according to the first device information, where M is a positive integer less than or equal to N. The cloud platform sends the M levels of parameters and the computational graph of the first neural network model to the client. The client loads the M levels of parameters and the computational graph of the first neural network model.
[0008] In the technical solution provided by this application, the full-scale parameters of the neural network model stored in the cloud platform are classified into multiple levels of parameters. These multiple levels of parameters are all sub-levels of the full-scale parameters of the neural network model, and there is no intersection between any two levels of these multiple levels of parameters. The cloud platform can return the corresponding level of parameters to the client according to the current network situation and computing resource situation of the client, which can avoid as much as possible the problem of large parameter transmission delay caused by the poor network condition of the client, and can also avoid as much as possible the problem that the client cannot normally load parameters due to less computing resources of the client.
[0009] In addition, by classifying and compressing the parameters of the neural network model and storing them hierarchically, it is possible to manage and update each level of parameters of the neural network model separately.
[0010] In a possible implementation, the method further includes:
[0011] The cloud platform determines the loading priority of each level of parameters in the N levels of parameters. Correspondingly, the cloud platform selects M levels of parameters from the N levels of parameters according to the first device information, including:
[0012] The cloud platform determines M loading priorities among the loading priorities corresponding to the N levels of parameters according to the first device information. The cloud platform selects M levels of parameters corresponding to the M loading priorities from the N levels of parameters.
[0013] In the technical solution provided by this application, each level of parameters of the neural network model can correspond to a loading priority. For example, the parameters of the neural network model are compressed twice to obtain three levels of parameters. Among them, the three levels of parameters are the remaining parameters after compression, the parameters compressed for the first time, and the parameters compressed for the second time. The loading priority of the remaining parameters after compression is the highest, the loading priority of the parameters compressed for the first time is the second, and the loading priority of the parameters compressed for the second time is the lowest. On this basis, the cloud platform can select parameters according to the loading priority corresponding to each level of parameters of the neural network model, and preferentially select parameters with a high loading priority.
[0014] In a possible implementation, the cloud platform determines M loading priorities among the loading priorities corresponding to the N levels of parameters according to the first device information, including:
[0015] The cloud platform determines M loading priorities corresponding to the first device information in the correspondence between device information and loading priorities.
[0016] In the technical solution provided by this application, the cloud platform can pre-establish a correspondence between device information and loading priorities. After receiving a request to obtain the first neural network model, it extracts the first device information carried therein. Furthermore, it can query the above correspondence to determine M loading priorities corresponding to the first device information.
[0017] In a possible implementation, the cloud platform determines M loading priorities from the loading priorities corresponding to the hierarchically stored parameters of the first neural network model according to at least one of the first network information and the first computing resource information, including:
[0018] The cloud platform inputs the first device information into the parameter selection model to obtain M loading priorities.
[0019] In the technical solution provided by this application, the cloud platform may deploy a parameter selection model, which is also a neural network model. After receiving a request for obtaining the first neural network model, the first device information carried therein is extracted. Furthermore, the cloud platform can input the first device information into the parameter selection model to obtain M loading priorities output by the parameter selection model.
[0020] In a possible implementation, the method further includes:
[0021] The cloud platform hierarchically compresses the parameters of the second neural network model into P-level parameters. If any one of the P-level parameters has the same digest as any one of the stored N-level parameters, the cloud platform uses the stored any one of the parameters and the parameters of the P-level parameters other than the any one of the parameters as the parameters after hierarchical classification of the second neural network model.
[0022] In the technical solution provided by this application, if at least one level of parameters is the same after hierarchical compression of the first neural network model and the second neural network model, then for this at least one level of parameters, only one copy can be stored in the file system and shared by the first neural network model and the second neural network model, which can effectively save storage space.
[0023] In a possible implementation, the method further includes:
[0024] The client sends a second request for obtaining the second neural network model to the cloud platform, where the second device information of the client is carried in the second request. The cloud platform selects X-level parameters from the P-level parameters of the second neural network model according to the second device information. If any one of the X-level parameters has the same digest as any one of the M-level parameters that the client has already obtained, the cloud platform sends the computation graph of the second neural network model and the parameters of the X-level parameters other than the any one of the parameters to the client.
[0025] In the technical solution provided by this application, in order to save the storage resources of the client, reduce the amount of data transmitted between the cloud platform and the client, and save the network resources of the client. When there is at least one level of parameters that are the same between the first neural network model and the second neural network model, and the terminal has obtained the Y-level parameters among this at least one level of parameters of the first neural network model, when the terminal requests the second neural network model from the cloud platform, the cloud platform may no longer send these Y-level parameters to the terminal.
[0026] In a possible implementation, the network information includes at least one of the available bandwidth and latency of the edge where the above client is installed.
[0027] In a possible implementation, the computing resource information includes at least one of the utilization rate of the CPU at the edge, the available memory at the edge, the utilization rate of the GPU at the edge, and the available video memory of the client.
[0028] In a second aspect, a method for loading a neural network model is provided. The method is applied to a cloud platform. The cloud platform stores the full amount of parameters of the first neural network model and the computational graph of the first neural network model. The method includes:
[0029] Classify the full amount of parameters of the first neural network model into N levels of parameters, where N is an integer greater than 1, and there is no intersection between any two levels of parameters among the N levels of parameters;
[0030] Store the N levels of parameters separately;
[0031] Receive a first acquisition request for the first neural network model sent by the client. Among them, the first device information of the client is carried in the first acquisition request, and the first device information includes at least one of the first network information and the first computing resource information;
[0032] According to the first device information, select M levels of parameters from the N levels of parameters, where M is a positive integer less than or equal to N;
[0033] Send the M levels of parameters and the computational graph of the first neural network model to the client.
[0034] In a possible implementation, the method further includes:
[0035] Determine the loading priority of each level of parameters in the N levels of parameters;
[0036] According to the first device information, select M levels of parameters from the N levels of parameters:
[0037] According to the first device information, determine M loading priorities among the loading priorities corresponding to the N levels of parameters respectively;
[0038] Among the N-level parameters, select the M-level parameters corresponding to the M loading priorities.
[0039] In a possible implementation, determining M loading priorities among the loading priorities corresponding to the hierarchically stored parameters of the first neural network model according to the first device information includes:
[0040] Determine the M loading priorities corresponding to the first device information in the correspondence between device information and loading priorities.
[0041] In a possible implementation, determining M loading priorities among the loading priorities corresponding to the hierarchically stored parameters of the first neural network model according to the first device information includes:
[0042] Input the first device information into a parameter selection model to obtain M loading priorities.
[0043] In a possible implementation, the method further includes:
[0044] Hierarchically compress the parameters of the second neural network model into P-level parameters, where P is an integer greater than 1;
[0045] If any one of the P-level parameters has the same digest as any one of the N-level parameters that have been stored, the cloud platform uses the stored any one of the N-level parameters and the parameters of the P-level parameters other than the any one of the P-level parameters as the parameters after hierarchical classification of the second neural network model.
[0046] In a possible implementation, the method further includes:
[0047] Receive a second acquisition request for the second neural network model sent by the client, where the second device information of the client is carried in the first acquisition request;
[0048] According to the second device information, select X-level parameters from the P-level parameters;
[0049] If any one of the X-level parameters has the same digest as any one of the M-level parameters that the client has already acquired, the cloud platform sends the computational graph of the second neural network model and the parameters of the X-level parameters other than the any one of the X-level parameters to the client.
[0050] In a possible implementation, the network information includes at least one of the available bandwidth and latency of the client edge where the above client is installed.
[0051] In a possible implementation, the computing resource information includes at least one of the utilization rate of the CPU at the edge, the available memory at the edge, the utilization rate of the GPU at the edge, and the available video memory of the client.
[0052] In a third aspect, a device for loading a neural network model is provided. The device is applied to a cloud platform, and the cloud platform stores the full parameters of a first neural network model and the computational graph of the first neural network model. The device includes:
[0053] A hierarchical compression module, configured to hierarchically compress the full parameters of the first neural network model into N levels of parameters, where N is an integer greater than 1, and there is no intersection between any two levels of parameters in the N levels of parameters;
[0054] A hierarchical storage module, configured to store the N levels of parameters respectively; receive a first acquisition request for the first neural network model sent by a client, where the first acquisition request carries first device information of the client, and the first device information includes at least one of first network information and first computing resource information; select M levels of parameters from the N levels of parameters according to the first device information, where M is a positive integer less than or equal to N; and send the M levels of parameters and the computational graph of the first neural network model to the client.
[0055] In a possible implementation, the hierarchical storage module is configured to:
[0056] Determine the loading priority of each level of parameters in the N levels of parameters;
[0057] Determine M loading priorities from the loading priorities corresponding to the N levels of parameters according to the first device information;
[0058] Select M levels of parameters corresponding to the M loading priorities from the N levels of parameters.
[0059] In a possible implementation, the hierarchical storage module is configured to:
[0060] Determine M loading priorities corresponding to the first device information in the correspondence between device information and loading priorities.
[0061] In a possible implementation, the hierarchical storage module is configured to:
[0062] Input the first device information into a parameter selection model to obtain M loading priorities.
[0063] In a possible implementation, the hierarchical compression module is further configured to:
[0064] The parameters of the second neural network model are hierarchically compressed into P-level parameters, where P is an integer greater than 1;
[0065] The hierarchical storage module is further configured to:
[0066] If any one of the P-level parameters has the same digest as any one of the stored N-level parameters, the cloud platform uses the stored any one of the N-level parameters and the parameters of the P-level parameters other than the any one of the P-level parameters as the parameters after hierarchical classification of the second neural network model.
[0067] In a possible implementation, the hierarchical storage module is further configured to:
[0068] Receive a second acquisition request for the second neural network model sent by the client, where the second device information of the client is carried in the first acquisition request;
[0069] According to the second device information, select X-level parameters from the P-level parameters;
[0070] If any one of the X-level parameters has the same digest as any one of the M-level parameters already acquired by the client, the cloud platform sends the computational graph of the second neural network model and the parameters of the X-level parameters other than the any one of the X-level parameters to the client.
[0071] In a possible implementation, the network information includes at least one of the available bandwidth and latency of the edge where the above client is installed.
[0072] In a possible implementation, the computing resource information includes at least one of the utilization rate of the CPU of the edge, the available memory of the edge, the utilization rate of the GPU of the edge, and the available video memory of the client.
[0073] Fourthly, a computing device cluster is provided. The computing device cluster includes at least one computing device, and each computing device includes a processor and a memory;
[0074] The processor of the at least one computing device is configured to execute instructions stored in the memory of the device, so that the computing device cluster executes the method for loading a neural network model as described in the second aspect above.
[0075] Fifthly, a computer program product containing instructions is provided. When the instructions are run on a computing device cluster, the computing device cluster is made to execute the method for loading a neural network model as described in the second aspect above.
[0076] In a sixth aspect, there is provided a computer-readable storage medium including computer program instructions which, when executed by a computing device, cause the computing device to execute the method for loading a neural network model as described in the second aspect above. BRIEF DESCRIPTION OF THE DRAWINGS
[0077] Figure 1 FIG. is a schematic diagram of a system for loading a neural network model provided by an embodiment of the present application;
[0078] Figure 2 FIG. is a schematic flowchart of a method for loading a neural network model provided by an embodiment of the present application;
[0079] Figure 3 FIG. is a schematic diagram of hierarchical storage of a neural network model provided by an embodiment of the present application;
[0080] Figure 4 FIG. is a schematic flowchart of a method for loading a neural network model provided by an embodiment of the present application;
[0081] Figure 5 FIG. is a schematic diagram of loading a neural network model provided by an embodiment of the present application;
[0082] Figure 6 FIG. is a schematic structural diagram of a device for loading a neural network model provided by an embodiment of the present application;
[0083] Figure 7 FIG. is a schematic structural diagram of a computing device provided by an embodiment of the present application;
[0084] Figure 8 FIG. is a schematic structural diagram of a computing device cluster provided by an embodiment of the present application;
[0085] Figure 9 FIG. is a schematic structural diagram of a computing device cluster provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0086] For ease of understanding of the embodiments of the present application, the following first explains the terms involved in the embodiments of the present application.
[0087] Parameters of the neural network model:
[0088] The parameters in the neural network model refer to the variables that can be learned in the neural network model, including weights and biases.
[0089] Weights:
[0090] Weights characterize the strength of the connections between neurons in the neural network, and they determine the degree of influence when the input signal is transmitted between neurons.
[0091] Bias:
[0092] The bias is the offset or threshold of a neuron, which allows the neuron to activate even without input.
[0093] Neuron:
[0094] The basic unit of a neural network that contains a single data feature.
[0095] Computation graph:
[0096] It includes the connection relationships of each neuron and records information such as the operation type, input, output, and attributes of each neuron.
[0097] The following describes the method for loading a neural network model provided by the embodiments of the present application with reference to the accompanying drawings.
[0098] The embodiments of the present application provide a method for loading a neural network model. This method can be implemented by a cloud platform and a client. The cloud platform can be a computing device or a cluster of computing devices. The client can be software installed at the edge for interacting with the cloud services provided by the cloud platform. The edge can be devices such as mobile phones, computers, robots, servers, etc.
[0099] See Figure 1 , which shows the system architecture diagram of a method for loading a neural network model provided by the embodiments of the present application. As Figure 1 shown, the cloud platform includes underlying hardware and cloud services. Among them, the underlying hardware includes components such as processors and memories. The processor can run the cloud services, and the memory can store the full-scale parameters of the neural network model and the computation graph of the neural network model. The edge also includes underlying hardware and a client. Among them, the underlying hardware includes components such as processors and memories. The processor can run the client, and the client is used to interact with the cloud services provided by the cloud platform. The memory can store the parameters and computation graph of the neural network model obtained by the client from the cloud platform.
[0100] A neural network model can be composed of a computation graph and parameters. For the obtained neural network model, the cloud platform can perform hierarchical compression on its full-scale parameters. For example, Figure 1The total number of parameters of the neural network model 1 is 1,100. The parameters of the neural network model 1 are hierarchically compressed into three levels of parameters. The first-level parameters include 600 parameters, the second-level parameters include 300 parameters, and the third-level parameters include 200 parameters. Among them, these three levels of parameters are all subsets of the parameters of the neural network model 1, and there is no intersection between any two levels of these three levels of parameters. The total number of parameters of the neural network model 2 is 700. The parameters of the neural network model 2 are hierarchically compressed into two levels of parameters. The first-level parameters include 500 parameters, and the second-level parameters include 200 parameters. Among them, these two levels of parameters are all subsets of the parameters of the neural network model 2. When the client needs to use the neural network model, it can send a request to obtain the neural network model to the cloud platform. For example, if the client needs to use the neural network model 1, it can send a request to obtain the neural network model 1 to the cloud platform, and carry the network information and computing resource information of the client in the request. After receiving the request to obtain the neural network model 1, the cloud platform selects at least one level of parameters that meet the network requirements and computing resource requirements from the three levels of parameters of the neural network model 1 according to the network information and computing resource information of the client, such as selecting the first two levels of parameters. Then, the selected first two levels of parameters and the computational graph of the neural network model 1 are returned to the client, and the client loads the first two levels of parameters and the computational graph of the neural network model 1. In this way, the cloud platform can return the corresponding level of parameters to the client according to the current network situation and computing resource situation of the client, which can avoid the problem of large parameter transmission delay caused by the poor network condition of the client as much as possible, and can also avoid the problem that the client cannot normally load the parameters due to less computing resources of the client.
[0101] See Figure 2 shows a flowchart of a method for loading a neural network model provided by an embodiment of the present application. As Figure 2 shown, the processing of this method may include the following steps:
[0102] Step 201, the platform hierarchically classifies the parameters of the first neural network model into N levels of parameters.
[0103] Among them, the first neural network model is any neural network model obtained by the cloud platform, and N is an integer greater than 1.
[0104] In implementation, the cloud platform can obtain the neural network model through various means such as user upload, in-house research and development by manufacturers, and public resources. For the obtained first neural network model, the cloud platform can hierarchically compress the parameters of the first neural network model into N-level parameters. Here, the hierarchical compression can be implemented by compression algorithms such as pruning, quantization, and distillation. The specific number of compression levels can be configured by relevant personnel considering factors such as the total amount of parameters of the neural network model and storage resources, or the number of compression levels can be automatically determined by the compression algorithm. The embodiments of the present application do not limit the compression method and the number of compression levels of the neural network model.
[0105] When the cloud platform hierarchically compresses the parameters of the first neural network model, the remaining parameters after compression are taken as one level of parameters alone, and the parameters compressed at each level are taken as one level of parameters respectively. Among them, in the first neural network model, the remaining parameters after compression can be used alone to complete inference, but the inference effect may be worse than that of using all the parameters.
[0106] For example, as Figure 1 in the neural network model 1, the total number of parameters of the neural network model 1 is 1100. The cloud platform hierarchically compresses the parameters of the neural network model 1 into three-level parameters, that is, the neural network model 1 is compressed twice. The first time, 300 parameters are compressed off, and the second time, 200 parameters are compressed off. After compression, 600 parameters remain. Then, the remaining 600 parameters after compression are taken as the first-level parameters, the 300 parameters compressed off for the first time are taken as the second-level parameters, and the 200 parameters compressed off for the second time are taken as the third-level parameters. Among them, in the neural network model 1, the remaining 600 parameters after compression can be used alone to complete inference, but the inference effect may be worse than that of using all 11000 parameters. In the neural network model 1, the remaining 600 parameters after compression and the 300 parameters compressed off for the first time can be used for inference, and the inference effect is better than that of using only the remaining parameters after compression. In the neural network model 1, the remaining 600 parameters after compression, the 300 parameters compressed off for the first time, and the 200 parameters compressed off for the second time are used for inference, that is, using all 1100 parameters for inference, which is better than the inference effect of using the remaining 600 parameters after compression and the 300 parameters compressed off for the first time.
[0107] In a possible implementation, when hierarchically compressing the parameters of the first neural network model, metadata for each level of parameters can also be generated. Among them, the metadata for each level of parameters can include the parameter identifier (layer id) of this level, the loading priority (priority), the data volume (size), the parameter identifier of the next level (next layer id), the globally unique identifier of the parameters of this level, the computing power required for loading, and so on. The following explains each item in the metadata.
[0108] Parameter identifier at this level: The layer id of each level of parameters is unique among the parameters of each level of the first neural network model. For example, taking the above neural network model 1 as an example, the layer ids of the three levels of parameters of neural network model 1 can be layer1, layer2, and layer3 respectively.
[0109] Loading priority: Used to indicate the priority when this level of parameters is loaded by the client. Among them, the parameters of each level are in the order of decreasing loading priority as follows: the remaining parameters after compression, the parameters compressed for the first time, the parameters compressed for the second time, and so on. For example, taking the above neural network model 1 as an example, the 600 remaining parameters after compression have the highest loading priority, and the loading priority is denoted as 1. The 300 parameters compressed for the first time have the second highest loading priority, and the loading priority is denoted as 2. The 200 parameters compressed for the second time have the lowest loading priority, and the lowest loading priority is denoted as 3.
[0110] Data volume: The data volume of the parameters at this level.
[0111] Next-level parameter identifier: The remaining parameters after compression are the first-level parameters, the parameters compressed for the first time are the second-level parameters, the parameters compressed for the second time are the third-level parameters, and so on. The next-level parameter identifier of the remaining parameters after compression is the parameter identifier at this level of the parameters compressed for the first time, and the next-level parameter identifier of the parameters compressed for the first time is the parameter identifier at this level of the parameters compressed for the second time. For example, taking the above neural network model 1 as an example, the next-level parameter identifier of the first-level parameters of neural network model 1 is layer2, and the next-level parameter identifier of the second-level parameters is layer3. The third-level parameters have no next-level parameters, so their next-level parameter identifier can be empty.
[0112] Computing power required for loading: The computing power required for the client to load and run the parameters at this level. The computing power here can be the computing power of the central processing unit (CPU), the computing power of the graphics processing unit (GPU), etc.
[0113] Global unique identifier of the parameters at this level: The global unique identifier refers to the unique identifier of the parameters at this level in the file system storing the parameters of the neural network model, which can be the digest of the parameters at this level. The digest algorithm can be a hash algorithm, a hashing algorithm, etc. If the global unique identifiers of two levels of parameters are the same, it means that these two levels of parameters are the same.
[0114] Step 202: The cloud platform stores the N-level parameters.
[0115] In implementation, the cloud platform can store the N-level parameters of the first neural network model through a unified file system respectively.
[0116] In a possible implementation, each level of parameter has a corresponding globally unique identifier. Before storing each level of parameter in the N-level parameters of the first neural network model, it is possible to first query whether there is a parameter corresponding to the target globally unique identifier in the file system according to the target globally unique identifier corresponding to this level of parameter. If it is queried that there is a parameter corresponding to the target globally unique identifier, then this level of parameter is not stored, and the stored parameter corresponding to the target globally unique identifier is used as this level of parameter of the first neural network model. If the parameter corresponding to the target globally unique identifier is not queried, then this level of parameter is stored in the file system.
[0117] For example, referring to Figure 3 , the neural network model 3 is hierarchically compressed into two levels of parameters. Among them, the globally unique identifier of the first level of parameter is digest1, and the globally unique identifier of the second level of parameter is digest2. The two levels of parameters of the neural network model 3 have been stored in the file system respectively. After hierarchically compressing the neural network model 4, three levels of parameters are obtained. The globally unique identifier of the first level of parameter is digest1, the globally unique identifier of the second level of parameter is digest3, and the globally unique identifier of the third level of parameter is digest4. When storing the first level of parameter of the neural network model 4, it is queried that there is a parameter with the globally unique identifier of digest1 stored in the file system, then the first level of parameter of the neural network model 4 is not stored, and the neural network model 4 and the neural network model 3 share the first level of parameter. When storing the second level of parameter of the neural network model 4, it is queried that there is no parameter with the globally unique identifier of digest3 stored in the file system, then the second level of parameter of the neural network model 4 is stored in the file system. When storing the third level of parameter of the neural network model 4, it is queried that there is no parameter with the globally unique identifier of digest4 stored in the file system, then the third level of parameter of the neural network model 4 is stored in the file system.
[0118] Step 203, the client sends a first acquisition request for the first neural network model to the cloud platform.
[0119] Among them, the first acquisition request carries the first device information of the client, and the first device information includes at least one of the first network information and the first computing resource information.
[0120] In implementation, the client can obtain the first neural network model from the cloud platform according to actual needs. Specifically, the client can send a first acquisition request for the first neural network model to the cloud platform, wherein the first acquisition request carries the first device information of the client, and the first device information includes at least one of the first network information and the first computing resource information. The first network information may include the client's current available bandwidth, latency, number of network connections, etc., and the first computing resource information may include the client's current CPU utilization, available memory, GPU utilization, GPU remaining computing power, available video memory, etc. In addition, the first device information may also include information such as device temperature and device operation log to describe the device's operating status.
[0121] Step 204: The cloud platform selects M-level parameters from the N-level parameters of the first neural network model according to the first device information.
[0122] Wherein, M is a positive integer less than or equal to N.
[0123] In implementation, after receiving the first acquisition request for the first neural network model sent by the client, the cloud platform selects M-level parameters from the N-level parameters of the first neural network model according to the first device information carried in the first acquisition request. There are many selection methods, and several of them are listed below for illustration.
[0124] Method 1:
[0125] For each neural network model, a correspondence between device information and loading priority can be established in advance. After receiving the first acquisition request for the first neural network model, the cloud platform queries the correspondence and determines the M loading priorities corresponding to the first device information, where each loading priority corresponds to a first-level parameter. Then, the M-level parameters corresponding to the M loading priorities can be selected.
[0126] Taking the above neural network model 1 as an example, the correspondence between device information and loading priority can be shown in Table 1 below.
[0127] Table 1
[0128]
[0129]
[0130] Among them, C1, C2, and C3 are the thresholds of CPU utilization, R1, R2, and R3 are the thresholds of available memory, G1, G2, and G3 are the thresholds of GPU utilization, T1, T2, and T3 are the thresholds of GPU remaining computing power, X1, X2, and X3 are the thresholds of available video memory, B1, B2, and B3 are the thresholds of available bandwidth, and D1, D2, and D3 are the thresholds of latency.
[0131] After the cloud platform receives the acquisition request for neural network model 1 sent by the terminal, it extracts the device information carried in the acquisition request, including the CPU utilization rate C, available memory R, GPU utilization rate G, remaining computing power T of the GPU, available video memory X, available loan B, and latency D of the terminal. Furthermore, it queries Table 1 to determine that the loading priorities corresponding to the device information are 1 and 2, and then selects the parameters with a loading priority of 1 and the parameters with a loading priority of 2 for neural network model 1.
[0132] Method 2:
[0133] The cloud platform can deploy a pre-trained parameter selection model, which can be a neural network model. After the cloud platform receives the first acquisition request for the first neural network model, it extracts the first device information carried in the first acquisition request. Then, it inputs the first device information, the data volume of each level of parameters of the first neural network model, and the computing power required for loading each level of parameters of the first neural network model into the parameter selection model, and the parameter selection model outputs M loading priorities. Furthermore, the neural network model obtains the M levels of parameters corresponding to the M loading priorities respectively.
[0134] Step 205: The cloud platform sends the above M levels of parameters and the computational graph of the first neural network model to the client.
[0135] In implementation, the cloud platform can read the M levels of parameters of the selected first neural network model in the file system and send the M levels of parameters of the first neural network model and the computational graph of the first neural network model to the terminal.
[0136] Step 206: The client loads the above M levels of parameters and the computational graph of the first neural network model.
[0137] In implementation, after the client receives the M levels of parameters of the first neural network model and the computational graph of the first neural network model sent by the cloud platform, it stores the M levels of parameters of the first neural network model and the computational graph of the first neural network model, and after receiving the loading and running instruction of the first neural model, it loads the M levels of parameters of the first neural network model and the computational graph of the first neural network model.
[0138] In a possible implementation, to save the storage resources of the client and reduce the amount of transmitted data between the cloud platform and the client, when there is at least one level of parameters that are the same between the first neural network model and the second neural model, and the terminal has obtained Y levels of parameters among these at least one level of parameters of the first neural network model, when the terminal requests the second neural network model from the cloud platform, the cloud platform can no longer send these Y levels of parameters to the terminal. The following describes how the client obtains the second neural network model. As Figure 4 shown, the processing can include the following steps:
[0139] Step 401: The client sends a second acquisition request for the second neural network model to the cloud platform.
[0140] Among them, the first acquisition request carries the second device information of the client and the first globally unique identifiers respectively corresponding to the parameters at all levels of the neural network model stored locally.
[0141] In implementation, when the client acquires the neural network model, in addition to sending the parameters of the neural network model and the computational graph of the neural network model to the terminal, the cloud platform can also send the first globally unique identifiers respectively corresponding to these parameters to the terminal.
[0142] Correspondingly, when the client sends a second acquisition request for the second neural network model to the cloud platform, in addition to sending the second device information of the client, it can also send the globally unique identifiers respectively corresponding to the parameters at all levels of the neural network model stored locally.
[0143] Step 402: The cloud platform selects X-level parameters from the P-level parameters of the second neural network model according to the second device information.
[0144] The specific processing of this step 402 is the same as the specific processing of the above step 204, and will not be elaborated here.
[0145] Step 403: If the X-level parameters include the Y-level parameters in the W-level parameters stored locally by the client, the cloud platform sends the computational graph of the second neural network model and the X - Y-level parameters other than the Y-level parameters in the X parameters to the client.
[0146] In implementation, when the cloud platform selects the X-level parameters, it can also obtain the globally unique identifiers respectively corresponding to these X-level parameters. If Y of the globally unique identifiers respectively corresponding to these X-level parameters are the same as the Y globally unique identifiers corresponding to the Y-level parameters of the neural network model stored locally by the client, it means that the parameters corresponding to these Y globally unique identifiers of the second neural network are the same as the Y-level parameters of the neural network model stored locally by the client. In this case, the cloud platform does not need to send these Y-level parameters to the client anymore. Then the cloud platform can send the computational graph of the second neural network model and the X - Y-level parameters other than the above Y-level parameters in the X parameters to the client. In addition, the cloud platform can also send the above X globally unique identifiers to the client, so that the client can find the above Y-level parameters that the cloud platform did not send to the client locally according to the X globally unique identifiers.
[0147] For example, as Figure 5As shown, the neural network model 3 is hierarchically compressed into two levels of parameters. Among them, the globally unique identifier of the first-level parameters is digest1, and the globally unique identifier of the second-level parameters is digest2. The neural network model 4 is hierarchically compressed into three levels of parameters. The globally unique identifier of the first-level parameters is digest1, the globally unique identifier of the second-level parameters is digest3, and the globally unique identifier of the third-level parameters is digest4. The client has obtained the neural network model 3, and the first-level parameters of the neural network model 3 are stored locally on the client. When the client sends a request to obtain the neural network model 4 to the cloud platform, the globally unique identifier digest1 of the first-level parameters of the neural network model 3 and the device information of the client are carried in the obtain request. The cloud platform selects the first-level parameters and the second-level parameters of the neural network model 4 according to the device information of the client. Then, it is determined that the globally unique identifier of the first-level parameters of the neural network model 4 is also digest1. Then, the cloud platform does not need to send the first-level parameters with the globally unique identifier digest1 of the neural network model 4 to the client, and only needs to send the second-level parameters with the globally unique identifier digest3 of the neural network model 4 and the computational graph of the neural network model 4.
[0148] In this way, the client does not need to repeatedly store the second-level parameters of the neural network model 4, saving the computing resources of the client. The cloud platform also does not need to send the second-level parameters of the neural network model 4 to the client, reducing the amount of transmitted data, improving the transmission efficiency, and reducing the network resource occupancy.
[0149] The embodiment of the present application also provides a device for loading a neural network model. This device can be a cloud platform, such as Figure 6 As shown, this device can include a hierarchical compression module 410 and a hierarchical storage module 420, where:
[0150] The hierarchical compression module 410 is used to hierarchically classify the full amount of parameters of the first neural network model into N levels of parameters, where N is an integer greater than 1, and there is no intersection between any two levels of parameters in the N levels of parameters;
[0151] The hierarchical storage module 420 is used to store the N levels of parameters; receive a first obtain request for the first neural network model sent by the client, where the first obtain request carries the first device information of the client, and the first device information includes at least one of the first network information and the first computing resource information; according to the first device information, select M levels of parameters from the N levels of parameters, where M is a positive integer less than or equal to N; send the M levels of parameters and the computational graph of the first neural network model to the client.
[0152] In a possible implementation, the hierarchical storage module 420 is used for:
[0153] Determine the loading priority of each level of parameters among the N-level parameters;
[0154] According to the first device information, determine M loading priorities among the loading priorities corresponding to the N-level parameters respectively;
[0155] Among the N-level parameters, select the M-level parameters corresponding to the M loading priorities.
[0156] In a possible implementation, the hierarchical storage module 420 is used for:
[0157] In the correspondence relationship between device information and loading priorities, determine the M loading priorities corresponding to the first device information.
[0158] In a possible implementation, the hierarchical storage module 420 is used for:
[0159] Input the first device information into the parameter selection model to obtain M loading priorities.
[0160] In a possible implementation, the hierarchical compression module 410 is further used for:
[0161] Hierarchically compress the parameters of the second neural network model into P-level parameters, where P is an integer greater than 1;
[0162] The hierarchical storage module 420 is further used for:
[0163] If any level of parameter in the P-level parameters has the same digest as any level of parameter in the stored N-level parameters, the cloud platform will use the stored any level of parameter and the parameters in the P-level parameters other than the any level of parameter together as the parameters after hierarchical classification of the second neural network model.
[0164] In a possible implementation, the hierarchical storage module 420 is further used for:
[0165] Receive a second acquisition request for the second neural network model sent by the client, where the first acquisition request carries the second device information of the client;
[0166] According to the second device information, select X-level parameters from the P-level parameters;
[0167] If any level of parameter in the X-level parameters has the same digest as any level of parameter in the M-level parameters that the client has already acquired, the cloud platform sends the computational graph of the second neural network model and the parameters in the X-level parameters other than the any level of parameter to the client.
[0168] In a possible implementation, the network information includes at least one of the available bandwidth and latency for installing the above-mentioned client edge.
[0169] In a possible implementation, the computing resource information includes at least one of the utilization rate of the CPU of the edge, the available memory of the edge, the utilization rate of the GPU of the edge, and the available video memory of the client.
[0170] Both the above-mentioned hierarchical compression module 410 and the hierarchical storage module 420 can be implemented by software or by hardware. Exemplarily, the implementation manner of the hierarchical compression module 410 will be introduced next. Similarly, the implementation manner of the hierarchical storage module 420 can refer to the implementation manner of the hierarchical compression module 410.
[0171] As an example of a software functional unit, the hierarchical compression module 410 may include code running on a computing instance. Among them, the computing instance can be at least one of computing devices such as a physical host (computing device), a virtual machine, and a container. Further, the above-mentioned computing device can be one or more. For example, the hierarchical compression module 410 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers for running this application can be distributed in the same region or in different regions. The multiple hosts / virtual machines / containers for running this code can be distributed in the same AZ or in different AZs, and each AZ includes one data center or multiple geographically proximate data centers. Among them, generally one region can include multiple AZs.
[0172] Similarly, the multiple hosts / virtual machines / containers for running this code can be distributed in the same VPC or in multiple VPCs. Among them, generally one VPC is set within one region. For cross-region communication between two VPCs within the same region and between VPCs in different regions, a communication gateway needs to be set in each VPC, and the interconnection between VPCs is achieved through the communication gateway.
[0173] As an example of a hardware functional unit, the hierarchical compression module 410 may include at least one computing device, such as a server, etc. Or, the hierarchical compression module 410 can also be a device implemented using ASIC or PLD, etc. Among them, the above-mentioned PLD can be implemented by CPLD, FPGA, GAL or any combination thereof.
[0174] The multiple computing devices included in the hierarchical compression module 410 can be distributed in the same region or in different regions. The multiple computing devices included in the hierarchical compression module 4100 can be distributed in the same availability zone (AZ) or in different AZs. Similarly, the multiple computing devices included in the management module 410 can be distributed in the same virtual private cloud (VPC) or in multiple VPCs. Among them, the multiple computing devices can be any combination of computing devices such as servers, application-specific integrated circuits (ASICs), programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), and generic array logic (GALs).
[0175] This application also provides a computing device 100. As Figure 7 shown, the computing device 100 includes: a bus 102, a processor 104, a memory 106, and a communication interface 108. The processor 104, the memory 106, and the communication interface 108 communicate with each other through the bus 102. The computing device 100 can be a server or a terminal device. It should be understood that this application does not limit the number of processors and memories in the computing device 100.
[0176] The bus 102 can be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience in representation, Figure 7 only one line is shown herein, but it does not mean that there is only one bus or one type of bus. The bus 102 can include a path for transmitting information between various components of the computing device 100 (for example, the memory 106, the processor 104, and the communication interface 108).
[0177] The processor 104 can include any one or more of processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).
[0178] The memory 106 may include volatile memory, such as random access memory (RAM). The memory 106 may also include non-volatile memory, such as read-only memory (ROM), flash memory, a hard disk drive (HDD), or a solid state drive (SSD).
[0179] Executable code is stored in the memory 106, and the processor 104 executes the executable code to implement the functions of the foregoing hierarchical compression module 410 and hierarchical storage module 420 respectively, so as to implement the method for loading a neural network model. That is, instructions for executing the method for loading a neural network model are stored on the memory 106.
[0180] The communication interface 108 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 100 and other devices or a communication network.
[0181] An embodiment of this application also provides a computing device cluster. The computing device cluster includes at least one computing device. The computing device may be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device may also be a terminal device such as a desktop computer, a laptop computer, or a smart phone.
[0182] As Figure 8 shown, the computing device cluster includes at least one computing device 100. Instructions for executing the method for loading a neural network model may be stored in the memory 106 of one or more of the computing devices 100 in the computing device cluster.
[0183] In some possible implementation manners, partial instructions for executing the method for loading a neural network model may also be stored in the memory 106 of one or more of the computing devices 100 in the computing device cluster. In other words, a combination of one or more computing devices 100 may jointly execute the instructions for executing the method for loading a neural network model.
[0184] It should be noted that the memories 106 in different computing devices 100 in the computing device cluster may store different instructions for respectively executing partial functions of the method for loading a neural network model.
[0185] In some possible implementations, the memory 106 of one or more computing devices 100 in the computing device cluster may also store some instructions for executing the method of loading a neural network model respectively. In other words, the combination of one or more computing devices 100 can jointly execute the instructions for executing the method of loading a neural network model.
[0186] It should be noted that the memories 106 in different computing devices 100 in the computing device cluster may store different instructions for implementing partial functions of the method of loading a neural network model. That is, the instructions stored in the memories 106 of different computing devices 100 can implement the functions of one or more nodes in the hierarchical compression module 410 and the hierarchical storage module 420.
[0187] In some possible implementations, one or more computing devices in the computing device cluster may be connected through a network. Among them, the network may be a wide area network or a local area network, etc. Figure 9 A possible implementation is shown. As Figure 9 shown, two computing devices 100A and 100B are connected through a network. Specifically, they are connected to the network through the communication interfaces in each computing device. In this type of possible implementation, the memory 106 in the computing device 100A stores instructions for implementing the function of the hierarchical compression module 410. At the same time, the memory 106 in the computing device 100B stores instructions for implementing the function of the hierarchical storage module 420. It should be understood that Figure 9 the function of the computing device 100A shown in
[0188] can also be completed by multiple computing devices 100. Similarly, the function of the computing device 100B can also be completed by multiple computing devices 100. Figure 8 and Figure 9 the connection manner of the described computing device cluster. The difference is that the memories 106 of one or more computing devices 100 in this computing device cluster may store the same instructions for executing the method of loading a neural network model.
[0189] In some possible implementations, the memories 106 of one or more computing devices 100 in the computing device cluster may also store some instructions for executing the method of loading a neural network model respectively. In other words, the combination of one or more computing devices 100 can jointly execute the instructions for executing the method of loading a neural network model.
[0190] It should be noted that the memories 106 in different computing devices 100 in the computing device cluster can store different instructions for performing some functions of the apparatus for loading a neural network model. That is, the instructions stored in the memories 106 in different computing devices 100 can implement the functions of one or more nodes in the hierarchical compression module 410 and the hierarchical storage module 420.
[0191] An embodiment of the present application also provides a computer program product including instructions. The computer program product may be a software or program product including instructions that can run on a computing device or be stored in any available medium. When the computer program product runs on a computing device, at least one computing device is caused to execute the method for loading a neural network model provided by the embodiment of the present application.
[0192] An embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium may be any available medium that can be stored by a computing device or a data storage device such as a data center including one or more available media. The available medium may be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid state drive), etc. The computer-readable storage medium includes instructions that instruct a computing device to execute the method for loading a neural network model provided by the embodiment of the present application.
[0193] In the present application, terms such as "first" and "second" are used to distinguish identical or similar items with basically the same functions. It should be understood that there is no logical or temporal dependency between "first" and "second", nor are the quantity and execution order limited. It should also be understood that although the following description uses terms "first" and "second" to describe various elements, these elements should not be limited by the terms. These terms are only used to distinguish one element from another. For example, without departing from the scope of various examples, the first neural network model may be referred to as the second neural network model, and similarly, the second neural network model may be referred to as the first neural network model. The first neural network model and the second neural network model can both be collectively referred to as the neural network model, and in some cases, may be separate and different feature vectors.
[0194] In the present application, the meaning of the term "at least one" refers to one or more, and the meaning of the term "a plurality of" refers to two or more.
[0195] The above description is only a specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.
Claims
1. A method for loading a neural network model, characterized in that: The method is applied to a cloud platform and a client, wherein the cloud platform stores all parameters and a calculation graph of a first neural network model, and the method includes: The cloud platform classifies all parameters of the first neural network model into N-level parameters, where N is an integer greater than 1, and there is no intersection between any two levels of parameters in the N-level parameters; The cloud platform stores the N-level parameters; The client sends a first acquisition request for the first neural network model to the cloud platform, wherein the first acquisition request carries first device information of the client, and the first device information includes at least one of first network information and first computing resource information; The cloud platform selects M-level parameters from the N-level parameters according to the first device information, where M is a positive integer less than or equal to N; The cloud platform sends the M-level parameters and the calculation graph to the client; The client loads the M-level parameters and the calculation graph.
2. The method according to claim 1, characterized in that The method further comprises: The cloud platform determines the loading priority of each level of parameters in the N levels of parameters; The cloud platform selects M-level parameters from the N-level parameters according to the first device information, including: The cloud platform determines, according to the first device information, M loading priorities among the loading priorities respectively corresponding to the N level parameters; The cloud platform selects M-level parameters corresponding to the M loading priorities from the N-level parameters.
3. The method according to claim 2, characterized in that The cloud platform determines M loading priorities among the loading priorities respectively corresponding to the N level parameters according to the first device information, including: The cloud platform determines M loading priorities corresponding to the first device information in the corresponding relationship between device information and loading priorities.
4. The method according to claim 2, characterized in that: The cloud platform determines M loading priorities among the loading priorities respectively corresponding to the N level parameters according to the first device information, including: The cloud platform inputs the first device information into a parameter selection model to obtain M loading priorities.
5. The method according to any one of claims 1 to 4, characterized in that The method further comprises: The cloud platform classifies the parameters of the second neural network model into P-level parameters, wherein P is an integer greater than 1; If any level parameter in the P-level parameters has the same summary as any level parameter in the stored N-level parameters, the cloud platform will use the stored any level parameter and the parameters in the P-level parameters except the any level parameter as the graded parameters of the second neural network model.
6. The method according to claim 5, characterized in that The method further comprises: The client sends a second acquisition request for the second neural network model to the cloud platform, wherein the second acquisition request carries second device information of the client; The cloud platform selects X-level parameters from the P-level parameters according to the second device information; If any level parameter in the X-level parameters and any level parameter in the M-level parameters already acquired by the client have the same summary, the cloud platform sends the computation graph of the second neural network model and the parameters in the X-level parameters except the any level parameter to the client.
7. The method according to any one of claims 1 to 6, characterized in that The first network information includes at least one of the available bandwidth and latency of the edge where the client is installed; or, the first computing resource information includes at least one of the utilization of the central processing unit CPU of the edge, the available memory, the utilization of the graphics processing unit GPU, and the available video memory.
8. A method for loading a neural network model, characterized in that: The method is applied to a cloud platform, the cloud platform stores all parameters of a first neural network model and a calculation graph of the first neural network model, and the method includes: Classifying all parameters of the first neural network model into N-level parameters, where N is an integer greater than 1, and there is no intersection between any two levels of parameters in the N-level parameters; Storing the N-level parameters; Receiving a first acquisition request for the first neural network model sent by a client, wherein the first acquisition request carries first device information of the client, and the first device information includes at least one of first network information and first computing resource information; According to the first device information, select M-level parameters from the N-level parameters, where M is a positive integer less than or equal to N; Send the M-level parameters and a computational graph of the first neural network model to the client.
9. The method according to claim 8, characterized in that The method further comprises: Determining the loading priority of each level of parameters in the N levels of parameters; Selecting M-level parameters from the N-level parameters according to the first device information includes: Determine, according to the first device information, M loading priorities among the loading priorities respectively corresponding to the N level parameters; Among the N levels of parameters, select M levels of parameters corresponding to the M loading priorities.
10. The method according to claim 9, characterized in that The step of determining, according to the first device information, M loading priorities from the loading priorities corresponding to the hierarchically stored parameters of the first neural network model comprises: In the correspondence between device information and loading priorities, M loading priorities corresponding to the first device information are determined.
11. The method according to claim 9, characterized in that The step of determining, according to the first device information, M loading priorities from the loading priorities corresponding to the hierarchically stored parameters of the first neural network model comprises: The first device information is input into a parameter selection model to obtain M loading priorities.
12. The method according to any one of claims 8 to 11, characterized in that The method further comprises: The parameters of the second neural network model are hierarchically compressed into P-level parameters, wherein P is an integer greater than 1; If any level parameter in the P-level parameters has the same summary as any level parameter in the stored N-level parameters, the cloud platform will use the stored any level parameter and the parameters in the P-level parameters except the any level parameter as the graded parameters of the second neural network model.
13. The method according to any one of claims 8 to 12, characterized in that: The method further comprises: receiving a second acquisition request for the second neural network model sent by the client, wherein the second acquisition request carries second device information of the client; Selecting, according to the second device information, X-level parameters from among the P-level parameters; If any level parameter in the X-level parameters and any level parameter in the M-level parameters already acquired by the client have the same summary, the cloud platform sends the computation graph of the second neural network model and the parameters in the X-level parameters except the any level parameter to the client.
14. The method according to any one of claims 8 to 13, characterized in that: The first network information includes at least one of the available bandwidth and latency of the edge where the client is installed; or, the first computing resource information includes at least one of the utilization of the central processing unit CPU of the edge, the available memory, the utilization of the graphics processing unit GPU, and the available video memory.
15. A device for loading a neural network model, characterized in that: The device is applied to a cloud platform, the cloud platform stores all parameters of a first neural network model and a calculation graph of the first neural network model, and the device includes: A hierarchical compression module, used for hierarchically compressing all parameters of the first neural network model into N-level parameters, where N is an integer greater than 1, and there is no intersection between any two levels of parameters in the N-level parameters; A hierarchical storage module is used to store the N-level parameters; receive a first acquisition request for the first neural network model sent by a client, wherein the first acquisition request carries first device information of the client, and the first device information includes at least one of first network information and first computing resource information; select M-level parameters from the N-level parameters according to the first device information, wherein M is a positive integer less than or equal to N; and send the M-level parameters and a calculation graph of the first neural network model to the client.
16. A computing device cluster, characterized in that: The computing device cluster includes at least one computing device, each computing device including a processor and a memory; The processor of at least one computing device is used to execute instructions stored in the memory of the device, so that the computing device cluster executes the method for loading a neural network model as described in any one of claims 8 to 14.
17. A computer program product comprising instructions, characterized in that When the instruction is executed by a computing device cluster, the computing device cluster executes the method for loading a neural network model as described in any one of claims 8 to 14.
18. A computer-readable storage medium, characterized in that: It includes computer program instructions. When the computer program instructions are executed by a client, the computing device cluster executes the method for loading a neural network model as described in any one of claims 8 to 14.