Method and apparatus for loading neural network model, and device

By grading the parameters of the neural network model and selecting appropriate levels based on client resources, the problems of transmission delay and insufficient computing resources caused by large model parameters are solved, and efficient model transmission and edge operation are achieved.

WO2025130239A1PCT designated stage expired Publication Date: 2025-06-26HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/121962
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-30
Filing Date
2024-09-27
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

In the prior art, the increase in the number of parameters of the neural network model leads to a large amount of data, an increase in transmission delay, and the inability to operate normally when the edge-end computing resources are insufficient.

Method used

By grading the full parameters of the neural network model into multi-level parameters, and selecting appropriate parameter levels for transmission and loading based on the client's network and computing resource information, we ensure that the transmission delay is small and meets the requirements of edge computing resources.

Benefits of technology

It effectively reduces the transmission delay of neural network models and ensures that the model can be loaded and run normally on edge devices, improving the performance and reliability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024121962_26062025_PF_FP_ABST
    Figure CN2024121962_26062025_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of the cloud. Provided are a method and apparatus for loading a neural network model, and a device. The method comprises: a cloud platform dividing parameters of a first neural network model into N levels of parameters, and storing the N levels of parameters; an edge end sending to the cloud platform a first acquisition request for the first neural network model, wherein the first acquisition request carries first device information of the edge end, which first device information comprises at least one of first network information and first computing resource information; the cloud platform selecting M levels of parameters from among the N levels of parameters on the basis of the first device information, and the cloud platform sending to the edge end the M levels of parameters and a computational graph of the first neural network model; and the edge end loading the M levels of parameters and the computational graph. By means of the present application, the requirements of an edge end for the transmission and loading of a neural network model can be met.
Need to check novelty before this filing date? Find Prior Art

Description

Method, device and equipment for loading neural network model

[0001] This application claims priority to Chinese patent application No. 202311758255.X, filed on December 19, 2023, with the invention name “Method, device and equipment for loading a neural network model”, and Chinese patent application No. 202410381682.9, filed on March 30, 2024, with the invention name “Method, device and equipment for loading a neural network model”, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] The present application relates to the field of cloud technology, and in particular to a method, apparatus, and device for loading a neural network model. Background Art

[0003] With the development of artificial intelligence (AI), AI has been widely applied in various industries, including automotive, robotics, and healthcare. Various neural network models play a key role in AI. Neural network models can achieve functions such as object recognition, object segmentation, and scene understanding.

[0004] However, the number of parameters in neural network models is increasing, which makes the amount of data in neural network models larger and larger. When the edge network is poor, the delay in transmitting the neural network model from the cloud to the edge will be large. In addition, when the computing resources at the edge are insufficient, the neural network model with large data volume cannot operate normally.

[0005] Summary of the Invention

[0006] This application provides a method, apparatus, and device for loading a neural network model, which can better meet the client's requirements for transmitting and loading the neural network model. The corresponding technical solutions are as follows:

[0007] In a first aspect, a method for loading a neural network model is provided, the method comprising:

[0008] The cloud platform classifies all parameters of the first neural network model into N-level parameters and stores the N-level parameters. The client sends a first acquisition request for the first neural network model to the cloud platform, wherein the first acquisition request carries the client's first device information, and the first device information includes at least one of the first network information and the first computing resource information. Based on the first device information, the cloud platform selects M-level parameters from the N-level parameters, where M is a positive integer less than or equal to N. The cloud platform sends the M-level parameters and the computational graph of the first neural network model to the client. The client loads the M-level parameters and the computational graph of the first neural network model.

[0009] In the technical solution provided in the present application, the full parameters of the neural network model stored in the cloud platform are classified into multi-level parameters. These multi-level parameters are all sub-levels of the full parameters of the neural network model, and there is no intersection between any two levels of parameters in these multi-level parameters. The cloud platform can return parameters of the corresponding level to the client based on the client's current network conditions and computing resource conditions, which can minimize the problem of long parameter transmission delay caused by poor client network conditions, and can also minimize the problem of the client being unable to load parameters normally due to insufficient client computing resources.

[0010] In addition, by performing hierarchical compression and hierarchical storage on the parameters of the neural network model, it is possible to manage and update the parameters of each level of the neural network model separately.

[0011] In one possible implementation, the method further includes:

[0012] The cloud platform determines the loading priority of each level of parameters in the N-level parameters. Accordingly, the cloud platform selects M-level parameters from the N-level parameters based on the first device information, including:

[0013] The cloud platform determines M loading priorities from the loading priorities corresponding to the N level parameters according to the first device information. The cloud platform selects M level parameters corresponding to the M loading priorities from the N level parameters.

[0014] In the technical solution provided in this application, each level of parameters of the neural network model can correspond to a loading priority. For example, the parameters of the neural network model are compressed twice to obtain three levels of parameters, where the three levels of parameters are the remaining parameters after compression, the parameters compressed for the first time, and the parameters compressed for the second time. The loading priority of the remaining parameters after compression is the highest, the loading priority of the parameters compressed for the first time is the second, and the loading priority of the parameters compressed for the second time is the lowest. On this basis, the cloud platform can select parameters according to the loading priority corresponding to each level of parameters of the neural network model, giving priority to parameters with high loading priority.

[0015] In a possible implementation, the cloud platform determines M loading priorities from the loading priorities corresponding to the N level parameters according to the first device information, including:

[0016] The cloud platform determines M loading priorities corresponding to the first device information in the correspondence between the device information and the loading priorities.

[0017] In the technical solution provided in the present application, the cloud platform can pre-establish a correspondence between device information and loading priority. After receiving a request to obtain the first neural network model, the first device information carried therein is extracted. Then, the above correspondence can be queried to determine the M loading priorities corresponding to the first device information.

[0018] In one possible implementation, the cloud platform determines, based on at least one of the first network information and the first computing resource information, M loading priorities from the loading priorities corresponding to the hierarchically stored parameters of the first neural network model, including:

[0019] The cloud platform inputs the first device information into the parameter selection model to obtain M loading priorities.

[0020] In the technical solution provided in this application, the cloud platform can deploy a parameter selection model, which is also a neural network model. After receiving a request to obtain the first neural network model, the cloud platform extracts the first device information carried therein. Furthermore, the cloud platform can input the first device information into the parameter selection model to obtain M loading priorities output by the parameter selection model.

[0021] In one possible implementation, the method further includes:

[0022] The cloud platform compresses the parameters of the second neural network model into P-level parameters in a hierarchical manner. If any level parameter in the P-level parameters has the same summary as any level parameter in the stored N-level parameters, the cloud platform uses the stored level parameter and the parameters in the P-level parameters other than the level parameter as the hierarchical parameters of the second neural network model.

[0023] In the technical solution provided in the present application, if the first neural network model and the second neural network model have at least one level of parameters in common after hierarchical compression, then only one copy of the at least one level of parameters can be stored in the file system and shared by the first neural network model and the second neural network model, which can effectively save storage space.

[0024] In one possible implementation, the method further includes:

[0025] The client sends a second acquisition request for the second neural network model to the cloud platform, wherein the second acquisition request carries the client's second device information. The cloud platform selects X-level parameters from the P-level parameters of the second neural network model based on the second device information. If any level parameter in the X-level parameters has the same digest as any level parameter in the M-level parameters already acquired by the client, the cloud platform sends the computation graph of the second neural network model and the parameters in the X-level parameters other than the level parameter to the client.

[0026] In the technical solution provided in the present application, in order to save the client's storage resources, reduce the amount of data transmitted between the cloud platform and the client, and save client network resources, when the first neural network model and the second neural model have at least one level of parameters in common, and the terminal has obtained Y-level parameters among the at least one level of parameters of the first neural network model, when the terminal obtains the second neural network model from the cloud platform, the cloud platform may no longer send the Y-level parameters to the terminal.

[0027] In a possible implementation, the network information includes at least one of available bandwidth and latency of the client edge.

[0028] In a possible implementation, the computing resource information includes at least one of the utilization of the CPU of the edge, the available memory of the edge, the utilization of the GPU of the edge, and the available video memory of the client.

[0029] In a second aspect, a method for loading a neural network model is provided. The method is applied to a cloud platform, where the cloud platform stores all parameters of a first neural network model and a computational graph of the first neural network model. The method includes:

[0030] Classifying all parameters of the first neural network model into N-level parameters, where N is an integer greater than 1, and there is no intersection between any two levels of parameters in the N-level parameters;

[0031] Storing the N levels of parameters respectively;

[0032] Receiving a first acquisition request for the first neural network model sent by a client, wherein the first acquisition request carries first device information of the client, and the first device information includes at least one of first network information and first computing resource information;

[0033] Selecting M-level parameters from the N-level parameters according to the first device information, where M is a positive integer less than or equal to N;

[0034] Send the M-level parameters and the computation graph of the first neural network model to the client.

[0035] In one possible implementation, the method further includes:

[0036] Determining the loading priority of each level of parameters in the N levels of parameters;

[0037] According to the first device information, select M-level parameters from the N-level parameters:

[0038] Determining, according to the first device information, M loading priorities among the loading priorities corresponding to the N levels of parameters;

[0039] Among the N levels of parameters, M level parameters corresponding to the M loading priorities are selected.

[0040] In a possible implementation, determining, based on the first device information, M loading priorities from the loading priorities corresponding to the hierarchically stored parameters of the first neural network model includes:

[0041] In the correspondence between device information and loading priorities, M loading priorities corresponding to the first device information are determined.

[0042] In a possible implementation, determining, based on the first device information, M loading priorities from the loading priorities corresponding to the hierarchically stored parameters of the first neural network model includes:

[0043] The first device information is input into a parameter selection model to obtain M loading priorities.

[0044] In one possible implementation, the method further includes:

[0045] Hierarchically compressing the parameters of the second neural network model into P-level parameters, where P is an integer greater than 1;

[0046] If any level parameter in the P-level parameters has the same summary as any level parameter in the stored N-level parameters, the cloud platform will use the stored any level parameter and the parameters in the P-level parameters except the any level parameter as the parameters after classification of the second neural network model.

[0047] In one possible implementation, the method further includes:

[0048] Receiving a second acquisition request for a second neural network model sent by a client, wherein the first acquisition request carries second device information of the client;

[0049] Selecting X-level parameters from the P-level parameters according to the second device information;

[0050] If any level parameter in the X-level parameters and any level parameter in the M-level parameters already obtained by the client have the same summary, the cloud platform sends the computation graph of the second neural network model and the parameters in the X-level parameters other than the any level parameter to the client.

[0051] In a possible implementation, the network information includes at least one of available bandwidth and latency of the client edge.

[0052] In a possible implementation, the computing resource information includes at least one of the utilization of the CPU of the edge, the available memory of the edge, the utilization of the GPU of the edge, and the available video memory of the client.

[0053] In a third aspect, a device for loading a neural network model is provided, the device being applied to a cloud platform, the cloud platform storing all parameters of a first neural network model and a computational graph of the first neural network model, the device comprising:

[0054] a hierarchical compression module, configured to classify all parameters of the first neural network model into N-level parameters, where N is an integer greater than 1, and no intersection exists between any two parameters of the N-level parameters;

[0055] A hierarchical storage module is used to store the N-level parameters separately; receive a first acquisition request for the first neural network model sent by a client, wherein the first acquisition request carries first device information of the client, and the first device information includes at least one of first network information and first computing resource information; according to the first device information, select M-level parameters from the N-level parameters, wherein M is a positive integer less than or equal to N; and send the M-level parameters and a computational graph of the first neural network model to the client.

[0056] In a possible implementation, the hierarchical storage module is configured to:

[0057] Determining the loading priority of each level of parameters in the N levels of parameters;

[0058] Determining, according to the first device information, M loading priorities among the loading priorities corresponding to the N levels of parameters;

[0059] Among the N levels of parameters, M level parameters corresponding to the M loading priorities are selected.

[0060] In a possible implementation, the hierarchical storage module is configured to:

[0061] In the correspondence between device information and loading priorities, M loading priorities corresponding to the first device information are determined.

[0062] In a possible implementation, the hierarchical storage module is configured to:

[0063] The first device information is input into a parameter selection model to obtain M loading priorities.

[0064] In a possible implementation, the hierarchical compression module is further configured to:

[0065] Hierarchically compressing the parameters of the second neural network model into P-level parameters, where P is an integer greater than 1;

[0066] The hierarchical storage module is further used for:

[0067] If any level parameter in the P-level parameters has the same summary as any level parameter in the stored N-level parameters, the cloud platform will use the stored any level parameter and the parameters in the P-level parameters except the any level parameter as the parameters after classification of the second neural network model.

[0068] In a possible implementation, the hierarchical storage module is further configured to:

[0069] receiving a second acquisition request for the second neural network model sent by the client, wherein the first acquisition request carries second device information of the client;

[0070] Selecting X-level parameters from the P-level parameters according to the second device information;

[0071] If any level parameter in the X-level parameters and any level parameter in the M-level parameters already obtained by the client have the same summary, the cloud platform sends the computation graph of the second neural network model and the parameters in the X-level parameters other than the any level parameter to the client.

[0072] In a possible implementation, the network information includes at least one of available bandwidth and latency of the client edge.

[0073] In a possible implementation, the computing resource information includes at least one of the utilization of the CPU of the edge, the available memory of the edge, the utilization of the GPU of the edge, and the available video memory of the client.

[0074] In a fourth aspect, a computing device cluster is provided, the computing device cluster comprising at least one computing device, each computing device comprising a processor and a memory;

[0075] The processor of the at least one computing device is used to execute instructions stored in the memory of the device, so that the computing device cluster executes the method for loading the neural network model as described in the second aspect above.

[0076] In a fifth aspect, a computer program product comprising instructions is provided. When the instructions are executed by a computing device cluster, the computing device cluster executes the method for loading a neural network model as described in the second aspect above.

[0077] In a sixth aspect, a computer-readable storage medium is provided, comprising computer program instructions. When the computer program instructions are executed by a computing device, the computing device executes the method for loading a neural network model as described in the second aspect above. BRIEF DESCRIPTION OF THE DRAWINGS

[0078] FIG1 is a schematic diagram of a system for loading a neural network model provided in an embodiment of the present application;

[0079] FIG2 is a flow chart of a method for loading a neural network model provided in an embodiment of the present application;

[0080] FIG3 is a schematic diagram of hierarchical storage of a neural network model provided in an embodiment of the present application;

[0081] FIG4 is a flow chart of a method for loading a neural network model provided in an embodiment of the present application;

[0082] FIG5 is a schematic diagram of a neural network model loading provided in an embodiment of the present application;

[0083] FIG6 is a schematic diagram of the structure of a device for loading a neural network model provided in an embodiment of the present application;

[0084] FIG7 is a schematic diagram of the structure of a computing device provided in an embodiment of the present application;

[0085] FIG8 is a schematic diagram of the structure of a computing device cluster provided in an embodiment of the present application;

[0086] FIG9 is a schematic diagram of the structure of a computing device cluster provided in an embodiment of the present application. DETAILED DESCRIPTION

[0087] In order to facilitate the understanding of the embodiments of the present application, the terms involved in the embodiments of the present application are explained below.

[0088] Parameters of the neural network model:

[0089] Parameters in a neural network model refer to the variables that can be learned in the neural network model, including weights and biases.

[0090] Weight:

[0091] Weights represent the strength of the connections between neurons in a neural network. They determine the influence of input signals when they are transmitted between neurons.

[0092] Bias:

[0093] The bias is the offset or threshold of a neuron, which allows it to fire even when there is no input.

[0094] Neurons:

[0095] The basic unit of a neural network that contains a single data feature.

[0096] Computational graph:

[0097] It includes the connection relationship of each neuron and records information such as the operation type, input, output, and properties of each neuron.

[0098] The following describes the method for loading a neural network model provided in an embodiment of the present application with reference to the accompanying drawings.

[0099] An embodiment of the present application provides a method for loading a neural network model, which can be implemented by a cloud platform and a client. The cloud platform can be a computing device or a cluster of computing devices. The client can be software installed on the edge for interacting with the cloud service provided by the cloud platform. The edge can be a mobile phone, computer, robot, server and other devices.

[0100] Referring to Figure 1, a system architecture diagram of a neural network model loading method provided by an embodiment of the present application is shown. As shown in Figure 1, the cloud platform includes underlying hardware and cloud services, wherein the underlying hardware such as processors, memories, etc., the processor can run cloud services, the memory can store the full parameters of the neural network model and the calculation graph of the neural network model, and the edge also includes underlying hardware and clients, wherein the underlying hardware such as processors, memories, etc., the processor can run the client, the client is used to interact with the cloud services provided by the cloud platform, and the memory can store the parameters and calculation graph of the neural network model obtained by the client from the cloud platform.

[0101] A neural network model can consist of a computational graph and parameters. The cloud platform can perform hierarchical compression on all parameters of a obtained neural network model. For example, neural network model 1 in Figure 1 has a total of 1100 parameters. These parameters are hierarchically compressed into three levels: the first level consists of 600 parameters, the second level consists of 300 parameters, and the third level consists of 200 parameters. These three levels of parameters are subsets of the parameters of neural network model 1, and no two of these three levels of parameters intersect. Neural network model 2 has a total of 700 parameters. These parameters are hierarchically compressed into two levels: the first level consists of 500 parameters, and the second level consists of 200 parameters. These two levels of parameters are subsets of the parameters of neural network model 2. When a client needs to use a neural network model, it can send a request to the cloud platform to obtain the neural network model. For example, if a client needs to use neural network model 1, it can send a request to the cloud platform to obtain neural network model 1. The request contains the client's network information and computing resource information. After receiving a request to obtain neural network model 1, the cloud platform selects at least one level of parameters from the three levels of neural network model 1 that meets the network and computing resource requirements based on the client's network information and computing resource information, such as selecting the first two levels of parameters. The selected first two levels of parameters and the computational graph of neural network model 1 are then returned to the client, which then loads the first two levels of parameters and the computational graph of neural network model 1. This allows the cloud platform to return the corresponding level of parameters to the client based on the client's current network and computing resource conditions, minimizing the risk of long parameter transmission delays due to poor client network conditions and the inability of the client to load parameters due to insufficient client computing resources.

[0102] Referring to FIG2 , a flowchart of a method for loading a neural network model provided by an embodiment of the present application is shown. As shown in FIG2 , the processing of the method may include the following steps:

[0103] Step 201: The platform classifies the parameters of the first neural network model into N levels of parameters.

[0104] Among them, the first neural network model is any neural network model obtained by the cloud platform, and N is an integer greater than 1.

[0105] In implementation, the cloud platform can obtain the neural network model through various methods such as user upload, manufacturer self-development, and public resources. For the first neural network model obtained, the cloud platform can hierarchically compress the parameters of the first neural network model into N levels of parameters. Here, the hierarchical compression can be implemented using compression algorithms such as pruning, quantization, and distillation. The specific number of compression levels can be configured by relevant personnel based on a comprehensive consideration of factors such as the total number of parameters of the neural network model and storage resources, or the number of compression levels can be automatically determined by the compression algorithm. The embodiments of this application do not limit the compression method and compression level of the neural network model.

[0106] When the cloud platform performs hierarchical compression on the parameters of the first neural network model, the remaining parameters after compression are used as first-level parameters separately, and the parameters compressed at each level are used as first-level parameters separately. Among them, the reasoning can also be completed by using the remaining parameters after compression in the first neural network model alone, but the reasoning effect may be worse than using the full amount of parameters.

[0107] For example, as shown in Figure 1, neural network model 1 has a total of 1,100 parameters. The cloud platform compresses the parameters of neural network model 1 into three levels of parameters. That is, neural network model 1 is compressed twice, with 300 parameters compressed in the first compression and 200 parameters compressed in the second compression, leaving 600 parameters. The remaining 600 parameters are used as the first-level parameters, the 300 parameters compressed in the first compression are used as the second-level parameters, and the 200 parameters compressed in the second compression are used as the third-level parameters. Inference can be completed using the remaining 600 parameters in neural network model 1 alone, but the inference effect may be worse than using the full 11,000 parameters. Inference can be performed using the remaining 600 parameters and the 300 parameters compressed in the first compression in neural network model 1, and the inference effect is better than using the remaining parameters alone. In neural network model 1, using the remaining 600 parameters after compression, the 300 parameters compressed in the first round, and the 200 parameters compressed in the second round for inference, that is, using the full 1100 parameters for inference, achieves better inference results than using the remaining 600 parameters after compression and the 300 parameters compressed in the first round for inference.

[0108] In one possible implementation, when hierarchically compressing the parameters of the first neural network model, metadata for each level of parameters may also be generated. The metadata for each level of parameters may include a parameter identifier for this level (layer id), a loading priority (priority), a data size (size), a parameter identifier for the next level (next layer id), a globally unique identifier for the parameter for this level, the computing power required for loading, and the like. Each item in the metadata is described below.

[0109] Parameter identification for this level: The layer ID of each level parameter is unique among all levels of parameters in the first neural network model. For example, taking the above neural network model 1 as an example, the layer IDs of the three levels of parameters in the neural network model 1 can be layer1, layer2, and layer3 respectively.

[0110] Loading priority: This indicates the priority of the parameters at this level when loaded by the client. The order of loading priority for each level is as follows: parameters remaining after compression, parameters lost in the first compression, parameters lost in the second compression, etc. For example, taking the neural network model 1 above as an example, the 600 parameters remaining after compression have the highest loading priority, which is recorded as 1. The 300 parameters lost in the first compression have the second highest loading priority, which is recorded as 2. The 200 parameters lost in the second compression have the lowest loading priority, which is recorded as 3.

[0111] Data volume: the data volume of the parameters at this level.

[0112] Next-level parameter identifier: The parameters remaining after compression are first-level parameters, the parameters compressed for the first time are second-level parameters, the parameters compressed for the second time are third-level parameters, and so on. The next-level parameter identifier of the parameters remaining after compression is the current-level parameter identifier of the parameters compressed for the first time, and the next-level parameter identifier of the parameters compressed for the first time is the current-level parameter identifier of the parameters compressed for the second time. For example, taking the above-mentioned neural network model 1 as an example, the next-level parameter identifier of the first-level parameter of the neural network model 1 is layer2, the next-level parameter identifier of the second-level parameter is layer3, and if the third-level parameter has no next-level parameter, its next-level parameter identifier can be empty.

[0113] Loading required computing power: The client loads the computing power required to run the parameters at this level. The computing power here can be the computing power of the central processing unit (CPU) or graphics processing unit (GPU).

[0114] Globally unique identifier of the parameters at this level: The globally unique identifier refers to the unique identifier of the parameters at this level in the file system that stores the parameters of the neural network model. It can be the digest of the parameters at this level. The digest algorithm can be a hash algorithm, a hash algorithm, etc. If the globally unique identifiers of the parameters at two levels are the same, it means that the parameters at the two levels are the same.

[0115] Step 202: The cloud platform stores the N-level parameters.

[0116] During implementation, the cloud platform can store the N-level parameters of the first neural network model separately through a joint file system.

[0117] In one possible implementation, each level of parameters has a corresponding globally unique identifier. Before storing each level of parameters in the N levels of parameters of the first neural network model, the target globally unique identifier corresponding to the level parameter can be used to query whether the file system stores the parameters corresponding to the target globally unique identifier. If the parameters corresponding to the target globally unique identifier are found, the parameters for that level are not stored, and the stored parameters corresponding to the target globally unique identifier are used as the parameters for that level of the first neural network model. If the parameters corresponding to the target globally unique identifier are not found, the parameters for that level are stored in the file system.

[0118] For example, referring to FIG3 , neural network model 3 is hierarchically compressed into two-level parameters, wherein the globally unique identifier of the first-level parameters is digest1, and the globally unique identifier of the second-level parameters is digest2. Both levels of parameters of neural network model 3 have been stored separately in the file system. After hierarchically compressing neural network model 4, three levels of parameters are obtained, wherein the globally unique identifier of the first-level parameters is digest1, the globally unique identifier of the second-level parameters is digest3, and the globally unique identifier of the third-level parameters is digest4. When storing the first-level parameters of neural network model 4, if the file system is queried to find parameters with the globally unique identifier digest1 stored, the first-level parameters of neural network model 4 are not stored. Neural network model 4 and neural network model 3 share the same first-level parameters. When storing the second-level parameters of neural network model 4, if the file system is queried to find parameters with the globally unique identifier digest3 not stored, the second-level parameters of neural network model 4 are stored in the file system. When storing the third-level parameters of neural network model 4, if the file system is queried to find parameters with the globally unique identifier digest4 not stored, the third-level parameters of neural network model 4 are stored in the file system.

[0119] Step 203: The client sends a first acquisition request for the first neural network model to the cloud platform.

[0120] The first acquisition request carries the first device information of the client, and the first device information includes at least one of the first network information and the first computing resource information.

[0121] During implementation, the client can obtain the first neural network model from the cloud platform based on actual needs. Specifically, the client can send a first acquisition request for the first neural network model to the cloud platform, wherein the first acquisition request carries the client's first device information, and the first device information includes at least one of the first network information and the first computing resource information. The first network information may include the client's current available bandwidth, latency, number of network connections, etc. The first computing resource information may include the client's current CPU utilization, available memory, GPU utilization, GPU remaining computing power, available video memory, etc. In addition, the first device information may also include information such as device temperature and device operation log to describe the device's operating status.

[0122] Step 204: The cloud platform selects M-level parameters from the N-level parameters of the first neural network model based on the first device information.

[0123] Wherein, M is a positive integer less than or equal to N.

[0124] In implementation, after receiving a first acquisition request for a first neural network model from a client, the cloud platform selects M-level parameters from the N-level parameters of the first neural network model based on the first device information carried in the first acquisition request. There are multiple possible selection methods, several of which are exemplified below.

[0125] Method 1:

[0126] For each neural network model, a correspondence between device information and loading priorities can be pre-established. Upon receiving a first acquisition request for a first neural network model, the cloud platform queries this correspondence and determines M loading priorities corresponding to the first device information, where each loading priority corresponds to a first-level parameter. Furthermore, M-level parameters corresponding to these M loading priorities can be selected.

[0127] Taking the above neural network model 1 as an example, the correspondence between device information and loading priority can be shown in Table 1 below.

[0128] Table 1

[0129] Among them, C1, C2, and C3 are the CPU utilization thresholds, R1, R2, and R3 are the available memory thresholds, G1, G2, and G3 are the GPU utilization thresholds, T1, T2, and T3 are the GPU remaining computing power thresholds, X1, X2, and X3 are the available video memory thresholds, B1, B2, and B3 are the available bandwidth thresholds, and D1, D2, and D3 are the latency thresholds.

[0130] After receiving the acquisition request for neural network model 1 sent by the terminal, the cloud platform extracts the device information carried in the acquisition request, including the terminal's CPU utilization rate C, available memory R, ​​GPU utilization rate G, GPU remaining computing power T, available video memory X, available loan B, and latency D. Then, it queries Table 1 to determine that the loading priorities corresponding to the device information are 1 and 2, and then selects the parameters with loading priority 1 and loading priority 2 of neural network model 1.

[0131] Method 2:

[0132] The cloud platform can deploy a pre-trained parameter selection model, which can be a neural network model. Upon receiving a first acquisition request for a first neural network model, the cloud platform extracts the first device information contained in the first acquisition request. The cloud platform then inputs the first device information, the data volume of each level of parameters of the first neural network model, and the computing power required to load each level of parameters of the first neural network model into the parameter selection model. The parameter selection model then outputs M loading priorities. The neural network model then obtains M levels of parameters corresponding to each of the M loading priorities.

[0133] Step 205: The cloud platform sends the above-mentioned M-level parameters and the calculation graph of the first neural network model to the client.

[0134] During implementation, the cloud platform can read the M-level parameters of the selected first neural network model in the file system, and send the M-level parameters of the first neural network model and the computational graph of the first neural network model to the terminal.

[0135] Step 206: The client loads the above-mentioned M-level parameters and the calculation graph of the first neural network model.

[0136] During implementation, after receiving the M-level parameters of the first neural network model and the computational graph of the first neural network model sent by the cloud platform, the client stores the M-level parameters of the first neural network model and the computational graph of the first neural network model, and after receiving the loading and running instructions of the first neural model, loads the M-level parameters of the first neural network model and the computational graph of the first neural network model.

[0137] In one possible implementation, in order to save client storage resources and reduce the amount of data transmitted between the cloud platform and the client, if there is at least one level of parameters in common between the first neural network model and the second neural network model, and the terminal has already obtained Y level parameters of the at least one level of parameters of the first neural network model, when the terminal obtains the second neural network model from the cloud platform, the cloud platform may no longer send these Y level parameters to the terminal. The following describes how the client obtains the second neural network model. As shown in Figure 4, the process may include the following steps:

[0138] Step 401: The client sends a second acquisition request for a second neural network model to the cloud platform.

[0139] The first acquisition request carries the second device information of the client and the first globally unique identifier corresponding to each level of parameters of the locally stored neural network model.

[0140] In implementation, when the client obtains the neural network model, the cloud platform can send the first globally unique identifier corresponding to these parameters to the terminal in addition to sending the parameters of the neural network model and the calculation graph of the neural network model to the terminal.

[0141] Correspondingly, when the client sends a second acquisition request for the second neural network model to the cloud platform, in addition to sending the second device information of the client, it can also send the globally unique identifiers corresponding to the parameters of each level of the locally stored neural network model.

[0142] Step 402: The cloud platform selects X-level parameters from the P-level parameters of the second neural network model based on the second device information.

[0143] The specific processing of step 402 is the same as that of step 204 above, and will not be described in detail here.

[0144] Step 403: If the X-level parameters include the Y-level parameters in the W-level parameters stored locally on the client, the cloud platform sends the computation graph of the second neural network model and the XY-level parameters in the X parameters excluding the Y-level parameters to the client.

[0145] During implementation, when the cloud platform selects X-level parameters, it can also obtain the globally unique identifiers corresponding to these X-level parameters. If Y of the globally unique identifiers corresponding to these X-level parameters are the same as the Y globally unique identifiers corresponding to the Y-level parameters of the neural network model stored locally on the client, then it means that the parameters corresponding to these Y globally unique identifiers of the second neural network are the same as the Y-level parameters of the neural network model stored locally on the client. In this case, the cloud platform does not need to send these Y-level parameters to the client. Instead, the cloud platform can send the computation graph of the second neural network model and the XY-level parameters of the X parameters other than the Y-level parameters. In addition, the cloud platform can also send the X globally unique identifiers to the client so that the client can locally find the Y-level parameters that the cloud platform did not send to the client this time based on the X globally unique identifiers.

[0146] For example, as shown in Figure 5, neural network model 3 is hierarchically compressed into two levels of parameters, where the first-level parameters have a globally unique identifier of digest1 and the second-level parameters have a globally unique identifier of digest2. Neural network model 4 is hierarchically compressed into three levels of parameters, where the first-level parameters have a globally unique identifier of digest1, the second-level parameters have a globally unique identifier of digest3, and the third-level parameters have a globally unique identifier of digest4. The client has already obtained neural network model 3 and has the first-level parameters of neural network model 3 stored locally on the client. When the client sends a request to the cloud platform to obtain neural network model 4, it includes the globally unique identifier of digest1 for the first-level parameters of neural network model 3 and the client's device information. The cloud platform selects the first-level and second-level parameters of neural network model 4 based on the client's device information. If it then determines that the first-level parameters of neural network model 4 also have a globally unique identifier of digest1, the cloud platform does not need to send the first-level parameters of neural network model 4, which have a globally unique identifier of digest1, to the client. Instead, it only needs to send the second-level parameters of neural network model 4, which have a globally unique identifier of digest3, and the computational graph of neural network model 4.

[0147] In this way, the client does not need to repeatedly store the second-level parameters of the neural network model 4, which saves the client's computing resources. The cloud platform also does not need to send the second-level parameters of the neural network model 4 to the client, which reduces the amount of transmitted data, improves transmission efficiency, and reduces network resource usage.

[0148] The embodiment of the present application further provides a device for loading a neural network model, which may be a cloud platform. As shown in FIG6 , the device may include a hierarchical compression module 410 and a hierarchical storage module 420 , wherein:

[0149] A hierarchical compression module 410 is configured to hierarchically compress all parameters of the first neural network model into N-level parameters, where N is an integer greater than 1, and no two parameters in the N-level parameters intersect.

[0150] A hierarchical storage module 420 is used to store the N-level parameters; receive a first acquisition request for the first neural network model sent by a client, wherein the first acquisition request carries first device information of the client, and the first device information includes at least one of first network information and first computing resource information; according to the first device information, select M-level parameters from the N-level parameters, wherein M is a positive integer less than or equal to N; and send the M-level parameters and a computational graph of the first neural network model to the client.

[0151] In a possible implementation, the hierarchical storage module 420 is configured to:

[0152] Determining the loading priority of each level of parameters in the N levels of parameters;

[0153] Determining, according to the first device information, M loading priorities among the loading priorities corresponding to the N levels of parameters;

[0154] Among the N levels of parameters, M level parameters corresponding to the M loading priorities are selected.

[0155] In a possible implementation, the hierarchical storage module 420 is configured to:

[0156] In the correspondence between device information and loading priorities, M loading priorities corresponding to the first device information are determined.

[0157] In a possible implementation, the hierarchical storage module 420 is configured to:

[0158] The first device information is input into a parameter selection model to obtain M loading priorities.

[0159] In a possible implementation, the hierarchical compression module 410 is further configured to:

[0160] Hierarchically compressing the parameters of the second neural network model into P-level parameters, where P is an integer greater than 1;

[0161] The hierarchical storage module 420 is further configured to:

[0162] If any level parameter in the P-level parameters has the same summary as any level parameter in the stored N-level parameters, the cloud platform will use the stored any level parameter and the parameters in the P-level parameters except the any level parameter as the parameters after classification of the second neural network model.

[0163] In a possible implementation, the hierarchical storage module 420 is further configured to:

[0164] receiving a second acquisition request for the second neural network model sent by the client, wherein the first acquisition request carries second device information of the client;

[0165] Selecting X-level parameters from the P-level parameters according to the second device information;

[0166] If any level parameter in the X-level parameters and any level parameter in the M-level parameters already obtained by the client have the same summary, the cloud platform sends the computation graph of the second neural network model and the parameters in the X-level parameters other than the any level parameter to the client.

[0167] In a possible implementation, the network information includes at least one of available bandwidth and latency of the client edge.

[0168] In a possible implementation, the computing resource information includes at least one of the utilization of the CPU of the edge, the available memory of the edge, the utilization of the GPU of the edge, and the available video memory of the client.

[0169] The hierarchical compression module 410 and the hierarchical storage module 420 can be implemented in software or hardware. For example, the implementation of the hierarchical compression module 410 is described below. Similarly, the implementation of the hierarchical storage module 420 can refer to the implementation of the hierarchical compression module 410.

[0170] As an example of a software functional unit, the hierarchical compression module 410 may include code running on a computing instance. The computing instance may be at least one of a physical host (computing device), a virtual machine, a container, and other computing devices. Furthermore, the above-mentioned computing device may be one or more. For example, the hierarchical compression module 410 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the application may be distributed in the same region or in different regions. The multiple hosts / virtual machines / containers used to run the code may be distributed in the same AZ or in different AZs, and each AZ includes one data center or multiple data centers with close geographical locations. Generally, a region may include multiple AZs.

[0171] Similarly, the multiple hosts / virtual machines / containers running the code can be distributed within the same VPC or across multiple VPCs. Typically, a VPC is located within a region. Cross-region communication between two VPCs within the same region, or between VPCs in different regions, requires a communication gateway within each VPC to interconnect the VPCs.

[0172] As an example of a hardware functional unit, hierarchical compression module 410 may include at least one computing device, such as a server. Alternatively, hierarchical compression module 410 may be implemented using an ASIC or a PLD. The PLD may be implemented using a CPLD, FPGA, GAL, or any combination thereof.

[0173] The multiple computing devices included in the hierarchical compression module 410 can be distributed in the same region or in different regions. The multiple computing devices included in the hierarchical compression module 4100 can be distributed in the same AZ or in different AZs. Similarly, the multiple computing devices included in the management module 410 can be distributed in the same VPC or in multiple VPCs. The multiple computing devices can be any combination of servers, ASICs, PLDs, CPLDs, FPGAs, GALs, and other computing devices.

[0174] This application also provides a computing device 100. As shown in FIG7 , computing device 100 includes a bus 102, a processor 104, a memory 106, and a communication interface 108. Processor 104, memory 106, and communication interface 108 communicate with each other via bus 102. Computing device 100 may be a server or a terminal device. It should be understood that this application does not limit the number of processors and memories in computing device 100.

[0175] Bus 102 may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, among others. Buses may be classified as address buses, data buses, control buses, and the like. For ease of illustration, FIG7 illustrates a single bus line, but this does not imply a single bus or type of bus. Bus 102 may include a path for transmitting information between various components of computing device 100 (e.g., memory 106, processor 104, and communication interface 108).

[0176] The processor 104 may include any one or more processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).

[0177] The memory 106 may include volatile memory, such as random access memory (RAM). The memory 106 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD).

[0178] Memory 106 stores executable code, which processor 104 executes to implement the functions of the aforementioned hierarchical compression module 410 and hierarchical storage module 420, thereby implementing the method for loading a neural network model. In other words, memory 106 stores instructions for executing the method for loading a neural network model.

[0179] The communication interface 108 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 100 and other devices or a communication network.

[0180] Embodiments of the present application also provide a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smartphone.

[0181] As shown in Figure 8, the computing device cluster includes at least one computing device 100. The memory 106 in one or more computing devices 100 in the computing device cluster may store the same instructions for executing the method for loading a neural network model.

[0182] In some possible implementations, the memory 106 of one or more computing devices 100 in the computing device cluster may also store partial instructions for executing the method for loading the neural network model. In other words, the combination of one or more computing devices 100 can jointly execute the instructions for executing the method for loading the neural network model.

[0183] It should be noted that the memory 106 in different computing devices 100 in the computing device cluster can store different instructions, which are respectively used to execute part of the functions of the method for loading the neural network model.

[0184] In some possible implementations, the memory 106 of one or more computing devices 100 in the computing device cluster may also store partial instructions for executing the method for loading the neural network model. In other words, the combination of one or more computing devices 100 can jointly execute the instructions for executing the method for loading the neural network model.

[0185] It should be noted that the memory 106 in different computing devices 100 in the computing device cluster can store different instructions for executing portions of the neural network model loading method. In other words, the instructions stored in the memory 106 in different computing devices 100 can implement the functions of one or more nodes in the hierarchical compression module 410 and the hierarchical storage module 420.

[0186] In some possible implementations, one or more computing devices in a computing device cluster may be connected via a network. The network may be a wide area network or a local area network, etc. FIG. 9 shows a possible implementation. As shown in FIG. 9 , two computing devices 100A and 100B are connected via a network. Specifically, the network is connected via a communication interface in each computing device. In this type of possible implementation, the memory 106 in the computing device 100A stores instructions for executing the functions of the hierarchical compression module 410. At the same time, the memory 106 in the computing device 100B stores instructions for executing the functions of the hierarchical storage module 420. It should be understood that the functions of the computing device 100A shown in FIG. 9 may also be performed by multiple computing devices 100. Similarly, the functions of the computing device 100B may also be performed by multiple computing devices 100.

[0187] The present application also provides another computing device cluster. The connection relationship between the computing devices in the computing device cluster can be similar to the connection method of the computing device cluster described in Figures 8 and 9. The difference is that the memory 106 of one or more computing devices 100 in the computing device cluster can store the same instructions for executing the method for loading the neural network model.

[0188] In some possible implementations, the memory 106 of one or more computing devices 100 in the computing device cluster may also store partial instructions for executing the method for loading the neural network model. In other words, the combination of one or more computing devices 100 can jointly execute the instructions for executing the method for loading the neural network model.

[0189] It should be noted that the memory 106 in different computing devices 100 in the computing device cluster can store different instructions for executing partial functions of the apparatus for loading the neural network model. In other words, the instructions stored in the memory 106 in different computing devices 100 can implement the functions of one or more nodes in the hierarchical compression module 410 and the hierarchical storage module 420.

[0190] The present application also provides a computer program product comprising instructions. The computer program product may be software or a program product comprising instructions that can be run on a computing device or stored in any available medium. When the computer program product is run on a computing device, it causes at least one computing device to execute the method for loading a neural network model provided in the present application.

[0191] The embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that can be stored by a computing device or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid-state drive). The computer-readable storage medium includes instructions that instruct the computing device to execute the method for loading the neural network model provided in the embodiment of the present application.

[0192] In this application, the terms "first", "second", etc. are used to distinguish between identical or similar items that have substantially the same effects and functions. It should be understood that there is no logical or temporal dependency between "first" and "second", nor is there any limitation on quantity and execution order. It should also be understood that although the following description uses the terms first, second, etc. to describe various elements, these elements should not be limited by the terms. These terms are only used to distinguish one element from another. For example, without departing from the scope of the various examples, a first neural network model may be referred to as a second neural network model, and similarly, a second neural network model may be referred to as a first neural network model. Both the first neural network model and the second neural network model may be collectively referred to as neural network models, and in some cases, may be separate and different feature vectors.

[0193] The term "at least one" in this application means one or more, and the term "plurality" in this application means two or more.

[0194] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and such modifications or substitutions should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.

Claims

1. A method for loading a neural network model, characterized in that: The method is applied to a cloud platform and a client, wherein the cloud platform stores all parameters and a calculation graph of a first neural network model, and the method includes: The cloud platform classifies all parameters of the first neural network model into N-level parameters, where N is an integer greater than 1, and there is no intersection between any two levels of parameters in the N-level parameters; The cloud platform stores the N-level parameters; The client sends a first acquisition request for the first neural network model to the cloud platform, wherein the first acquisition request carries first device information of the client, and the first device information includes at least one of first network information and first computing resource information; The cloud platform selects M-level parameters from the N-level parameters according to the first device information, where M is a positive integer less than or equal to N; The cloud platform sends the M-level parameters and the calculation graph to the client; The client loads the M-level parameters and the calculation graph.

2. The method according to claim 1, characterized in that The method further comprises: The cloud platform determines the loading priority of each level of parameters in the N levels of parameters; The cloud platform selects M-level parameters from the N-level parameters according to the first device information, including: The cloud platform determines, according to the first device information, M loading priorities among the loading priorities respectively corresponding to the N level parameters; The cloud platform selects M-level parameters corresponding to the M loading priorities from the N-level parameters.

3. The method according to claim 2, characterized in that The cloud platform determines M loading priorities among the loading priorities respectively corresponding to the N level parameters according to the first device information, including: The cloud platform determines M loading priorities corresponding to the first device information in the corresponding relationship between the device information and the loading priorities.

4. The method according to claim 2, characterized in that: The cloud platform determines M loading priorities among the loading priorities respectively corresponding to the N level parameters according to the first device information, including: The cloud platform inputs the first device information into a parameter selection model to obtain M loading priorities.

5. The method according to any one of claims 1 to 4, characterized in that The method further comprises: The cloud platform classifies the parameters of the second neural network model into P-level parameters, wherein P is an integer greater than 1; If any level parameter in the P-level parameters has the same summary as any level parameter in the stored N-level parameters, the cloud platform will use the stored any level parameter and the parameters in the P-level parameters except the any level parameter as the graded parameters of the second neural network model.

6. The method according to claim 5, characterized in that The method further comprises: The client sends a second acquisition request for the second neural network model to the cloud platform, wherein the second acquisition request carries second device information of the client; The cloud platform selects X-level parameters from the P-level parameters according to the second device information; If any level parameter in the X-level parameters and any level parameter in the M-level parameters already acquired by the client have the same summary, the cloud platform sends the computation graph of the second neural network model and the parameters in the X-level parameters except the any level parameter to the client.

7. The method according to any one of claims 1 to 6, characterized in that The first network information includes at least one of the available bandwidth and latency of the edge where the client is installed; or, the first computing resource information includes at least one of the utilization of the central processing unit CPU of the edge, the available memory, the utilization of the graphics processing unit GPU, and the available video memory.

8. A method for loading a neural network model, characterized in that: The method is applied to a cloud platform, the cloud platform stores all parameters of a first neural network model and a calculation graph of the first neural network model, and the method includes: Classifying all parameters of the first neural network model into N-level parameters, where N is an integer greater than 1, and there is no intersection between any two levels of parameters in the N-level parameters; Storing the N-level parameters; Receiving a first acquisition request for the first neural network model sent by a client, wherein the first acquisition request carries first device information of the client, and the first device information includes at least one of first network information and first computing resource information; According to the first device information, select M-level parameters from the N-level parameters, where M is a positive integer less than or equal to N; Send the M-level parameters and a computational graph of the first neural network model to the client.

9. The method according to claim 8, characterized in that The method further comprises: Determining the loading priority of each level of parameters in the N levels of parameters; Selecting M-level parameters from the N-level parameters according to the first device information includes: Determine, according to the first device information, M loading priorities among the loading priorities respectively corresponding to the N level parameters; Among the N levels of parameters, select M levels of parameters corresponding to the M loading priorities.

10. The method according to claim 9, characterized in that The step of determining, according to the first device information, M loading priorities among the loading priorities corresponding to the hierarchically stored parameters of the first neural network model, comprises: In the correspondence between device information and loading priorities, M loading priorities corresponding to the first device information are determined.

11. The method according to claim 9, characterized in that The step of determining, according to the first device information, M loading priorities from the loading priorities corresponding to the hierarchically stored parameters of the first neural network model comprises: The first device information is input into a parameter selection model to obtain M loading priorities.

12. The method according to any one of claims 8 to 11, characterized in that The method further comprises: The parameters of the second neural network model are hierarchically compressed into P-level parameters, wherein P is an integer greater than 1; If any level parameter in the P-level parameters has the same summary as any level parameter in the stored N-level parameters, the cloud platform will use the stored any level parameter and the parameters in the P-level parameters except the any level parameter as the graded parameters of the second neural network model.

13. The method according to any one of claims 8 to 12, characterized in that: The method further comprises: receiving a second acquisition request for the second neural network model sent by the client, wherein the second acquisition request carries second device information of the client; Selecting, according to the second device information, X-level parameters from among the P-level parameters; If any level parameter in the X-level parameters and any level parameter in the M-level parameters already acquired by the client have the same summary, the cloud platform sends the computation graph of the second neural network model and the parameters in the X-level parameters except the any level parameter to the client.

14. The method according to any one of claims 8 to 13, characterized in that: The first network information includes at least one of the available bandwidth and latency of the edge where the client is installed; or, the first computing resource information includes at least one of the utilization of the central processing unit CPU of the edge, the available memory, the utilization of the graphics processing unit GPU, and the available video memory.

15. A device for loading a neural network model, characterized in that: The device is applied to a cloud platform, the cloud platform stores all parameters of a first neural network model and a calculation graph of the first neural network model, and the device includes: A hierarchical compression module, used for hierarchically compressing all parameters of the first neural network model into N-level parameters, where N is an integer greater than 1, and there is no intersection between any two levels of parameters in the N-level parameters; The hierarchical storage module is configured to store the N-level parameters; receive a first acquisition request for the first neural network model sent by a client, wherein the first acquisition request carries first device information of the client, and the first device information includes a first network at least one of the network information and the first computing resource information; according to the first device information, selecting M-level parameters from the N-level parameters, where M is a positive integer less than or equal to N; and sending the M-level parameters and the calculation graph of the first neural network model to the client.

16. A computing device cluster, characterized in that: The computing device cluster includes at least one computing device, each computing device including a processor and a memory; The processor of at least one computing device is used to execute instructions stored in the memory of the device, so that the computing device cluster executes the method for loading a neural network model as described in any one of claims 8 to 14.

17. A computer program product comprising instructions, characterized in that When the instruction is executed by a computing device cluster, the computing device cluster executes the method for loading a neural network model as described in any one of claims 8 to 14.

18. A computer-readable storage medium, characterized in that: It includes computer program instructions. When the computer program instructions are executed by a client, the computing device cluster executes the method for loading a neural network model as described in any one of claims 8 to 14.

Citation Information

Patent Citations

  • Data processing method, end side device, cloud side device and end-cloud cooperative system

    CN108243216A

  • Model performance optimization method and related equipment

    CN116720566A

  • A method, an apparatus and a computer program product for neural networks

    EP3683733A1

  • Delivery of compressed neural networks

    EP3767548A1

  • Delivery of compressed neural networks

    EP3767549A1