Machine learning model training and deployment method and device, equipment and storage medium
By using grid search and cross-validation methods on the cloud computing platform for distributed training, combined with model compression and inference performance optimization, the deployment and stable operation of machine learning models in the production environment is solved, and efficient optimization and deployment is achieved.
Patent Information
- Application Number
- CN202510096043.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-23
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The efficient deployment and stable operation of machine learning models in production environments faces challenges, especially in how to optimize model performance and ensure its stable operation.
The initial model is distributedly trained in combination with training data to obtain the optimal model parameters; then the initial model is compressed and inference performance optimized to obtain the second machine learning model; finally, the second machine learning model is deployed on the cloud computing platform using a microservice architecture.
It realizes efficient optimization and deployment of machine learning models, improves the performance and accuracy of the model, and ensures its stable operation in production environment.
Smart Images

Figure CN120031107A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of machine learning technology, and in particular to a method, apparatus, device and storage medium for training and deploying a machine learning model. Background Art
[0002] Machine learning is one of the core branches of artificial intelligence, which enables computers to learn and improve through data without explicit programming. The optimization of machine learning models is crucial to the performance and accuracy of the models, and is also the basis for subsequent deployment and optimization. At the same time, how to efficiently deploy machine learning models to production environments and ensure their stable operation also faces many challenges. Summary of the invention
[0003] The present application aims to solve one of the technical problems in the related art at least to some extent.
[0004] In the first aspect, the present application proposes a method for training and deploying a machine learning model, which is applied to a cloud computing platform. The method includes: using grid search and cross-validation methods in combination with training data to perform distributed training on the initial model to obtain optimal model parameters; obtaining a first machine learning model based on the optimal model parameters; performing model compression and inference performance optimization on the first machine learning model to obtain a second machine learning model; and deploying the second machine learning model on a cloud computing platform using a microservice architecture.
[0005] In one implementation, the grid search and cross-validation method is combined with training data to perform distributed training on the initial model to obtain optimal model parameters, including: dividing the training data into at least one sub-data set; dividing the initial model into at least one sub-model unit; obtaining a first value based on the number of the sub-data sets, the number of the sub-model units, the total amount of data of the training data and the total amount of model of the initial model; determining a target parallel scheme based on the first value and a preset first threshold; and performing distributed training on the sub-model units based on the target parallel scheme and using the grid search and cross-validation method.
[0006] In an optional implementation, the training data includes a training set and a validation set, and the grid search and cross-validation method is combined with the training data to perform distributed training on the initial model, and obtaining the optimal model parameters can be expressed as follows:
[0007]
[0008] Among them, θ * are the optimal model parameters, k is the number of cross-validation folds, is the validation set for the i-th fold.
[0009] In an optional implementation, the first value is obtained based on the number of the sub-data sets, the number of the sub-model units, the total amount of the training data, and the total amount of the initial model using the following formula:
[0010]
[0011] Wherein, V is the first value, S is the total amount of the training data, s is the number of the sub-datasets, M is the total amount of the initial model, m is the number of the sub-model units, and Z 1 and Z 2 is the preset parameter value.
[0012] In an optional implementation, the distributed training of the sub-model units is performed based on the target parallel scheme and adopts a grid search and cross-validation method, including: based on a parameter server, distributed training is performed using the target parallel scheme and a gradient compression method.
[0013] In one implementation, the inference performance optimization method adopted for the inference performance optimization includes at least one of model quantization, model pruning and hardware acceleration, and the model compression and inference performance optimization of the first machine learning model to obtain the second machine learning model includes: model compression of the first machine learning model; obtaining the weight value corresponding to each of the inference performance optimization methods; determining the use probability of each of the inference performance optimization methods based on the weight value corresponding to each of the inference performance optimization methods; based on the use probability and the inference performance optimization method, performing inference performance optimization on the first machine learning model to obtain the second machine learning model.
[0014] In the second aspect, the present application proposes a training and deployment device for a machine learning model, which is applied to a cloud computing platform, and the device includes: a first processing module, which is used to perform distributed training on the initial model using grid search and cross-validation methods in combination with training data to obtain optimal model parameters; a second processing module, which is used to obtain a first machine learning model based on the optimal model parameters; a third processing module, which is used to perform model compression and inference performance optimization on the first machine learning model to obtain a second machine learning model; and a deployment module, which is used to deploy the second machine learning model on a cloud computing platform using a microservice architecture.
[0015] In one implementation, the first processing module is specifically used to: divide the training data into at least one sub-data set; divide the initial model into at least one sub-model unit; obtain a first value based on the number of the sub-data sets, the number of the sub-model units, the total data volume of the training data, and the total model volume of the initial model; determine a target parallel scheme based on the first value and a preset first threshold; and perform distributed training on the sub-model units based on the target parallel scheme and using grid search and cross-validation methods.
[0016] In an optional implementation, the training data includes a training set and a validation set, and the grid search and cross-validation method is combined with the training data to perform distributed training on the initial model, and obtaining the optimal model parameters can be expressed as follows:
[0017]
[0018] Among them, θ * are the optimal model parameters, k is the number of cross-validation folds, is the validation set for the i-th fold.
[0019] In an optional implementation, the first processing module uses the following formula to obtain the first value:
[0020]
[0021] Wherein, V is the first value, S is the total amount of the training data, s is the number of the sub-datasets, M is the total amount of the initial model, m is the number of the sub-model units, and Z 1 and Z 2 is the preset parameter value.
[0022] In an optional implementation, the first processing module is specifically used to: perform distributed training based on a parameter server using the target parallel scheme and gradient compression method.
[0023] In one implementation, the inference performance optimization method adopted for the inference performance optimization includes at least one of model quantization, model pruning and hardware acceleration, and the model compression and inference performance optimization of the first machine learning model to obtain the second machine learning model includes: model compression of the first machine learning model; obtaining the weight value corresponding to each of the inference performance optimization methods; determining the use probability of each of the inference performance optimization methods based on the weight value corresponding to each of the inference performance optimization methods; based on the use probability and the inference performance optimization method, performing inference performance optimization on the first machine learning model to obtain the second machine learning model.
[0024] In a third aspect, the present application proposes an electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the training and deployment method of the machine learning model as described in the first aspect.
[0025] In a fourth aspect, the present application proposes a computer-readable storage medium for storing instructions, which, when executed, enables the method described in the first aspect to be implemented.
[0026] In a fifth aspect, the present application proposes a computer program product, comprising a computer program, which, when executed by a processor, implements the steps of the method for training and deploying a machine learning model as described in the first aspect.
[0027] The training and deployment method, device, equipment and storage medium of the machine learning model provided in this application can use grid search and cross-validation methods in combination with training data to perform distributed training on the initial machine learning model to obtain optimal machine learning model parameters, so as to obtain a first machine learning model based on the optimal machine learning model parameters, and perform machine learning model compression and reasoning performance optimization on the first machine learning model to obtain a second machine learning model, and deploy the second machine learning model on a cloud computing platform using a microservice architecture. Efficient optimization and deployment of machine learning models can be achieved.
[0028] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become apparent from the description below, or will be learned through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:
[0030] Figure 1 It is a flowchart of a method for training and deploying a machine learning model provided in an embodiment of the present application;
[0031] Figure 2 is a connection diagram of a convolutional neural network provided in an embodiment of the present application;
[0032] Figure 3 It is a flowchart of another method for training and deploying a machine learning model provided in an embodiment of the present application;
[0033] Figure 4 It is a structural diagram of a training and deployment device for a machine learning model provided in an embodiment of the present application;
[0034] Figure 5 It is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0035] Embodiments of the present application are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present application, and should not be construed as limiting the present application.
[0036] The following describes the training and deployment method and device of the machine learning model of the embodiment of the present application with reference to the accompanying drawings.
[0037] Figure 1 1 is a flow chart of a method for training and deploying a machine learning model provided in an embodiment of the present application. The method can be applied to a cloud computing platform. Figure 1 As shown, the method may include but is not limited to the following steps:
[0038] Step S101: Use grid search and cross-validation methods in combination with training data to perform distributed training on the initial model to obtain optimal model parameters.
[0039] In one implementation, a grid search and cross-validation method is used in combination with training data to perform distributed training on the initial model to obtain optimal model parameters, including: dividing the training data into at least one sub-data set; dividing the initial model into at least one sub-model unit; obtaining a first value based on the number of sub-data sets, the number of sub-model units, the total amount of training data and the total amount of models of the initial model; determining a target parallel scheme based on the first value and a preset first threshold; and performing distributed training on the sub-model units based on the target parallel scheme and using a grid search and cross-validation method.
[0040] Among them, in an embodiment of the present application, the above-mentioned target parallel scheme is one of data parallelism and model parallelism.
[0041] In one implementation, the first value is obtained based on the number of sub-data sets, the number of sub-model units, the total amount of training data, and the total amount of models of the initial model using the following formula:
[0042]
[0043] Where v is the first value, s is the number of sub-datasets, S is the total amount of training data, m is the number of sub-model units, M is the total amount of models of the initial model, and Z 1 and Z 2is the preset initial value. Taking the first threshold as α as an example, if v≤α, the data parallel solution is adopted; if v>α, the model parallel solution is adopted.
[0044] It should be noted that the total model volume can be the total number of parameters of the machine learning model, that is, the total number of all parameters in the model.
[0045] In one implementation, the grid search and cross-validation methods are combined with training data to perform distributed training on the initial model, and the optimal model parameters can be expressed as follows:
[0046]
[0047] Among them, θ * is the optimal model parameter, k is the number of cross-validation folds, is the validation set for the i-th fold.
[0048] Step S102: Obtain a first machine learning model based on optimal model parameters.
[0049] In some embodiments, a CNN (Convolutional neural network) model architecture can be used to establish a first machine learning model based on optimal model parameters.
[0050] For example, the convolutional layer formula of CNN can be expressed as follows:
[0051]
[0052] The output of the pooling layer can be expressed as follows:
[0053] p=pooling(O)
[0054] in, Represents the output feature map of the lth convolutional layer at position (i, j). The value of the l-1 layer input feature map at position (i, j), is the weight of the l-layer convolution kernel at position (m,n), b l is the variance of layer l, σ is the activation function, m, n are the coordinates of the convolution kernel, and pooling represents the pooling operation. As an example, see Figure 2 , Figure 2 It is a connection diagram of a convolutional neural network provided in an embodiment of the present application.
[0055] Step S103: Perform model compression and inference performance optimization on the first machine learning model to obtain a second machine learning model.
[0056] Exemplarily, the first machine learning model is quantized and pruned to reduce the size and inference time of the model, and the inference performance of the first machine learning model after model compression is optimized to obtain a second machine learning model.
[0057] In an optional implementation, the quantization operation can be expressed as:
[0058]
[0059] Where Q is the quantization weight, W is the initial weight, S is the size of the quantization step, and Z is the quantization initial value.
[0060] In an optional implementation, the pruning operation can be expressed as:
[0061] W′=W·M
[0062] Among them, W′ is the pruned weight, W is the initial weight, and M is the pruning mask.
[0063] Step S104: Deploy the second machine learning model on the cloud computing platform using a microservice architecture.
[0064] By implementing the embodiments of the present application, the grid search and cross-validation methods can be combined with training data to perform distributed training on the initial machine learning model, and the optimal machine learning model parameters can be obtained to obtain the first machine learning model based on the optimal machine learning model parameters, and the first machine learning model can be compressed and the reasoning performance optimized to obtain the second machine learning model, and the second machine learning model can be deployed on the cloud computing platform using a microservice architecture. Efficient optimization and deployment of machine learning models can be achieved.
[0065] In some embodiments, a corresponding method may be selected from a plurality of inference performance optimization methods to optimize the inference performance of the first machine learning model. As an example, see Figure 3 , Figure 3 FIG. 1 is a flow chart of another method for training and deploying a machine learning model provided in an embodiment of the present application. The method can be applied to a cloud computing platform. Figure 3 As shown, the method may include but is not limited to the following steps:
[0066] Step S301: Use grid search and cross-validation methods in combination with training data to perform distributed training on the initial model to obtain optimal model parameters.
[0067] In the embodiments of the present application, step S301 can be implemented by any of the methods in the embodiments of the present application, and the embodiments of the present application do not limit this and will not be described in detail.
[0068] Step S302: Obtain a first machine learning model based on optimal model parameters.
[0069] In the embodiments of the present application, step S302 can be implemented by any of the methods in the embodiments of the present application, and the embodiments of the present application do not limit this and will not be described in detail.
[0070] Step S303: compress the first machine learning model.
[0071] In the embodiments of the present application, step S303 can be implemented by any of the methods in the embodiments of the present application, and the embodiments of the present application do not limit this and will not be repeated.
[0072] Step S304: Determine the use probability of each reasoning performance optimization method based on the weight value corresponding to each reasoning performance optimization method.
[0073] Among them, in an embodiment of the present application, the above-mentioned reasoning performance optimization method includes model quantization, model pruning and hardware acceleration.
[0074] Exemplarily, the following formula may be used to determine the use probability of each reasoning performance optimization method based on the weight value corresponding to each reasoning performance optimization method.
[0075]
[0076] Among them, p1, p2, p3 are the weight values corresponding to model quantization, model pruning and hardware acceleration respectively, i = 1, 2, 3, and p(n) is the usage probability corresponding to model quantization, model pruning and hardware acceleration.
[0077] As an example, if p1 is the weight value corresponding to the model quantization, then when pi=p1, the p(n) obtained is the usage probability corresponding to the model quantization.
[0078] Step S305: Based on the usage probability and the inference performance optimization method, the inference performance of the first machine learning model is optimized to obtain a second machine learning model.
[0079] Exemplarily, the inference performance optimization method with the highest probability is selected to optimize the inference performance of the first machine learning model to obtain the second machine learning model.
[0080] Step S306: Deploy the second machine learning model on the cloud computing platform using a microservice architecture.
[0081] By implementing the embodiments of the present application, the use probability of each inference performance optimization method can be obtained based on the weight value corresponding to each inference performance optimization method, so as to select the corresponding inference performance optimization method based on the use probability to optimize the inference performance of the first machine learning model to obtain the second machine learning model. Efficient optimization and deployment of machine learning models can be achieved.
[0082] See also Figure 4 , Figure 4 Schematic diagram of a training and deployment device for a machine learning model provided in an embodiment of the present application. The device can be applied to a cloud computing platform. Figure 4 As shown, the device 400 includes: a first processing module 401, which is used to perform distributed training on the initial model using grid search and cross-validation methods in combination with training data to obtain optimal model parameters; a second processing module 402, which is used to obtain a first machine learning model based on the optimal model parameters; a third processing module 403, which is used to perform model compression and inference performance optimization on the first machine learning model to obtain a second machine learning model; and a deployment module 404, which is used to deploy the second machine learning model on a cloud computing platform using a microservice architecture.
[0083] In one implementation, the first processing module 401 is specifically used to: divide the training data into at least one sub-data set; divide the initial model into at least one sub-model unit; obtain a first numerical value based on the number of sub-data sets, the number of sub-model units, the total amount of training data and the total amount of models of the initial model; determine a target parallel scheme based on the first numerical value and a preset first threshold; and perform distributed training on the sub-model units based on the target parallel scheme and using grid search and cross-validation methods.
[0084] In an optional implementation, the training data includes a training set and a validation set. The grid search and cross-validation methods are combined with the training data to perform distributed training on the initial model. The optimal model parameters can be expressed as follows:
[0085]
[0086] Among them, θ * is the optimal model parameter, k is the number of cross-validation folds, is the validation set for the i-th fold.
[0087] In an optional implementation, the first processing module 401 uses the following formula to obtain the first value:
[0088]
[0089] Where V is the first value, S is the total amount of training data, s is the number of sub-datasets, M is the total amount of the initial model, m is the number of sub-model units, and Z is the total amount of training data. 1 and Z 2 is the preset parameter value.
[0090] In an optional implementation, the first processing module 401 is specifically used to: perform distributed training based on a parameter server using a target parallel scheme and a gradient compression method.
[0091] In one implementation, the inference performance optimization method used for inference performance optimization includes at least one of model quantization, model pruning and hardware acceleration, and the third processing module 403 is specifically used to: perform model compression on the first machine learning model; obtain the weight value corresponding to each inference performance optimization method; determine the use probability of each inference performance optimization method based on the weight value corresponding to each inference performance optimization method; perform inference performance optimization on the first machine learning model based on the use probability and the inference performance optimization method to obtain the second machine learning model.
[0092] Through the device of the embodiment of the present application, the grid search and cross-validation methods can be combined with training data to perform distributed training on the initial machine learning model to obtain the optimal machine learning model parameters, and then obtain the first machine learning model based on the optimal machine learning model parameters, and perform machine learning model compression and reasoning performance optimization on the first machine learning model to obtain the second machine learning model, and then deploy the second machine learning model on the cloud computing platform using a microservice architecture. Efficient optimization and deployment of machine learning models can be achieved.
[0093] It should be noted that the aforementioned explanation of the embodiment of the training and deployment method of the machine learning model is also applicable to the training and deployment device of the machine learning model of this embodiment, and will not be repeated here.
[0094] In order to implement the above embodiment, the present application also proposes an electronic device. Figure 5 , Figure 5 Schematic diagram of the structure of the electronic device provided in the embodiment of the present application. Figure 5 As shown, the electronic device 500 includes: a processor 501, and a memory 502 communicatively connected to the processor 501; the memory 502 stores computer-executable instructions; the processor 501 executes the computer-executable instructions stored in the memory to implement the method provided in the aforementioned embodiment.
[0095] In order to implement the above embodiments, the present application also proposes a computer-readable storage medium, in which computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, they are used to implement the methods provided by the above embodiments.
[0096] In order to implement the above embodiments, the present application also proposes a computer program product, including a computer program, which implements the methods provided by the above embodiments when executed by a processor.
[0097] In the description of this application, unless otherwise specified, “ / ” means or, for example, A / B can mean A or B; “and / or” in this article is merely a way to describe the association relationship of associated objects, indicating that there can be three relationships, for example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone.
[0098] In the description of the aforementioned embodiments, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, without contradiction.
[0099] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of the features. In the description of this application, the meaning of "plurality" is at least two, such as two, three, etc., unless otherwise clearly and specifically defined.
[0100] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, fragment or portion of code comprising one or more executable instructions for implementing the steps of a custom logical function or process, and the scope of the preferred embodiments of the present application includes alternative implementations in which functions may not be performed in the order shown or discussed, including performing functions in a substantially simultaneous manner or in the reverse order depending on the functions involved, which should be understood by technicians in the technical field to which the embodiments of the present application belong.
[0101] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute the instructions), or in combination with these instruction execution systems, devices or apparatuses. For the purpose of this specification, "computer-readable medium" can be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, device or apparatus, or in combination with these instruction execution systems, devices or apparatuses. More specific examples of computer-readable media (a non-exhaustive list) include the following: an electrical connection with one or more wires (electronic device), a portable computer disk box (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), a fiber optic device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium and then editing, interpreting or processing in other suitable ways if necessary, and then stored in a computer memory.
[0102] It should be understood that the various parts of the present application can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0103] A person skilled in the art may understand that all or part of the steps in the method for implementing the above-mentioned embodiment may be completed by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, which, when executed, includes one or a combination of the steps of the method embodiment.
[0104] In addition, each functional unit in each embodiment of the present application may be integrated into a processing module, or each unit may exist physically separately, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.
[0105] The storage medium mentioned above may be a read-only memory, a magnetic disk or an optical disk, etc. Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and cannot be understood as limiting the present application. A person of ordinary skill in the art may change, modify, replace and modify the above embodiments within the scope of the present application.
Claims
1. A method for training and deploying a machine learning model, characterized in that: The method is applied to a cloud computing platform, and the method comprises: The grid search and cross-validation methods are used to combine the training data to perform distributed training on the initial model to obtain the optimal model parameters; Acquire a first machine learning model based on the optimal model parameters; Performing model compression and inference performance optimization on the first machine learning model to obtain a second machine learning model; The second machine learning model is deployed on a cloud computing platform using a microservice architecture.
2. The method according to claim 1, characterized in that The grid search and cross-validation method is used to perform distributed training on the initial model in combination with the training data to obtain the optimal model parameters, including: Dividing the training data into at least one sub-dataset; Dividing the initial model into at least one sub-model unit; Obtaining a first value based on the number of the sub-data sets, the number of the sub-model units, the total amount of data of the training data, and the total amount of models of the initial model; Determine a target parallel solution based on the first value and a preset first threshold; Based on the target parallel scheme, the sub-model units are distributedly trained by using grid search and cross-validation methods.
3. The method according to claim 2, characterized in that The training data includes a training set and a validation set. The grid search and cross-validation method is combined with the training data to perform distributed training on the initial model, and the optimal model parameters can be expressed as follows: Among them, θ * are the optimal model parameters, k is the number of cross-validation folds, is the validation set for the i-th fold.
4. The method according to claim 2, characterized in that The first value is obtained by using the following formula based on the number of the sub-data sets, the number of the sub-model units, the total amount of the training data, and the total amount of the initial model: Among them, V is the first value, S is the total data amount of the training data, s is the number of the sub-data sets, M is the total model amount of the initial model, m is the number of the sub-model units, and Z1 and z2 are preset parameter values.
5. The method according to claim 2, characterized in that The method of performing distributed training on the sub-model units based on the target parallel scheme and using a grid search and cross-validation method includes: Based on the parameter server, the target parallel scheme and gradient compression method are adopted to perform distributed training.
6. The method according to claim 1, characterized in that The inference performance optimization method adopted by the inference performance optimization includes at least one of model quantization, model pruning and hardware acceleration, and the first machine learning model is compressed and the inference performance is optimized to obtain the second machine learning model, including: Performing model compression on the first machine learning model; Obtaining a weight value corresponding to each of the inference performance optimization methods; Determining the use probability of each of the reasoning performance optimization methods based on the weight value corresponding to each of the reasoning performance optimization methods; Based on the usage probability and the reasoning performance optimization method, the reasoning performance of the first machine learning model is optimized to obtain the second machine learning model.
7. A training and deployment device for a machine learning model, characterized in that: The device is applied to a cloud computing platform and includes: The first processing module is used to perform distributed training on the initial model using a grid search and cross-validation method combined with training data to obtain optimal model parameters; A second processing module, configured to obtain a first machine learning model based on the optimal model parameters; A third processing module is used to perform model compression and inference performance optimization on the first machine learning model to obtain a second machine learning model; A deployment module is used to deploy the second machine learning model on a cloud computing platform using a microservice architecture.
8. An electronic device, characterized in that: include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory to implement the method according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are executed by a processor, they are used to implement the method according to any one of claims 1 to 6.
10. A computer program product, characterized in that The method comprises a computer program, which implements the method according to any one of claims 1 to 6 when being executed by a processor.
Citation Information
Patent Citations
Deep learning model training method and deep learning model training system
CN117669700A
Machine learning model training method and device, equipment and storage medium
CN118261222A