Large model automatic deployment method, device and equipment and storage medium

Through automated construction, deployment and testing processes, the complexity and inefficiency of large-model deployment are solved, efficient and reliable large-model automated deployment is achieved, and the complexity and error rate of manual operations are reduced.

CN120104334APending Publication Date: 2025-06-06SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510223867.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The deployment process of large models is complex, and the traditional manual deployment method is inefficient and error-prone, which increases the risk and cost of deployment.

Method used

By implementing the automated construction, deployment and testing process of the big model, detect the available GPU resources of the target server, determine the deployment estimate resources based on the preset model parameters, establish the GPU index, and modify the model loading script to automatically deploy the model, and conduct deployment completion tests.

Benefits of technology

Improve deployment efficiency and reliability, reduce the complexity and error rate of manual operations, and simplify the management and maintenance of the deployment process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120104334A_ABST
    Figure CN120104334A_ABST
Patent Text Reader

Abstract

The invention discloses a large model automatic deployment method and device, equipment and a storage medium, and relates to the technical field of computers, and the method comprises the steps: detecting current available GPU resources of a target server, determining deployment estimation resources needed by a deployment model through a preset continuous integration tool according to preset model parameters, establishing a GPU index for calling available GPU resources according to the deployment pre-estimation resources; modifying the preset model loading script according to the preset model parameters and the GPU index to obtain a target model loading script; and executing the target model loading script to deploy the target model in the target server based on the model data stored in the specified directory, and performing a deployment completion degree test on the target model after the target model is deployed. Therefore, the deployment efficiency and reliability can be improved by realizing the automatic construction, deployment and test process of the large model, and meanwhile, the complexity and error rate of manual operation are reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to a large model automatic deployment method, device, equipment and storage medium. Background Art

[0002] With the rapid development of artificial intelligence technology, large models have been widely used in natural language processing, computer vision and other fields, promoting the innovation and progress of related technologies. However, the deployment process of large models is complicated, involving multiple links such as code construction, model loading and reasoning, test verification and online deployment, each of which requires a lot of time and effort. The traditional manual deployment method is not only inefficient, but also prone to errors, increasing the risk and cost of deployment. Summary of the invention

[0003] In view of this, the purpose of the present invention is to provide a large model automatic deployment method, device, equipment and storage medium, which can improve the efficiency and reliability of deployment by realizing the automatic construction, deployment and testing process of the large model, while reducing the complexity and error rate of manual operation. The specific scheme is as follows:

[0004] In a first aspect, the present application discloses a large model automatic deployment method, comprising:

[0005] Detect the currently available GPU resources of the target server, and determine the deployment estimated resources required for the deployment model according to the preset model parameters through the preset continuous integration tool, and then establish a GPU index for calling the available GPU resources according to the deployment estimated resources;

[0006] Modifying the preset model loading script according to the preset model parameters and the GPU index to obtain a target model loading script;

[0007] The target model loading script is executed to deploy the target model on the target server based on the model data stored in the specified directory, and a deployment completion test is performed on the target model after the target model is deployed.

[0008] Optionally, before detecting the currently available GPU resources of the target server, determining the deployment estimated resources required for the deployment model according to preset model parameters through a preset continuous integration tool, and then establishing a GPU index for calling the available GPU resources according to the deployment estimated resources, the method further includes:

[0009] Execute a preset detection script to determine whether a preset port of the deployment model is occupied, and if the preset port is occupied, retrieve the name of the process occupying the preset port;

[0010] Based on the name, the corresponding target process is located, and the occupation of the preset port by the target process is released, and the model data downloaded by the preset continuous integration tool is saved to the designated directory of the target server through the preset security protocol.

[0011] Optionally, the large model automatic deployment method further includes:

[0012] Save the target model loading script;

[0013] If the target model deployment fails, port detection is performed to determine whether the preset port has been unoccupied;

[0014] If the preset port is not unoccupied, the target model's current occupation of the preset port is unoccupied; if the preset port is unoccupied, the target model loading script is re-executed to redeploy the target model.

[0015] Optionally, the currently available GPU resources of the target server are detected, and the estimated deployment resources required for the deployment model are determined according to the preset model parameters by a preset continuous integration tool, and then a GPU index for calling the available GPU resources is established according to the estimated deployment resources, including:

[0016] Perform resource detection on the target server to determine the current available GPU resources of the target server;

[0017] Determine the estimated deployment resources required for the deployment model through a preset continuous integration tool according to the preset tensor parallel size in the preset model parameters and the estimated video memory requirement corresponding to the model loading;

[0018] A target available GPU resource corresponding to the deployment estimated resource is selected from the available GPU resources, and a GPU index corresponding to the target available GPU resource is established.

[0019] Optionally, modifying the preset model loading script according to the preset model parameters and the GPU index to obtain a target model loading script includes:

[0020] Replacing the model parameters in the preset model loading script based on the preset model parameters to obtain a modified model loading script;

[0021] The GPU index is embedded into the modified model loading script to obtain a target model loading script.

[0022] Optionally, after the target model is deployed, the step of performing a deployment completion test on the target model includes:

[0023] sending a test request to the target model to input simulation data to the target model based on the test request;

[0024] A deployment completion test is performed on the target model based on the simulation data to determine the model performance of the target model, and the obtained test results are printed out.

[0025] Optionally, the executing the target model loading script to deploy the target model on the target server based on the model data stored in the specified directory, and after the target model is deployed, performing a deployment completion test on the target model, further includes:

[0026] If the target model passes the deployment completion test, providing model services through the target model and monitoring the running status of the target model;

[0027] If the operating state is abnormal, an abnormal alarm is issued and the target model is restarted.

[0028] In a second aspect, the present application discloses a large model automatic deployment device, comprising:

[0029] An index building module is used to detect the currently available GPU resources of the target server, and determine the deployment estimated resources required for the deployment model according to the preset model parameters through a preset continuous integration tool, and then build a GPU index for calling the available GPU resources according to the deployment estimated resources;

[0030] A script modification module, used to modify the preset model loading script according to the preset model parameters and the GPU index to obtain a target model loading script;

[0031] The model deployment module is used to execute the target model loading script to deploy the target model on the target server based on the model data stored in the specified directory, and to perform a deployment completion test on the target model after the target model is deployed.

[0032] In a third aspect, the present application discloses an electronic device, comprising:

[0033] Memory, used to store computer programs;

[0034] A processor is used to execute the computer program to implement the large model automatic deployment method as mentioned above.

[0035] In a fourth aspect, the present application discloses a computer-readable storage medium for storing a computer program, wherein the computer program, when executed by a processor, implements the aforementioned large model automatic deployment method.

[0036] In the present application, the method of the present application can detect the current available GPU resources of the target server, and determine the deployment estimated resources required for the deployment model according to the preset model parameters through the preset continuous integration tool, and then establish a GPU index for calling the available GPU resources according to the deployment estimated resources; modify the preset model loading script according to the preset model parameters and the GPU index to obtain the target model loading script; execute the target model loading script to deploy the target model on the target server based on the model data saved in the specified directory, and perform a deployment completion test on the target model after the deployment of the target model is completed.

[0037] It can be seen that through the method of the present application, after detecting the current available GPU resources of the target server and determining the deployment estimated resources required for the deployment model according to the preset model parameters, a GPU index for calling available GPU resources can be established according to the deployment estimated resources, and then the preset model loading script can be modified according to the preset model parameters and the GPU index, and then the modified target model loading script is executed to deploy the target model on the target server based on the model data saved in the specified directory, and the model is tested for deployment completion after deployment. In this way, the efficiency and reliability of deployment can be improved by realizing the automated construction, deployment and testing process of large models, while reducing the complexity and error rate of manual operations. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.

[0039] Figure 1 A flow chart of a large model automatic deployment method disclosed in this application;

[0040] Figure 2 A flowchart of a specific large model automatic deployment method disclosed in this application;

[0041] Figure 3 A schematic diagram of a large model automated deployment architecture disclosed in this application;

[0042] Figure 4 This is a schematic diagram of the structure of a large-scale model automatic deployment device disclosed in this application;

[0043] Figure 5 This is a structural diagram of an electronic device disclosed in this application. DETAILED DESCRIPTION

[0044] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0045] In the existing technology, the deployment process of large models is complicated and each link requires a lot of time and energy investment. In addition, the traditional manual deployment method is not only inefficient but also prone to errors, which increases the risk and cost of deployment.

[0046] In order to overcome the above-mentioned technical problems, the present application discloses a large model automatic deployment method, device, equipment and storage medium, which can improve the efficiency and reliability of deployment by realizing the automatic construction, deployment and testing process of large models, while reducing the complexity and error rate of manual operations.

[0047] The large model automated deployment method in this application is implemented by setting up a deployment script, and the deployment script includes a preset detection script and a preset model loading script. All of its automated processes are implemented by running the deployment script, and the deployment script is designed through Jenkins Pipeline.

[0048] See also Figure 1 As shown, an embodiment of the present invention discloses a large model automatic deployment method, including:

[0049] Step S11, detecting the currently available GPU resources of the target server, and determining the deployment estimated resources required for the deployment model according to the preset model parameters through the preset continuous integration tool, and then establishing a GPU index for calling the available GPU resources according to the deployment estimated resources.

[0050] In this embodiment, before detecting the current available resources of the target server, it is necessary to determine whether the preset port of the deployment model is occupied by executing a preset detection script. In the case where the preset port is occupied, it is necessary to retrieve the process name occupying the preset port, locate the corresponding target process based on the name, and release the occupation of the preset port by the target process. Then, the latest code and model files downloaded through Jenkins are efficiently and securely transferred to the specified directory of the target server through a preset security protocol, such as the SCP (Secure Copy Protocol) protocol. It should be noted that when Jenkins downloads code and model files, it is necessary to download the model code from the preset Git repository and download the model file from the preset OSS (Object Storage Service) bucket. However, since pulling dependencies through official sources may encounter network stability problems, in order to avoid obtaining code and model files in a single way, the preset code and model files can also be saved in a private server or private warehouse to achieve persistent storage, and each version of the code and model file can be saved in a private server or private warehouse to facilitate model rollback. In this way, on the one hand, resources can be released for the deployment and startup of the target model to avoid model deployment failure or operation failure due to insufficient memory resources; on the other hand, by setting up a private server to save code and model files, data acquisition failure due to network instability can be avoided, and the network consumption for each acquisition of code and model files can be reduced.

[0051] It should be noted that after confirming the occupancy of the preset port, the target server can be monitored for resources to determine the current available GPU (Graphics Processing Unit) resources of the target server, and then the preset continuous integration tool can be used to determine the deployment estimated resources required for the deployment model according to the preset tensor parallel size in the preset model parameters and the estimated video memory requirement corresponding to the model loading, wherein the preset tensor parallel size is the model tensor-parallel-size, and the preset continuous integration tool is Jenkins. Further, the target available GPU resources corresponding to the deployment estimated resources can be selected from the available GPU resources. For example, the available GPU resources in the current target server are GPU0, GPU1, and GPU2, but according to the deployment estimated resources, only two processors are required for model deployment, so GPU0 and GPU1 can be selected for model deployment, so GPU0 and GPU1 can be used as target available GPU resources, and then a GPU index corresponding to the target available GPU resource can be established, that is, a GPU index corresponding to GPU0 and GPU1 can be established.

[0052] In this way, by detecting and reasonably allocating GPU resources, we can ensure that large models can fully utilize hardware resources for efficient reasoning, thereby improving resource utilization and model performance. In addition, through the collaborative work of automated scripts and JenkinsPipeline, manual intervention is reduced, the deployment efficiency of large models is significantly improved, and the cycle from development to launch is shortened.

[0053] Step S12: modifying the preset model loading script according to the preset model parameters and the GPU index to obtain a target model loading script.

[0054] In this embodiment, the preset model loading script needs to be modified by the preset model parameters built by the preset continuous integration tool Jenkins to obtain the modified model loading script, and then the constructed GPU index needs to be embedded in the modified model loading script to obtain the target model loading script, so that the corresponding GPU can be directly called by the index when executing the target model loading script. It should be noted that the modification of script parameters requires the use of Jenkins continuous integration (CI) and continuous deployment (CD) mechanisms so that code changes can be frequently submitted to the shared version control repository. After each code submission is triggered, Jenkins automatically executes the build and automated testing process, thereby identifying and repairing defects in the integration process as early as possible, reducing the complexity of the problem, thereby improving code quality and accelerating development iterations.

[0055] It should be further explained that before model deployment, the target model loading script needs to be saved so that it can be quickly restored to a stable state when the model deployment fails or needs to be rolled back. Then the model is deployed, and if the model deployment fails, the port detection is performed again to determine whether the preset port has been unoccupied. If the preset port has not been unoccupied, the target model's current occupation of the preset port is released. If the preset port has been unoccupied, the target model loading script is re-executed according to the saved target model loading script to redeploy the target model. In this way, the complexity of manual operations is reduced through automated deployment and rollback mechanisms, the management and maintenance of the deployment process is simplified, and the operation and maintenance costs are reduced.

[0056] Step S13: execute the target model loading script to deploy the target model on the target server based on the model data stored in the specified directory, and perform a deployment completion test on the target model after the target model is deployed.

[0057] In this embodiment, after completing the script modification, the target model loading script needs to be executed. Then the saved model data can be obtained from the specified directory, and the model data includes the model file and the parameters required for deploying the model. Then the model can be deployed based on the obtained model data, and the deployment completion test can be performed after the deployment is completed. Specifically, a test request can be sent to the deployed model, and then the model loading function needs to be executed. The model service needs to be started according to the model startup command and the corresponding model parameters. Then, simulation data needs to be input to the target model to simulate the data input of the actual application scenario, so as to verify the performance indicators such as the model's reasoning accuracy and response speed according to the model's processing of the data. If all the performance indicators obtained pass the corresponding preset performance indicator thresholds, it means that the target model has passed the deployment completion test. Further, after the test is completed, in order to provide an intuitive basis for evaluating the model deployment effect, the test results need to be printed out to provide an intuitive basis for evaluating the model deployment effect. In this way, by testing the performance of the model line, it can be ensured that the model provides a complete model service when it goes online, and the printing of the model test results helps non-technical personnel understand the model performance, and can easily track the performance changes of the model at different time points, which is helpful for long-term analysis and improvement.

[0058] It can be seen that through the method of the present application, after detecting the current available GPU resources of the target server and determining the deployment estimated resources required for the deployment model according to the preset model parameters, a GPU index for calling available GPU resources can be established according to the deployment estimated resources, and then the preset model loading script can be modified according to the preset model parameters and the GPU index, and then the modified target model loading script is executed to deploy the target model on the target server based on the model data saved in the specified directory, and the model is tested for deployment completion after deployment. In this way, the efficiency and reliability of deployment can be improved by realizing the automated construction, deployment and testing process of large models, while reducing the complexity and error rate of manual operations.

[0059] Based on the above embodiments, it can be seen that the method disclosed in this application can realize automatic model deployment. However, after the model is successfully deployed, the running status of the model needs to be monitored. Therefore, this embodiment describes in detail how to monitor the model status. Figure 2 As shown, an embodiment of the present invention discloses a large model automatic deployment method, including:

[0060] Step S21, detect the current available GPU resources of the target server, and determine the deployment estimated resources required for the deployment model according to the preset model parameters through the preset continuous integration tool, and then establish a GPU index for calling the available GPU resources according to the deployment estimated resources.

[0061] Step S22: modify the preset model loading script according to the preset model parameters and the GPU index to obtain a target model loading script.

[0062] Step S23: execute the target model loading script to deploy the target model on the target server based on the model data stored in the specified directory, and perform a deployment completion test on the target model after the target model is deployed.

[0063] Step S24: If the target model passes the deployment completion test, model services are provided through the target model, and the running status of the target model is monitored.

[0064] In this embodiment, if the deployment of the target model is completed on the target server, and the target model has passed the deployment completion test, the model service can be provided through the target model, and the running status of the target model can be monitored. Specifically, after the model service of the target model is started, it is necessary to monitor the running status of the target model in real time, covering indicators such as port monitoring conditions and service response speed, to ensure the normal and stable operation of the service. Further, it is necessary to log the target model, including various links such as file transfer, port detection, resource allocation, and service startup. In this way, the recorded log information can provide a basis for subsequent troubleshooting, and the performance of the model can be deeply analyzed through the recorded log information, which helps developers optimize the model structure and deployment strategy, and further improves the reasoning efficiency and accuracy of the model.

[0065] Step S25: If the operating state is abnormal, an abnormal alarm is issued and the target model is restarted.

[0066] In this embodiment, if a service anomaly is found in the model when monitoring the running status of the model, the alarm mechanism in the script needs to be triggered to notify relevant personnel to perform model maintenance, and then try to automatically restart the model to ensure the high availability of the model service. Furthermore, if the model anomaly is serious, it is necessary to roll back the model according to the saved target model loading script, restore the model to the historical version corresponding to the target model loading script, and then monitor the service startup status in real time to ensure the normal operation of the model service.

[0067] It can be seen that in this embodiment, after the target model is loaded, if the target model passes the deployment completion test, the model service is provided through the target model, and the running status of the target model is monitored. If the running status is abnormal, an abnormal alarm is issued and the target model is restarted. In this way, the recorded log information can provide a basis for subsequent troubleshooting, and in-depth analysis of the performance of the model can help developers optimize the model structure and deployment strategy. The model can also automatically roll back the model, which helps to ensure the stable and continuous normal operation of the model service and effectively improve the user experience.

[0068] See also Figure 3 As shown, it is an architecture for automatic deployment of a large model disclosed in this application, in which it is necessary to build an automatic deployment script for the large model by using the Pipeline Script syntax of Jenkins. Then, the model code is downloaded from the preset Git repository through the Checkout step, and the model file is downloaded from the preset OSS (Object Storage Service) bucket, and placed in the Jenkins workspace so that the subsequent steps such as building, testing and deployment can use these codes and files. Then, the model needs to be deployed in the Deploy stage, and the model is deployed to the actual application environment of the server. Specifically, it is necessary to first accurately determine the deployment requirements of the large model according to the configuration parameters of the project construction, such as the reasoning framework type, the model deployment environment and the model startup parameters, such as ports, version numbers, etc. Then, the SCP protocol is used to log in to the target server, and a series of pre-written automatic deployment scripts are executed to realize the automatic deployment of the model, and the executed automatic script can realize the functions of port detection, process management, file backup, resource configuration, model service startup, daily monitoring, model testing and automatic rollback as described in the above embodiments.

[0069] See also Figure 4 As shown, the embodiment of the present invention discloses a large model automatic deployment device, including:

[0070] An index building module 11 is used to detect the currently available GPU resources of the target server, and determine the deployment estimated resources required for the deployment model according to the preset model parameters through a preset continuous integration tool, and then build a GPU index for calling the available GPU resources according to the deployment estimated resources;

[0071] A script modification module 12, used to modify the preset model loading script according to the preset model parameters and the GPU index to obtain a target model loading script;

[0072] The model deployment module 13 is used to execute the target model loading script to deploy the target model on the target server based on the model data stored in the specified directory, and to perform a deployment completion test on the target model after the target model is deployed.

[0073] In this embodiment, the method of the present application can detect the current available GPU resources of the target server, and determine the deployment estimated resources required for the deployment model according to the preset model parameters through the preset continuous integration tool, and then establish a GPU index that calls the available GPU resources according to the deployment estimated resources; modify the preset model loading script according to the preset model parameters and the GPU index to obtain the target model loading script; execute the target model loading script to deploy the target model on the target server based on the model data stored in the specified directory, and perform a deployment completion test on the target model after the target model is deployed. It can be seen that through the method of the present application, after the current available GPU resources of the target server are detected and the deployment estimated resources required for the deployment model are determined according to the preset model parameters, a GPU index that calls the available GPU resources can be established according to the deployment estimated resources, and then the preset model loading script can be modified according to the preset model parameters and the GPU index, and then the target model loading script obtained after the modification is executed to deploy the target model on the target server based on the model data stored in the specified directory, and the model is deployed after the deployment is completed. The completion test. In this way, the efficiency and reliability of deployment can be improved by realizing the automated construction, deployment and testing process of large models, while reducing the complexity and error rate of manual operations.

[0074] In some embodiments, the large model automatic deployment device may further include:

[0075] A port occupancy judgment unit, used to execute a preset detection script to determine whether a preset port of the deployment model is occupied, and if the preset port is occupied, retrieve the name of the process occupying the preset port;

[0076] A data storage unit is used to locate the corresponding target process based on the name, release the occupation of the preset port by the target process, and save the model data downloaded by the preset continuous integration tool to the designated directory of the target server through a preset security protocol.

[0077] In some embodiments, the large model automatic deployment device may further include:

[0078] A script saving unit, used for saving the target model loading script;

[0079] an occupancy release judgment unit, configured to perform port detection to determine whether the preset port has been released if the target model deployment fails;

[0080] The model deployment unit is used to release the occupation of the preset port by the target model if the preset port has not been released, and to re-execute the target model loading script to redeploy the target model if the preset port has been released.

[0081] In some embodiments, the index building module 11 may specifically include:

[0082] A resource detection unit, used to perform resource detection on a target server to determine currently available GPU resources of the target server;

[0083] A deployment resource estimation unit, used to determine the deployment estimated resources required for the deployment model according to the preset tensor parallel size in the preset model parameters and the estimated video memory requirement corresponding to the model loading through a preset continuous integration tool;

[0084] An index establishing unit is used to select a target available GPU resource corresponding to the deployment estimated resource from the available GPU resources, and establish a GPU index corresponding to the target available GPU resource.

[0085] In some embodiments, the script modification module 12 may specifically include:

[0086] A script modification unit, configured to replace the model parameters in the preset model loading script based on the preset model parameters to obtain a modified model loading script;

[0087] An index embedding unit is used to embed the GPU index into the modified model loading script to obtain a target model loading script.

[0088] In some embodiments, the model deployment module 13 may specifically include:

[0089] a data input unit, configured to send a test request to the target model, so as to input simulation data to the target model based on the test request;

[0090] A performance testing unit is used to perform a deployment completion test on the target model based on the simulation data to determine the model performance of the target model, and print out the obtained test results.

[0091] In some embodiments, the large model automatic deployment device may further include:

[0092] An operation status monitoring module, used for providing model services through the target model and monitoring the operation status of the target model if the target model passes the deployment completion test;

[0093] The model restart module is used to issue an abnormality alarm and restart the target model if the operating state is abnormal.

[0094] Furthermore, the present application also discloses an electronic device. Figure 5 This is a structural diagram of an electronic device 20 according to an exemplary embodiment. The content in the diagram cannot be regarded as any limitation on the scope of use of the present application.

[0095] Figure 5 A schematic diagram of the structure of an electronic device 20 provided in an embodiment of the present application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 is used to store a computer program, which is loaded and executed by the processor 21 to implement the relevant steps in the large model automatic deployment method disclosed in any of the aforementioned embodiments. In addition, the electronic device 20 in this embodiment may specifically be an electronic computer.

[0096] In this embodiment, the power supply 23 is used to provide working voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and the external device, and the communication protocol it follows is any communication protocol that can be applied to the technical solution of the present application, and is not specifically limited here; the input and output interface 25 is used to obtain external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs and is not specifically limited here.

[0097] In addition, the memory 22 as a carrier for storing resources may be a read-only memory, a random access memory, a disk or an optical disk, etc. The resources stored thereon may include an operating system 221, a computer program 222, etc., and the storage method may be temporary storage or permanent storage.

[0098] The operating system 221 is used to manage and control the hardware devices and computer program 222 on the electronic device 20, which can be Windows Server, Netware, Unix, Linux, etc. In addition to including a computer program that can be used to complete the large model automatic deployment method performed by the electronic device 20 disclosed in any of the aforementioned embodiments, the computer program 222 can further include a computer program that can be used to complete other specific tasks.

[0099] Furthermore, the present application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, the aforementioned large model automatic deployment method is implemented. The specific steps of the method can refer to the corresponding contents disclosed in the aforementioned embodiments, and will not be repeated here.

[0100] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.

[0101] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in the above description according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0102] The steps of the method or algorithm described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.

[0103] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.

[0104] The technical solution provided by the present application is introduced in detail above. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea. At the same time, for general technicians in this field, according to the idea of ​​the present application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.

Claims

1. A large model automatic deployment method, characterized in that: include: Detect the currently available GPU resources of the target server, and determine the deployment estimated resources required for the deployment model according to the preset model parameters through the preset continuous integration tool, and then establish a GPU index for calling the available GPU resources according to the deployment estimated resources; Modifying the preset model loading script according to the preset model parameters and the GPU index to obtain a target model loading script; The target model loading script is executed to deploy the target model on the target server based on the model data stored in the specified directory, and a deployment completion test is performed on the target model after the target model is deployed.

2. The large model automatic deployment method according to claim 1 is characterized in that: The method further includes: detecting the currently available GPU resources of the target server, determining the estimated deployment resources required for the deployment model according to the preset model parameters through the preset continuous integration tool, and then establishing a GPU index for calling the available GPU resources according to the estimated deployment resources. Execute a preset detection script to determine whether a preset port of the deployment model is occupied, and if the preset port is occupied, retrieve the name of the process occupying the preset port; Based on the name, the corresponding target process is located, and the occupation of the preset port by the target process is released, and the model data downloaded by the preset continuous integration tool is saved to the designated directory of the target server through the preset security protocol.

3. The large model automatic deployment method according to claim 2 is characterized in that: Also includes: Save the target model loading script; If the target model deployment fails, port detection is performed to determine whether the preset port has been unoccupied; If the preset port is not unoccupied, the target model's current occupation of the preset port is unoccupied; if the preset port is unoccupied, the target model loading script is re-executed to redeploy the target model.

4. The large model automatic deployment method according to claim 1 is characterized in that: The currently available GPU resources of the target server are detected, and the estimated deployment resources required for the deployment model are determined according to the preset model parameters by a preset continuous integration tool, and then a GPU index for calling the available GPU resources is established according to the estimated deployment resources, including Perform resource detection on the target server to determine the current available GPU resources of the target server; Determine the estimated deployment resources required for the deployment model through a preset continuous integration tool according to the preset tensor parallel size in the preset model parameters and the estimated video memory requirement corresponding to the model loading; A target available GPU resource corresponding to the deployment estimated resource is selected from the available GPU resources, and a GPU index corresponding to the target available GPU resource is established.

5. The large model automatic deployment method according to claim 1 is characterized in that: The step of modifying the preset model loading script according to the preset model parameters and the GPU index to obtain a target model loading script includes: Replacing the model parameters in the preset model loading script based on the preset model parameters to obtain a modified model loading script; The GPU index is embedded into the modified model loading script to obtain a target model loading script.

6. The large model automatic deployment method according to claim 1, characterized in that: After the target model is deployed, the deployment completion test of the target model includes: sending a test request to the target model to input simulation data to the target model based on the test request; A deployment completion test is performed on the target model based on the simulation data to determine the model performance of the target model, and the obtained test results are printed out.

7. The large model automatic deployment method according to any one of claims 1 to 6, characterized in that: The method of executing the target model loading script to deploy the target model on the target server based on the model data stored in the specified directory, and after the target model is deployed, performing a deployment completion test on the target model, further includes: If the target model passes the deployment completion test, providing model services through the target model and monitoring the running status of the target model; If the operating state is abnormal, an abnormal alarm is issued and the target model is restarted.

8. A large model automatic deployment device, characterized in that: include: An index building module is used to detect the currently available GPU resources of the target server, and determine the deployment estimated resources required for the deployment model according to the preset model parameters through the preset continuous integration tool, and then build a GPU index for calling the available GPU resources according to the deployment estimated resources; A script modification module, used to modify the preset model loading script according to the preset model parameters and the GPU index to obtain a target model loading script; The model deployment module is used to execute the target model loading script to deploy the target model on the target server based on the model data stored in the specified directory, and to perform a deployment completion test on the target model after the target model is deployed.

9. An electronic device, characterized in that: include: Memory, used to store computer programs; A processor is used to execute the computer program to implement the large model automatic deployment method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: Used to store a computer program, wherein when the computer program is executed by a processor, the large model automatic deployment method as described in any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • A neural network model loading implementation method and device and medium

    CN122547416A