Large model automatic deployment method and device, equipment and medium

By automatically determining the model scheduling information and building directional acyclic graph tasks, the efficient and automated deployment of large models on terminal devices is achieved, and the problem of low deployment efficiency and quality caused by differences in terminal device performance is solved.

CN120234016AActive Publication Date: 2025-07-01PENG CHENG LAB
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510711093.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-07-01
Estimated Expiration
2045-05-29

Smart Images

  • Figure CN120234016A_ABST
    Figure CN120234016A_ABST
Patent Text Reader

Abstract

The invention provides a large model automatic deployment method and device, equipment and a medium, by executing the large model automatic deployment method, an original large model, a lightweight large model, a target equipment type of target equipment and a target model architecture capable of running on the target equipment can be determined based on input of a user side; and in order to adapt to target equipment with different performance differences, scheduling information for the pre-registered model conversion service is automatically generated, and the model distillation service, the model distillation service and the model deployment service are pre-registered, so that manual intervention is not needed when the directed acyclic graph task information is generated, and the efficiency of generating the directed acyclic graph task information is improved. And the execution sequence of the services is restrained through the directed acyclic graph task information, after the directed acyclic graph task information is executed, the services can be scheduled in sequence, operations such as distillation and conversion can be automatically performed on the large model, and finally, the converted lightweight large model is deployed on the target equipment, so that the efficiency is improved. And the efficiency and the quality of large model deployment can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of large model deployment, and in particular, to a method, device, equipment and medium for automatically deploying large models. Background Art

[0002] With the development of artificial intelligence (AI) technology, large language models (LLMs) have achieved remarkable performance in fields such as natural language processing and computer vision. Large language models are also simply referred to as large models. During the deployment process of large models, huge amounts of parameters and complex computing requirements are needed. How to quickly deploy large models is crucial for the popularization and application of large models.

[0003] In related technologies, due to the limited performance of terminal devices, different degrees of adaptation are required during the process of deploying large models to terminal devices. However, due to the large performance differences among different terminal devices, it is often necessary to rely on manual operations for large model deployment according to the characteristics of each terminal device, which greatly reduces the efficiency and quality of large model deployment. Summary of the Invention

[0004] The main purpose of the embodiments of the present disclosure is to propose a method, device, equipment and medium for automatically deploying large models, which can improve the efficiency and quality of large model deployment.

[0005] To achieve the above object, the first aspect of the embodiments of the present disclosure proposes a method for automatically deploying a large model, including: Obtaining downstream task information input by a user side, and determining an original large model, a lightweight large model, a target device type of a target device, and a target model architecture that can run on the target device based on the downstream task information; When at least one of the following conditions is met: the distillation device type during the distillation of the original large model to the lightweight large model is inconsistent with the target device type, and the original model architecture of the lightweight large model is inconsistent with the target model architecture, generating scheduling information for a pre-registered model conversion service; Constructing scheduling information for a model distillation service for distilling the original large model to the lightweight large model pre-registered, and constructing scheduling information for a model deployment service for the pre-registered target device, and constructing directed acyclic graph task information based on the order of all the scheduling information; When the directed acyclic graph task information runs, based on each of the scheduling information, first schedule the model distillation service to distill the original large model onto the lightweight large model, then schedule the model conversion service to convert the distilled lightweight large model to the target device type or the target model architecture, and schedule the model deployment service to deploy the converted lightweight large model to the target device.

[0006] In some embodiments, constructing the directed acyclic graph task information based on the order of all the scheduling information includes: Configure the scheduling dependency relationships corresponding to adjacent services between the scheduling information of the model distillation service, the model conversion service, and the model deployment service in sequence, and construct the directed acyclic graph task information based on the order of all the scheduling information and the corresponding scheduling dependency relationships; The "when the directed acyclic graph task information runs, based on each of the scheduling information, first schedule the model distillation service to distill the original large model onto the lightweight large model, then schedule the model conversion service to convert the distilled lightweight large model to the target device type or the target model architecture, and schedule the model deployment service to deploy the converted lightweight large model to the target device" includes: When the directed acyclic graph task information runs, schedule the model distillation service, the model conversion service, and the model deployment service in sequence based on each of the scheduling information; When running the model distillation service, distill the original large model onto the lightweight large model; When running the model conversion service, determine whether the dependent model distillation service is completed, and after the model distillation service is completed, schedule the model conversion service to convert the distilled lightweight large model to the target device type or the target model architecture; When running the model deployment service, determine whether the dependent model conversion service is completed, and after the model conversion service is completed, schedule the model deployment service to deploy the converted lightweight large model to the target device.

[0007] In some embodiments, the "schedule the model distillation service to distill the original large model onto the lightweight large model, then schedule the model conversion service to convert the distilled lightweight large model to the target device type or the target model architecture, and schedule the model deployment service to deploy the converted lightweight large model to the target device" includes: After scheduling the model distillation service, a distillation container is created, the original large model and the lightweight large model are mounted on the distillation container, and the original large model is distilled onto the lightweight large model within the distillation container; After scheduling the model conversion service, a conversion image is installed in the distillation container, and the distilled lightweight large model is converted to the target device type or the target model architecture through the conversion image; After scheduling the model deployment service, the converted lightweight large model in the distillation container is deployed to the target device.

[0008] In some embodiments, determining the original large model, the lightweight large model, the target device type of the target device, and the target model architecture that can run on the target device based on the downstream task information includes: Extracting the target device and the task description that needs to be executed in the target device from the downstream task information; Obtaining the device description, the target device type, and the target model architecture that can run on the target device from a preset database; Extracting the first text feature of the task description and the second text feature of the device description, and fusing the first text feature and the second text feature to obtain a fusion feature; Predicting a plurality of preset candidate large models based on the fusion feature to obtain a prediction result, and selecting the original large model and the lightweight large model from the plurality of candidate large models based on the prediction result.

[0009] In some embodiments, predicting a plurality of preset candidate large models based on the fusion feature to obtain a prediction result, and selecting the original large model and the lightweight large model from the plurality of candidate large models based on the prediction result includes: Inputting the fusion feature into a preset first prediction model to predict the large models that meet the requirements of the task description among a plurality of candidate large models, obtaining a first prediction result, and selecting the original large model from the plurality of candidate large models based on the first prediction result; Inputting the fusion feature into a preset second prediction model to predict the large models that meet the requirements of the device description among the plurality of candidate large models, obtaining a second prediction result, and selecting the lightweight large model from the plurality of candidate large models based on the second prediction result.

[0010] In some embodiments, scheduling the model distillation service to distill the original large model onto the lightweight large model includes: Determining a distillation dataset based on the downstream task information; Schedule the model distillation service, and distill the original large model onto the lightweight model through the distillation dataset.

[0011] In some embodiments, the large model automatic deployment method further includes: When the distillation device type during the distillation of the original large model onto the lightweight model is consistent with the target device type, and the original model architecture of the lightweight model is consistent with the target model architecture, construct directed acyclic graph task information based on the order of the scheduling information of the model distillation service and the scheduling information of the model deployment service; When the directed acyclic graph task information runs, first schedule the model distillation service to distill the original large model onto the lightweight model based on each piece of the scheduling information, and schedule the model deployment service to deploy the distilled lightweight model onto the target device.

[0012] To achieve the above object, a second aspect of the embodiments of the present disclosure provides a large model automatic deployment device, including: An information receiving module, configured to obtain downstream task information input by a user terminal, and determine an original large model, a lightweight model, a target device type of the target device, and a target model architecture that can run on the target device based on the downstream task information; A scheduling confirmation module, configured to generate scheduling information for a pre-registered model conversion service when at least one of the following conditions is met: the distillation device type during the distillation of the original large model onto the lightweight model is inconsistent with the target device type, and the original model architecture of the lightweight model is inconsistent with the target model architecture; A directed acyclic graph construction module, configured to construct scheduling information for the model distillation service for distilling the original large model onto the lightweight model that is pre-registered, and construct scheduling information for the model deployment service for the pre-registered target device, and construct directed acyclic graph task information based on the order of all the scheduling information; An automatic deployment module, configured to, when the directed acyclic graph task information runs, first schedule the model distillation service to distill the original large model onto the lightweight model based on each piece of the scheduling information, then schedule the model conversion service to convert the distilled lightweight model to the target device type or the target model architecture, and schedule the model deployment service to deploy the converted lightweight model onto the target device.

[0013] To achieve the above object, a third aspect of the embodiments of the present disclosure provides an electronic device, the electronic device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, it implements the large model automatic deployment method described in the first aspect embodiments above.

[0014] To achieve the above object, a fourth aspect of the embodiments of the present disclosure provides a storage medium, which is a computer-readable storage medium. The storage medium stores a computer program, and when the computer program is executed by a processor, it implements the large model automatic deployment method described in the first aspect embodiments above.

[0015] The large model automatic deployment method, device, equipment and medium proposed by the embodiments of the present disclosure. The large model automatic deployment method can be applied in a large model automatic deployment device. By executing the large model automatic deployment method, the original large model, lightweight large model, target device type of the target device, and target model architecture that can run on the target device can be determined based on the input of the user side. And to adapt to target devices with different performance differences, scheduling information for the pre-registered model conversion service is automatically generated. Since the model distillation service, model distillation service, and model deployment service are all pre-registered, no manual intervention is required when generating the directed acyclic graph task information. And by the directed acyclic graph task information, the execution order between each service is constrained. After executing the directed acyclic graph task information, each service can be scheduled in turn to automatically perform operations such as distilling and converting the large model, and finally deploy the converted lightweight large model to the target device, which can improve the efficiency and quality of large model deployment. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 is a schematic diagram of an implementation environment provided by the embodiments of the present disclosure; Figure 2 is a flowchart of the large model automatic deployment method provided by the embodiments of the present disclosure; Figure 3 is Figure 2 a further flowchart included in step S104 in Figure 4 is Figure 2 another flowchart included in step S104 in Figure 5 is Figure 2 a further flowchart included in step S101 in Figure 6 is Figure 5 a further flowchart included in step S404 in Figure 7 is Figure 2 another flowchart included in step S104 in Figure 8 is another flowchart of the large model automatic deployment method provided by the embodiments of the present disclosure; Figure 9It is a schematic diagram of the large model automatic deployment system provided by the embodiments of the present disclosure; Figure 10 It is a schematic diagram of constructing acyclic graph task information provided by the embodiments of the present disclosure; Figure 11 It is a schematic diagram of the large model automatic deployment process provided by the embodiments of the present disclosure; Figure 12 It is a schematic diagram of the functional modules of the large model automatic deployment device provided by the embodiments of the present disclosure; Figure 13 It is a schematic diagram of the hardware structure of the electronic device provided by the embodiments of the present disclosure. Detailed implementation manners

[0017] In order to make the purpose, technical solutions and advantages of the present disclosure clearer, the present disclosure will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present disclosure, and are not used to limit the present disclosure.

[0018] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present disclosure belongs. The terms used herein are only for the purpose of describing the embodiments of the present disclosure, and are not intended to limit the present disclosure.

[0019] First, several nouns involved in the present disclosure are analyzed: Artificial intelligence (AI): It is a new technical science that studies, develops theories, methods, technologies and application systems for simulating, extending and expanding human intelligence; artificial intelligence is a branch of computer science. Artificial intelligence attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. The research in this field includes robots, speech recognition, image recognition, natural language processing and expert systems, etc. Artificial intelligence can simulate the information process of human consciousness and thinking. Artificial intelligence also uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results of theories, methods, technologies and application systems.

[0020] Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technologies, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, robotics, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0021] A large language model (LLM) is a deep learning model trained using a large amount of text data that can generate natural language text or understand the meaning of language text. Large language models can detect large models, handle a variety of natural language tasks such as text classification, question answering, dialogue, etc., and are an important approach to artificial intelligence.

[0022] Object Storage Service (OBS) is an object-based storage service that provides users with the ability to store massive, secure, highly reliable, and low-cost data. There is no need to consider capacity limitations during use, and multiple storage types are provided for selection.

[0023] A directed acyclic graph (DAG) is a directed graph where the edges have a direction and there are no loops in the graph.

[0024] With the development of artificial intelligence technology, large models have achieved remarkable performance in fields such as natural language processing and computer vision. Large models require a huge number of parameters and complex computational requirements during the deployment process. How to quickly deploy large models is crucial for the popularization and application of large models.

[0025] In related technologies, due to the limited performance of terminal devices, different degrees of adaptation are required during the process of deploying large models to terminal devices. However, due to the large performance differences among different terminal devices, it often relies on manual operations for large model deployment according to the characteristics of each terminal device, greatly reducing the efficiency and quality of large model deployment.

[0026] Based on this, the embodiments of the present disclosure provide a large model automatic deployment method, device, equipment, and medium, which can improve the efficiency and quality of large model deployment.

[0027] The large model automatic deployment method in the embodiments of the present disclosure can be illustrated by the following embodiments.

[0028] The embodiments of the present disclosure can acquire and process relevant data based on artificial intelligence technology.

[0029] Please refer to Figure 1 , Figure 1 which is a schematic diagram of an implementation environment provided by the embodiments of the present disclosure. The implementation environment includes a terminal and a server. Among them, the terminal and the server are connected through a communication network.

[0030] Exemplarily, the server may obtain the downstream task information sent by the terminal, and determine the original large model, the lightweight large model, the target device type of the target device, and the target model architecture that can run on the target device based on the downstream task information; when at least one of the following conditions is met: the distillation device type during the distillation of the original large model to the lightweight large model is inconsistent with the target device type, and the original model architecture of the lightweight large model is inconsistent with the target model architecture, generate scheduling information for the pre-registered model conversion service; construct scheduling information for the pre-registered model distillation service for distilling the original large model to the lightweight large model, and construct scheduling information for the pre-registered model deployment service for the target device, and construct a directed acyclic graph task information based on the order of all the scheduling information; when the directed acyclic graph task information runs, first schedule the model distillation service to distill the original large model to the lightweight large model based on each scheduling information, then schedule the model conversion service to convert the distilled lightweight large model to the target device type or the target model architecture, and schedule the model deployment service to deploy the converted lightweight large model to the target device.

[0031] The terminal can also be used as the target device and finally receive the converted lightweight large model sent by the server, and deploy and apply this large model.

[0032] In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, etc.; the server can be configured as an independent physical server, or can be configured as a server cluster or a distributed system composed of multiple physical servers, and can also be configured as a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the large model automatic deployment method can also be deployed in the software in the server, and the software can be an application that implements the large model automatic deployment method, etc., but is not limited to the above forms.

[0033] The present disclosure can be used in numerous general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. The present disclosure can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present disclosure can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0034] It should be noted that in each specific embodiment of the present disclosure, when obtaining downstream task information input by the user terminal for relevant processing, the user's permission or consent will be obtained first. Moreover, the collection, use, and processing of these data, etc., will comply with relevant laws, regulations, and standards. In addition, when the embodiments of the present disclosure need to obtain downstream task information, the user's separate permission or separate consent can be obtained through pop-up windows or by jumping to a confirmation page, etc. After clearly obtaining the user's separate permission or separate consent, the necessary downstream task information for the normal operation of the embodiments of the present disclosure can be obtained.

[0035] Please refer to Figure 2 , Figure 2 which is a schematic flowchart of the large model automatic deployment method provided by the embodiments of the present disclosure. Figure 2 The method in

[0036] Step S101: Obtain the downstream task information input by the user terminal, and determine the original large model, lightweight large model, target device type of the target device, and the target model architecture that can run on the target device based on the downstream task information; Step S102: When at least one of the following conditions is met: the distillation device type during the distillation of the original large model to the lightweight large model is inconsistent with the target device type, and the original model architecture of the lightweight large model is inconsistent with the target model architecture, generate scheduling information for the pre-registered model conversion service; Step S103: Construct scheduling information for the pre-registered model distillation service for distilling the original large model to the lightweight large model, and construct scheduling information for the pre-registered model deployment service for the target device, and construct directed acyclic graph task information based on the order of all the scheduling information; Step S104, when the directed acyclic graph task information is running, first schedule the model distillation service to distill the original large model onto the lightweight large model based on each scheduling information, then schedule the model conversion service to convert the distilled lightweight large model to the target device type or target model architecture, and schedule the model deployment service to deploy the converted lightweight large model to the target device.

[0037] For the above-mentioned step S101, the client can be a terminal or an input module on the server, and the embodiments of the present disclosure do not make specific limitations thereto. The client can be provided with a mobile application, a web interface or a command-line tool, through which the user can input specific downstream task information to represent some necessary information required for the downstream task. Exemplarily, the downstream task information can be some specific information input by the user for selecting during the large model deployment process, such as large model id, small model id, dataset id, version id, distillation component configuration, distillation device chip type, target device chip type, target device model architecture, target device address, inference algorithm, etc.; or, the downstream task information is a task description used to describe what the target device needs to execute.

[0038] In the embodiments of the present disclosure, the original large model, the lightweight large model, the target device type of the target device, and the target model architecture that can run on the target device can be determined based on the downstream task information. For example, if the downstream task information includes some specific information during the large model deployment process, the required large model can be directly selected as the original large model based on the input of the client, and the required large model can be selected as a small model, that is, the lightweight large model, and the target device type of the target device and the target model architecture that can run on the target device can be directly determined; or, if the downstream device information includes a task description used to describe what the target device needs to execute, operations such as parsing and analyzing the downstream device information can be performed to obtain the required original large model, lightweight large model, target device type of the target device, and target model architecture that can run on the target device.

[0039] It should be noted that the original large model is a pre-trained large model with relatively powerful and complete functions, which can execute complex tasks. However, due to its large number of parameters, it is not suitable for deployment in some target devices with poor performance. The lightweight large model is a large model with a smaller number of parameters compared to the original large model, and can also be called a small model, which is suitable for deployment in some target devices with poor performance. Exemplarily, the lightweight large model can also be formed after compressing the original large model. However, due to its relatively poor performance, it needs to be adjusted before being deployed to the target device, including but not limited to operations such as model distillation and model conversion, and the embodiments of the present disclosure do not make specific limitations thereto.

[0040] Regarding the above step S102, embodiments of the present disclosure determine whether additional conversion of the lightweight large model is required to adapt to the characteristics and requirements of the target device, including determining whether at least one of the following is satisfied: the type of the distillation device during the distillation of the original large model into the lightweight large model is inconsistent with the type of the target device, and the original model architecture of the lightweight large model is inconsistent with the target model architecture.

[0041] Among them, the distillation device type is the type of device used when distilling the original large model into the lightweight large model. Different distillation devices have different computing capabilities, memory sizes, storage speeds, etc. Further, the distillation device type can be the distillation device chip type; the target device type is the type of device of the target device where the lightweight large model will ultimately be deployed, corresponding to the distillation device type. Further, the target device type can be the target device chip type. By comparing these two device types, it can be determined whether there are significant differences that will affect the performance of the lightweight large model on the target device. If there are significant differences between the distillation device type and the target device type in terms of computing power, memory limitations, or hardware acceleration support, etc., then further conversion of the lightweight large model is required.

[0042] Among them, the original model architecture is the architecture adopted by the lightweight large model during the distillation process, including the hierarchical structure of the model, the number and type of parameters, etc. The target model architecture is the model architecture that can run on the target device, which is restricted by the hardware and software environment of the target device. By comparing these two model architectures, it is determined whether there is an incompatible situation. For example, if the target device only supports specific types of neural network layers, while the lightweight large model contains incompatible layer types, then the lightweight large model needs to be converted to match the architecture requirements of the target device.

[0043] Model conversion services are predefined and registered. They can receive scheduling information and perform conversions on the lightweight large model based on this information. If any of the above checks finds an inconsistency, then embodiments of the present disclosure need to generate scheduling information for the pre-registered model conversion services, and this scheduling information will guide how the model conversion services perform the necessary conversions or optimizations on the lightweight large model.

[0044] Regarding the above step S103, embodiments of the present disclosure can construct scheduling information for the model distillation service, which is used to execute the distillation process; and can also construct scheduling information for the model conversion service to perform necessary conversions when the distilled lightweight large model is incompatible with the target device or the target model architecture.

[0045] Next, to ensure that each service is executed in the correct order and avoid circular dependencies, a directed acyclic graph task information is constructed. This task information is used to represent the directed acyclic graph required during the service execution process. The directed acyclic graph includes multiple nodes, and each node represents a service, such as a model distillation service, a model conversion service, a model deployment service, etc. The directed edges between the nodes represent the execution order and dependency relationships between the services. Finally, a unique identifier is assigned to each node in the directed acyclic graph.

[0046] Therefore, by constructing these scheduling information and directed acyclic graph task information, it can be ensured that each step in the automatic deployment process of the large model is executed according to the predetermined order and conditions, thereby improving the efficiency and reliability of the deployment, while reducing the need for manual intervention.

[0047] Regarding the above step S104, after obtaining the directed acyclic graph task information, the embodiments of the present disclosure can sequentially schedule each service in the directed acyclic graph based on the directed acyclic graph task information, including sequentially scheduling the model distillation service, the model conversion service, the model deployment service, etc.

[0048] Specifically, during the process of scheduling and executing the model distillation service, the original large model and the parameters required for distillation, such as the distillation algorithm and the dataset, can be passed as inputs to the model distillation service. The model distillation service will execute the distillation algorithm to distill the original large model into a lightweight large model, and the service will output the distilled lightweight large model, which will be used as the input for the subsequent services.

[0049] Subsequently, during the process of scheduling and executing the model conversion service, the distilled lightweight large model and the parameters required for conversion, such as the quantization bit number, the pruning ratio, the target device type, or the target model architecture, etc., can be used as inputs. The model conversion service will perform necessary conversions on the distilled lightweight large model according to the inputs to adapt to the type and architecture of the target device, and finally output the converted lightweight large model.

[0050] Finally, during the process of invoking and executing the model deployment service, the converted lightweight large model is deployed to the target device, including uploading the model file to the device, configuring the running environment on the device, and performing necessary tests or inferences to ensure that the model can run correctly. In this way, the automatic deployment of the large model is realized, thereby greatly improving the efficiency and quality of the deployment. Since the entire process is automated, the user only needs to input the downstream task information through the client during the deployment process. Therefore, the need for manual intervention can be significantly reduced, and the error rate during the deployment process can be lowered.

[0051] In summary, in the embodiments of the present disclosure, through steps S101 to S104, by executing the large model automatic deployment method, the original large model, the lightweight large model, the target device type of the target device, and the target model architecture that can run on the target device can be determined based on the input of the user side. In order to adapt to target devices with different performance differences, scheduling information for the pre-registered model conversion service is automatically generated. Since the model distillation service, the model conversion service, and the model deployment service are all pre-registered, no manual intervention is required when generating the directed acyclic graph task information. Moreover, the execution order among various services is constrained by the directed acyclic graph task information. After executing the directed acyclic graph task information, each service can be scheduled in sequence to automatically perform operations such as distilling and converting the large model, and finally deploy the converted lightweight large model to the target device, thereby improving the efficiency and quality of large model deployment.

[0052] In some embodiments, in the process of constructing the directed acyclic graph task information based on the order of all the scheduling information in step S103 above, it may further include: Configuring the scheduling dependency relationships corresponding to adjacent services among the scheduling information of the model distillation service, the model conversion service, and the model deployment service in sequence, and constructing the directed acyclic graph task information based on the order of all the scheduling information and the corresponding scheduling dependency relationships.

[0053] In the above steps, the scheduling dependency relationships corresponding to adjacent services are configured among the scheduling information of the model distillation service, the model conversion service, and the model deployment service, so that the execution of the subsequent service needs to wait until the execution of the dependent service is completed. Specifically, in the embodiments of the present disclosure, the model distillation service, the model conversion service, and the model deployment service are sequentially used as each node in the directed acyclic graph. Among them, the model conversion service has a scheduling dependency relationship with the previous model distillation service, so that after the model conversion service is scheduled, it needs to check whether the previous model distillation service has been executed, and only after the model distillation service is executed can the model conversion service run; in addition, the model deployment service has a scheduling dependency relationship with the previous model conversion service, so that after the model deployment service is scheduled, it needs to check whether the previous model conversion service has been executed, and only after the model conversion service is executed can the model deployment service run.

[0054] In addition, it should be noted that when there are many nodes in the directed acyclic graph, the service on any one node can also, through the scheduling dependency relationship, run only after any one of the previous services is executed, that is, the embodiments of the present disclosure are not limited to only depending on the execution of the previous service.

[0055] It should be noted that in the embodiments of the present disclosure, only model distillation service, model conversion service, and model deployment service are taken as examples for illustration. On the premise of meeting the requirements of the embodiments of the present disclosure, the directed acyclic graph task may further include other services as nodes. The embodiments of the present disclosure can configure scheduling dependencies based on the services on each node to construct the directed acyclic graph task information required by the embodiments of the present disclosure.

[0056] Please refer to Figure 3 , Figure 3 which Figure 2 is the schematic flowchart further included in step S104 in Step S201, when the directed acyclic graph task information runs, schedule the model distillation service, model conversion service, and model deployment service in sequence based on each scheduling information; Step S202, when running the model distillation service, distill the original large model onto the lightweight large model; Step S203, when running the model conversion service, determine whether the dependent model distillation service is completed, and after the model distillation service is completed, schedule the model conversion service to convert the distilled lightweight large model to the target device type or target model architecture; Step S204, when running the model deployment service, determine whether the dependent model conversion service is completed, and after the model conversion service is completed, schedule the model deployment service to deploy the converted lightweight large model to the target device.

[0057] In the above steps, when the directed acyclic graph task information starts to run, the embodiments of the present disclosure will schedule and execute each service in sequence according to the order of the nodes and edges defined in the indicated directed acyclic graph, including the model distillation service, model conversion service, and model deployment service.

[0058] Among them, in sequence, the model distillation service is scheduled and executed. The model distillation service will execute the distillation algorithm to distill the original large model onto the lightweight large model. Since the model conversion service depends on the execution of the model distillation service, therefore, before scheduling the model conversion service, the embodiments of the present disclosure will check whether the model distillation service it depends on has been completed. Only when the model distillation service is successfully completed, the model conversion service will be run, and the distilled lightweight large model will be converted to the target device type or target model architecture through the model conversion service to adapt to the target device. And the model deployment service depends on the execution of the model conversion service. Therefore, before scheduling the model deployment service, the embodiments of the present disclosure will check whether the model conversion service it depends on has been completed. Only when the model conversion service is successfully completed, the model deployment service will be run, and the converted lightweight large model will be deployed to the target device through the model deployment service.

[0059] In summary, through the execution of steps S201 to S204 in the embodiments of the present disclosure, the distillation, conversion, and deployment processes of the large model can be automatically completed. By configuring the scheduling dependencies between services, the smooth operation of each service can be ensured, and the subsequent services can be prevented from being started before the services that must be executed are completed, thereby greatly improving the efficiency and quality of the large model deployment. At the same time, since the entire process is automated, the need for manual intervention can be significantly reduced, and the error rate during the deployment process can be lowered.

[0060] Please refer to Figure 4 , Figure 4 which Figure 2 is another process schematic diagram further included in step S104 in . In some embodiments, step S104 may further include steps S301 to S303: Step S301, after scheduling the model distillation service, create a distillation container, mount the original large model and the lightweight large model onto the distillation container, and distill the original large model onto the lightweight large model within the distillation container; Step S302, after scheduling the model conversion service, set a conversion image in the distillation container, and convert the distilled lightweight large model to the target device type or target model architecture through the conversion image; Step S303, after scheduling the model deployment service, deploy the converted lightweight large model in the distillation container to the target device.

[0061] In the above steps, during the process of scheduling and running the model distillation service, the embodiments of the present disclosure will first create a distillation container, which is an independent and isolated environment for executing the model distillation task. Specifically, the embodiments of the present disclosure will use container technologies such as Docker to create and manage the distillation container. Container technologies allow applications and their dependencies to be packaged together and deployed and run as an independent unit. After creating the distillation container, the code or files of the original large model and the lightweight large model will be mounted into the distillation container, and the original large model will be distilled onto the lightweight large model within the distillation container.

[0062] Subsequently, when receiving the instruction to schedule the model conversion service, the embodiments of the present disclosure will set a conversion image in the distillation container. This conversion image is a predefined environment containing conversion tools or scripts for converting the distilled lightweight large model into a format or architecture suitable for running on the target device. Specifically, the embodiments of the present disclosure will use technologies such as Docker images to create and manage the conversion image. These images contain all the dependencies and tools required to execute the conversion task. Then, run the tools or scripts in the conversion image to convert the distilled lightweight large model into a format or architecture suitable for running on the target device, and finally obtain the converted lightweight large model.

[0063] Finally, after receiving the instruction to deploy the scheduling model service, the embodiment of the present disclosure will deploy the converted lightweight large model in the distillation container to the target device. Therefore, the embodiment of the present disclosure can automatically complete the distillation, conversion, and deployment processes of the large model, thereby greatly improving the efficiency and quality of large model deployment. At the same time, since the entire process is based on container and image technologies, the isolation and repeatability between different steps can be ensured, and the error rate during the deployment process can be reduced.

[0064] Please refer to Figure 5 , Figure 5 is Figure 2 the schematic flowchart further included in step S101 in Step S401, extracting the target device and the task description to be executed in the target device from the downstream task information; Step S402, obtaining the device description, target device type, and target model architecture that can run on the target device from a preset database; Step S403, extracting the first text feature of the task description and the second text feature of the device description, and fusing the first text feature and the second text feature to obtain a fusion feature; Step S404, predicting a preset plurality of candidate large models based on the fusion feature to obtain a prediction result, and selecting the original large model and the lightweight large model from the plurality of candidate large models based on the prediction result.

[0065] In the above steps, when the downstream task information is used to describe the task description to be executed by the target device, the embodiment of the present disclosure can extract and analyze the content in the downstream task information. Specifically, when the downstream task information describes what kind of task is expected to be performed in the target device. For example, the user can input input data such as "deploy a large model in device A to achieve an intelligent question and answer scenario in the technology field" through the input-output interface provided by the client. This input data can be used as the downstream task information. Among them, after parsing the downstream task information through natural language processing technology in the embodiment of the present disclosure, keywords or phrases about the target device and the text describing the task can be identified and extracted. It can be obtained that device A is the provided target device, and "deploy a large model to achieve an intelligent question and answer scenario in the technology field" is the task description to be executed in the target device.

[0066] Next, after determining the target device to be deployed in the embodiments of the present disclosure, the device description of the target device, the target device type, and the target model architecture that can run on the target device can be obtained from a preset database. The preset database is a business database provided in the embodiments of the present disclosure, which is used to pre-store various device information, and the target device is one of the devices. The database contains the device description of the target device, the target device type, and the target model architecture that can run on the target device. Among them, the device description can be the hardware specifications of the device, such as the CPU type, memory size, storage space, operating system information, supported software environment, etc. The target device type is the device type of the target device on which the lightweight large model is finally to be deployed, corresponding to the distillation device type, and can be the target device chip type; the target model architecture refers to the type of large model architecture that the device can support to run, such as the PyTorch model architecture or the TensorRt model architecture. In addition, the data can also contain other relevant information of the target device, such as device compatibility information, performance evaluation data, user feedback, etc.

[0067] Subsequently, the embodiments of the present disclosure can respectively extract text features from the task description and the device description, and obtain the first text feature of the task description and the second text feature of the device description. The first text feature and the second text feature can include lexical features, syntactic features, semantic features, etc. under the corresponding descriptions. For example, keywords can be extracted from the task description, such as "image", "recognition", "technology field", and these keywords can be converted into vector representations using word embedding techniques (such as Word2Vec). Also, descriptions of the target device are obtained from the preset database, such as "8-core CPU", "4GB RAM", "Android system", etc., and these descriptions are also converted into vector representations using word embedding or other text representation methods.

[0068] The embodiments of the present disclosure also fuse the extracted first text feature and second text feature together to form a fused feature vector, which contains both information about the task and information about the target device. Among them, the fusion method can be feature concatenation, weighted average, or using more complex deep learning models, such as attention mechanisms, neural network layers, etc. to fuse the features to generate a fused feature vector that combines task requirements and device characteristics.

[0069] It should be noted that fusing the first text feature and the second text feature can help understand the requirements of the task for the model performance and the computing power that the device can provide, so as to select a model that can meet the task requirements and does not exceed the device performance limit. Considering the task description or device description alone may lead to selection bias. By fusing the two features, factors of both the task and the device can be comprehensively considered. Moreover, as the AI technology continues to develop and new devices emerge, the content of the task description and device description may change. The method of fusing features can flexibly adapt to these changes. Therefore, in the embodiments of the present disclosure, by fusing the first text feature of the task description and the second text feature of the device description, the task requirements and the performance characteristics of the target device can be comprehensively considered, so as to more accurately select a suitable original large model and lightweight large model.

[0070] Finally, the embodiments of the present disclosure can predict a plurality of preset candidate large models based on the fused features to obtain a prediction result, and select the original large model and the lightweight large model from the plurality of candidate large models based on the prediction result. Specifically, the candidate large models are some large models pre-stored in the embodiments of the present disclosure, which can meet the deployment requirements of different scenarios. The embodiments of the present disclosure can use machine learning or deep learning algorithms to predict the most suitable large model according to the fused feature vector, including classifiers, regression models, etc., and the output of the algorithm is a score or ranking of one or more candidate large models. The most suitable original large model and lightweight large model are selected according to these scores or rankings. Finally, the most suitable original large model and lightweight large model for the target device and task are selected from the plurality of candidate large models, and these models will be used in the subsequent model distillation and deployment processes.

[0071] In summary, through the above steps, the embodiments of the present disclosure can automatically select the most suitable original large model and lightweight large model based on the downstream task information input by the user, and determine the type of the target device and the target model architecture that can be run, which provides important input information for the subsequent model distillation, conversion and deployment processes.

[0072] Please refer to Figure 6 , Figure 6 which Figure 5 is a schematic flowchart further included in step S404. In some embodiments, step S404 may include steps S501 to S502: Step S501, input the fused features into a preset first prediction model to predict the large models that meet the task description requirements among the plurality of candidate large models, obtain a first prediction result, and select the original large model from the plurality of candidate large models based on the first prediction result; Step S502: Input the fused features into a preset second prediction model to predict the large models among multiple candidate large models that meet the device description requirements, obtain a second prediction result, and select a lightweight large model from the multiple candidate large models based on the second prediction result.

[0073] In the above step, the first prediction model is a pre-trained classifier. Its role is to predict which candidate large models can meet specific task requirements based on the input fused features. This classifier is a deep learning model that has been trained to identify features related to the task description and classify candidate large models based on these features. When the fused features are input into the first prediction model, the model will output a first prediction result. The first prediction result is a probability distribution indicating the probability that each candidate large model meets the task description requirements, or the first prediction result is a direct classification result indicating which candidate large model is the most suitable. Based on the first prediction result, a candidate large model with the highest probability can be selected as the original large model for the subsequent model distillation process. In addition, if the first prediction result is a classification result, then directly select the candidate large model classified as the most suitable as the original large model.

[0074] Similarly, the second prediction model is a pre-trained classifier. Its role is to predict which candidate large models can meet specific device requirements based on the input fused features. This classifier is a deep learning model that has been trained to identify features related to the device description and classify candidate large models based on these features. When the fused features are input into the second prediction model, the model will output a second prediction result. The second prediction result is a probability distribution indicating the probability that each candidate large model meets the device description requirements, or the second prediction result is a direct classification result indicating which candidate large model is the most suitable. Based on the second prediction result, a candidate large model with the highest probability can be selected as the lightweight large model for the subsequent model distillation and deployment processes. In addition, if the second prediction result is a classification result, then directly select the candidate large model classified as the most suitable as the lightweight large model.

[0075] Please refer to Figure 7 , Figure 7 is Figure 2 Another process schematic diagram further included in step S104. In some embodiments, during the process of step S104 where the model distillation service schedules the distillation of the original large model to the lightweight large model, steps S601 to S602 may also be included: Step S601: Determine the distillation dataset based on the downstream task information; Step S602: Schedule the model distillation service and distill the original large model to the lightweight large model through the distillation dataset.

[0076] In the above steps, in the embodiments of the present disclosure, the distillation dataset required for distillation can also be determined based on the downstream task information. For example, when the downstream task information is some specific information input by the user side for selecting the deployment process of the large model, the embodiments of the present disclosure can determine the required dataset based on the input of the user side as the distillation dataset; or, when the downstream task information is a task description for describing the tasks that the target device needs to execute, the embodiments of the present disclosure can analyze based on the downstream task information to screen a dataset that meets the current task requirements from multiple alternative datasets as the distillation dataset. Finally, after determining the distillation dataset, the embodiments of the present disclosure can schedule the model distillation service and distill the original large model onto the lightweight large model through the distillation dataset.

[0077] Please refer to Figure 8 , Figure 8 which is another schematic flowchart of the large model automatic deployment method provided by the embodiments of the present disclosure. In some embodiments, the large model automatic deployment method may further include steps S701 to S702: Step S701, when the distillation device type in the process of distilling the original large model into the lightweight large model is the same as the target device type, and the original model architecture of the lightweight large model is the same as the target model architecture, construct the directed acyclic graph task information based on the order of the scheduling information of the model distillation service and the scheduling information of the model deployment service; Step S702, when the directed acyclic graph task information runs, first schedule the model distillation service to distill the original large model onto the lightweight large model based on each scheduling information, and schedule the model deployment service to deploy the distilled lightweight large model onto the target device.

[0078] In the above steps, the embodiments of the present disclosure determine whether additional conversion of the lightweight large model is required to adapt to the characteristics and requirements of the target device, including determining whether at least one of the following is satisfied: the distillation device type in the process of distilling the original large model into the lightweight large model is inconsistent with the target device type, and the original model architecture of the lightweight large model is inconsistent with the target model architecture. When the distillation device type in the process of distilling the original large model into the lightweight large model is the same as the target device type, and the original model architecture of the lightweight large model is the same as the target model architecture, it indicates that no model conversion service is required anymore. Therefore, the directed acyclic graph task information can be directly constructed based on the order of the scheduling information of services such as the model distillation service and the model deployment service.

[0079] Similarly, when the model conversion service is not required, the constructed directed acyclic graph task information is also used to represent the directed acyclic graph required during the service execution process. The directed acyclic graph includes multiple nodes, and each node represents a service, such as a model distillation service, a model deployment service, etc. The directed edges between the nodes represent the execution order and dependency relationships between the services. Finally, a unique identifier is assigned to each node in the directed acyclic graph.

[0080] Therefore, by constructing these scheduling information and directed acyclic graph task information, it can be ensured that each step of the large model automatic deployment process is executed in accordance with the predetermined order and conditions, thereby improving the efficiency and reliability of the deployment, and at the same time reducing the need for manual intervention.

[0081] Similarly, after obtaining the directed acyclic graph task information when the model conversion service is not required, the embodiments of the present disclosure can sequentially schedule each service in the directed acyclic graph based on the directed acyclic graph indicated by the directed acyclic graph task information, including sequentially scheduling the model distillation service, the model deployment service, etc., so as to first schedule the model distillation service to distill the original large model onto the lightweight large model based on each scheduling information, and schedule the model deployment service to deploy the distilled lightweight large model to the target device. The embodiments of the present disclosure will not elaborate on this.

[0082] Next, taking the downstream task information as an example of some specific information input by the user side for selecting during the large model deployment process, the above embodiments will be illustrated by examples: First, please refer to Figure 9 , Figure 9 which is a schematic diagram of the large model automatic deployment system provided by the embodiments of the present disclosure, showing the overall structure of the large model deployment. As Figure 9 shown, the overall steps of the large model deployment include: Step S801, the environment preparation stage; Among them, in the environment preparation stage, the present embodiment can prepare the system-built large model, built-in small model, distillation algorithm, running image, etc. related to the large model deployment process, and upload them to the system. Then, prepare the large model mapping table, small model mapping table, distillation algorithm mapping table, algorithm configuration mapping table and upload these mapping tables to the system. At the same time, the system provides pluggable management of information such as large models, algorithms, and computing platforms, as well as the associated mapping function between models, algorithms, and computing platforms. Next, the present embodiment can obtain the service names, service ports, service addresses, service protocols, health status, routing and other information of services such as the large model management service, distillation algorithm management service, large model distillation service, model conversion service, model deployment service, etc. in the system, and then register this service information in the registration center.

[0083] And as Figure 9As shown in the figure, business data tables, mapping tables, service information tables, and process task tables can be stored in the business database (MySQL); distillation algorithm files, large model files, small model files, scenario models, DAG configuration files, and log files can be stored in the object storage service (OBS).

[0084] Step S8021, store the deployment data file; among them, the Web management terminal can store the deployment data file into the data storage; Step S8022, submit the deployment information; among them, the Web management terminal can submit the deployment information to the system service, and the system service includes a task process service, a build configuration service, and a log monitoring service; Step S8023, store the deployment information; among them, the system service can store the deployment information into the data storage; Step S8024, obtain the mapping table; among them, the system service can obtain the mapping table from the data storage; Step S8025, obtain the service information; among them, the system service can obtain the service information from the registration center; Step S8026, build a directed acyclic graph configuration file; among them, the system service can build a DAG configuration file according to the obtained information; Among them, the specific processes of the above steps S8021 to S8026 can be referred to Figure 10 , Figure 10 is a schematic diagram of building acyclic graph task information provided by an embodiment of the present disclosure. Figure 10 The process in Step S901, upload the distillation data set; this step corresponds to the above step S8021; Step S902, deployment information; this step corresponds to the above step S8022; Step S903, store the deployment information; this step corresponds to the above step S8023; Step S904, obtain the mapping table; this step corresponds to the above step S8024; Step S905, obtain the service information; this step corresponds to the above step S8025; Step S906, build the configuration; this step corresponds to the above step S8026; Step S907, upload the configuration file; According to Figure 10 , in the process of building directed acyclic graph task information, that is, building a DAG configuration file, the user submits deployment information such as downstream task information, data, and computing platforms through the Web management terminal in steps S901 and S902, and the corresponding deployment data files are respectively uploaded to the object storage service (OBS) and the business database (MySQL).

[0085] Next, in step S903, the build configuration service determines based on the distillation device chip type and the target device chip type in the deployment information whether the deployment process requires model conversion, and obtains the PyTorch model architecture obtained by distillation and the TensorRt model architecture submitted by the user. When they are inconsistent, model conversion is required. Most different types of conversion images are built into the system. In step S904, the build configuration service then retrieves the mapping table and deployment information from the business database. In step S905, it obtains the service information from the registration center, that is, the registration information of the system services. If model conversion is required, according to the dependency order of the large model deployment process, the large model management service, the distillation algorithm management service, the model distillation service, the model conversion service, and the model deployment service are used as the execution order of the DAG configuration file. Otherwise, according to the dependency order of the large model deployment process, the large model management service, the distillation algorithm management service, the model distillation service, and the model deployment service are used as the execution order of the DAG configuration file, and are written into the DAG configuration file DAG.json. When executing step S906, the build configuration service writes the default parameters default_args of the system default automation engine into the directed acyclic graph configuration file, and uploads the generated DAG.json to the object storage service for storage through step S907.

[0086] Step S803, scheduling service; Step S8031, service output file, service status, service log; among them, the service output file, service status, and service log finally generated by the system can be stored with data; Among them, the process of the scheduling service can refer to Figure 11 , Figure 11 which is a schematic diagram of the large model automatic deployment process provided by the embodiments of the present disclosure. Figure 11 The process in includes: Step S1001, loading; Step S1002, creating a task workflow; Step S1003, scheduling; Step S1004, creating a task; Step S1005, updating task information and the scheduling service; Step S1006, storing to the object storage service; Step S1007, task retry; According to Figure 11, specifically, in this embodiment, the large model is automatically deployed through the automation engine. First, in step S1001, the automation engine loads the DAG configuration file to obtain the default parameters for engine initialization in the directed acyclic graph configuration file, that is, the parameters in default_args, and initializes the engine. In step S1002, the scheduler creates a task workflow according to the order of the large model deployment process in the directed acyclic graph configuration and stores it in the business database. Subsequently, the scheduler queries the task workflow from the business database. In step S1003, according to the configuration, the task nodes in the task workflow are scheduled to the task queue at regular intervals and the status flag of the task nodes is updated to waiting. In step S1004, the executor obtains the task node information in the task queue, creates a task execution unit to run the tasks, such as the distillation task A, the conversion task B, and the deployment task C in the figure, and updates the node status flag to running. In step S1005, the execution unit schedules the corresponding services according to the service information in the task information, such as the large model management service, the distillation algorithm management service, the large model distillation service, the model conversion service, and the model deployment service in the figure, and updates the node status information to success or failure according to the service execution status, and updates the service execution information and the node status to the business database. Finally, in step S1006, the execution unit stores the scheduling service operation result in the object storage service. In addition, when the task node execution fails, step S1007 is executed, and the automation engine retries the task node according to the initialization configuration and reloads it into the queue. In addition, by executing step S1008, the Web management terminal can obtain information such as the log information and status in the task workflow through the business database.

[0087] For example, if the information submitted by the user on the client page includes the large model id, small model id, dataset id, version id, distillation component configuration, distillation device chip type, target device chip type, target device model architecture, target device address, inference algorithm, etc., the user can select the yolov8_x large model as the original large model, and take the yolov8_l small model as the lightweight large model as an example, select the road disease dataset as the distillation dataset, the distillation device type is the distillation device chip type, and it is PowerEdgeT640 (RTX2080Ti), and the target device type is the target device chip type, which is T506s (Jetson Xavier NX).

[0088] When implementing the large model automatic deployment method in the embodiments of the present disclosure, first, the large model management service manages the models uploaded by users and the models built into the system. The purpose is to obtain the large and small model files through the model ID. The large model distillation service obtains the addresses of the large and small model files, as well as the distillation configuration file, uploads the large and small models to the parallel file system, and starts the distillation container according to the distillation running image to mount the model address and the dataset into the container directory, and then performs distillation. A PyTorch model architecture is obtained through distillation, and then it is determined whether the PyTorch model architecture, the TensorRt model architecture submitted by the user, the distillation device chip type, and the target device chip type are consistent. If they are inconsistent, model conversion is performed. The model conversion service first obtains the distillation device chip type (here, PowerEdge T640 (RTX2080Ti) is used as an example) and the original model architecture, and the target device chip type (here, PowerEdge T640 (RTX2080Ti) is used as an example) and the target model architecture submitted by the user. Through the built-in conversion image, the address of the distilled model is obtained for mounting into the container directory, and a TensorRt architecture model is obtained through conversion. Finally, the model deployment service pushes the converted TensorRt model and the inference algorithm to the target device for deployment.

[0089] During the execution of the DAG configuration file, automation causes the DAG configuration file to be loaded, and the service information set of tasks is obtained. Then, the set is traversed to obtain the task_id of each service information. If the task_id is not empty, it indicates that the current service has been created, and the execution of the next one is skipped. If it is empty, it is checked whether there are dependent services. If it is empty, the scheduler puts the current service information into the execution queue and randomly generates a UUID as the task_id. If it is not empty, the service status of the task_id of the dependent service is queried. If the dependent service has not ended or has failed, the current task is skipped and the next task is executed. When all the dependent services have been executed, the scheduler schedules the service information to the execution queue, thus ensuring the orderly execution of the services. The executor will create services according to the execution queue.

[0090] Step S804, log monitoring; finally, the Web management terminal can obtain relevant logs through log monitoring.

[0091] In the above multiple processes, except that the Web management terminal visualization page needs to be manually selected and submit the downstream task information for deployment, no manual intervention is required. Therefore, in this embodiment, the scheduling service and the functional service are completed by designing an automation engine, and the system realizes efficient and diversified one-stop large model deployment.

[0092] Please refer to Figure 12, an embodiment of the present disclosure further provides a large model automatic deployment device, which can implement the above large model automatic deployment method. The large model automatic deployment device includes: An information receiving module 1201, configured to obtain downstream task information input by a user terminal, and determine an original large model, a lightweight large model, a target device type of a target device, and a target model architecture that can run on the target device based on the downstream task information; A scheduling confirmation module 1202, configured to generate scheduling information for a pre-registered model conversion service when at least one of the following conditions is met: the distillation device type during the distillation of the original large model to the lightweight large model is inconsistent with the target device type, and the original model architecture of the lightweight large model is inconsistent with the target model architecture; A directed acyclic graph construction module 1203, configured to construct scheduling information for a pre-registered model distillation service for distilling the original large model to the lightweight large model, and construct scheduling information for a pre-registered model deployment service for the target device, and construct directed acyclic graph task information based on the order of all the scheduling information; An automatic deployment module 1204, configured to, when the directed acyclic graph task information runs, first schedule the model distillation service to distill the original large model to the lightweight large model based on each scheduling information, then schedule the model conversion service to convert the distilled lightweight large model to the target device type or the target model architecture, and schedule the model deployment service to deploy the converted lightweight large model to the target device.

[0093] In summary, by executing the large model automatic deployment method, the large model automatic deployment device can determine the original large model, the lightweight large model, the target device type of the target device, and the target model architecture that can run on the target device based on the input of the user terminal, and automatically generate scheduling information for a pre-registered model conversion service in order to adapt to target devices with different performance differences. Since the model distillation service, the model conversion service, and the model deployment service are all pre-registered, no manual intervention is required when generating the directed acyclic graph task information, and the execution order between the services is constrained by the directed acyclic graph task information. After executing the directed acyclic graph task information, each service can be scheduled in sequence to automatically perform operations such as distilling and converting the large model, and finally deploy the converted lightweight large model to the target device, thereby improving the efficiency and quality of large model deployment.

[0094] The specific implementation manner of the large model automatic deployment device is basically the same as the specific embodiment of the above large model automatic deployment method, and will not be elaborated here. On the premise of meeting the requirements of the embodiments of the present disclosure, other functional modules can be set in the large model automatic deployment device to implement the large model automatic deployment method in the above embodiments.

[0095] An embodiment of the present disclosure also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the above-mentioned large model automatic deployment method is implemented. The electronic device can be any intelligent terminal including a tablet computer, an in-vehicle computer, etc.

[0096] Please refer to Figure 13 , Figure 13 which schematically shows the hardware structure of an electronic device in another embodiment. The electronic device includes: A processor 1301, which can be implemented in ways such as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided by the embodiments of the present disclosure; A memory 1302, which can be implemented in forms such as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 1302 can store an operating device and other application programs. When implementing the technical solutions provided by the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1302, and the processor 1301 is called to execute the large model automatic deployment method of the embodiments of the present disclosure; An input / output interface 1303, which is used to implement information input and output; A communication interface 1304, which is used to implement communication interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or through wireless means (such as mobile network, WIFI, Bluetooth, etc.); A bus 1305, which transmits information between various components of the device (such as the processor 1301, the memory 1302, the input / output interface 1303, and the communication interface 1304); Among them, the processor 1301, the memory 1302, the input / output interface 1303, and the communication interface 1304 achieve communication connections with each other inside the device through the bus 1305.

[0097] An embodiment of the present disclosure also provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, the above-mentioned large model automatic deployment method is implemented.

[0098] The memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include a memory remotely disposed relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0099] The embodiments described in the embodiments of the present disclosure are for more clearly illustrating the technical solutions of the embodiments of the present disclosure, and do not constitute a limitation on the technical solutions provided by the embodiments of the present disclosure. Those skilled in the art will know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present disclosure are equally applicable to similar technical problems.

[0100] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present disclosure, and may include more or fewer steps than those shown in the figures, or combine certain steps, or different steps.

[0101] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0102] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the devices, can be implemented as software, firmware, hardware, and appropriate combinations thereof.

[0103] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present disclosure and the above drawings are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present disclosure described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, device, product, or equipment that includes a series of steps or units does not have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or equipment.

[0104] It should be understood that in this disclosure, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" may represent: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after. "At least one (one) of the following" or similar expressions refer to any combination of these items, including any combination of single items (ones) or plural items (ones). For example, at least one (one) of a, b, or c may represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0105] In several embodiments provided in this disclosure, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the above-mentioned division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of devices or units can be in electrical, mechanical or other forms.

[0106] The units described above as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0107] In addition, each functional unit in various embodiments of this disclosure can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0108] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present disclosure, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of the various embodiments of the present disclosure. The aforementioned storage medium includes: various media that can store programs such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.

[0109] The preferred embodiments of the embodiments of the present disclosure have been described above with reference to the accompanying drawings, and thus do not limit the scope of the rights of the embodiments of the present disclosure. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of the present disclosure shall be within the scope of the rights of the embodiments of the present disclosure.

Claims

1. An automatic deployment method for large models, characterized in that, Including: Obtain downstream task information input by the user terminal, and determine the original large model, lightweight large model, target device type of the target device, and the target model architecture that can run on the target device based on the downstream task information; When at least one of the following conditions is met: the distillation device type during the distillation of the original large model to the lightweight large model is inconsistent with the target device type, and the original model architecture of the lightweight large model is inconsistent with the target model architecture, generate scheduling information for a pre-registered model conversion service; Construct scheduling information for the model distillation service of distilling the original large model to the lightweight large model pre-registered, and construct scheduling information for the model deployment service of the target device pre-registered, and construct directed acyclic graph task information based on the order of all the scheduling information; When the directed acyclic graph task information runs, first schedule the model distillation service to distill the original large model to the lightweight large model based on each scheduling information, then schedule the model conversion service to convert the distilled lightweight large model to the target device type or the target model architecture, and schedule the model deployment service to deploy the converted lightweight large model to the target device.

2. The large model automatic deployment method according to claim 1, wherein The constructing directed acyclic graph task information based on the order of all the scheduling information includes: Configure the scheduling dependency relationship corresponding to adjacent services for each scheduling information of the model distillation service, the model conversion service, and the model deployment service in sequence, and construct directed acyclic graph task information based on the order of all the scheduling information and the corresponding scheduling dependency relationship; The when the directed acyclic graph task information runs, first schedule the model distillation service to distill the original large model to the lightweight large model based on each scheduling information, then schedule the model conversion service to convert the distilled lightweight large model to the target device type or the target model architecture, and schedule the model deployment service to deploy the converted lightweight large model to the target device includes: When the directed acyclic graph task information runs, schedule the model distillation service, the model conversion service, and the model deployment service in sequence based on each scheduling information; When running the model distillation service, distill the original large model to the lightweight large model; When running the model conversion service, determine whether the dependent model distillation service is completed, and after the model distillation service is completed, schedule the model conversion service to convert the distilled lightweight large model to the target device type or the target model architecture; When running the model deployment service, determine whether the dependent model conversion service is completed, and after the model conversion service is completed, schedule the model deployment service to deploy the converted lightweight large model to the target device.

3. The large model automatic deployment method according to claim 1, characterized in that The scheduling of the model distillation service distills the original large model onto the lightweight large model, then schedules the model conversion service to convert the distilled lightweight large model to the target device type or the target model architecture, and schedules the model deployment service to deploy the converted lightweight large model to the target device, including: After scheduling the model distillation service, create a distillation container, mount the original large model and the lightweight large model onto the distillation container, and distill the original large model onto the lightweight large model within the distillation container; After scheduling the model conversion service, install a conversion image in the distillation container, and convert the distilled lightweight large model to the target device type or the target model architecture through the conversion image; After scheduling the model deployment service, deploy the converted lightweight large model in the distillation container to the target device.

4. The large model automatic deployment method according to claim 1, wherein The determination of the original large model, the lightweight large model, the target device type of the target device, and the target model architecture that can run on the target device based on the downstream task information includes: Extract the target device and the task description that needs to be executed in the target device from the downstream task information; Obtain the device description, the target device type, and the target model architecture that can run on the target device from a preset database; Extract the first text feature of the task description and the second text feature of the device description, and fuse the first text feature and the second text feature to obtain a fused feature; Perform predictions on a preset number of candidate large models based on the fused feature to obtain a prediction result, and select the original large model and the lightweight large model from the multiple candidate large models based on the prediction result.

5. The large model automatic deployment method according to claim 4, wherein The performing of predictions on a preset number of candidate large models based on the fused feature to obtain a prediction result, and the selection of the original large model and the lightweight large model from the multiple candidate large models based on the prediction result includes: Input the fused feature into a preset first prediction model to predict the large models that meet the requirements of the task description among the multiple candidate large models, obtain a first prediction result, and select the original large model from the multiple candidate large models based on the first prediction result; Input the fused feature into a preset second prediction model to predict the large models that meet the requirements of the device description among the multiple candidate large models, obtain a second prediction result, and select the lightweight large model from the multiple candidate large models based on the second prediction result.

6. The large model automatic deployment method according to claim 1, wherein The scheduling of the model distillation service to distill the original large model onto the lightweight large model includes: Determine the distillation dataset based on the downstream task information; Schedule the model distillation service, and distill the original large model onto the lightweight large model through the distillation dataset.

7. The large model automatic deployment method according to claim 1, characterized in that, The large model automatic deployment method further includes: When the distillation device type in the process of distilling the original large model into the lightweight large model is consistent with the target device type, and the original model architecture of the lightweight large model is consistent with the target model architecture, construct the directed acyclic graph task information based on the order of the scheduling information of the model distillation service and the scheduling information of the model deployment service; When the directed acyclic graph task information runs, first schedule the model distillation service to distill the original large model onto the lightweight large model based on each of the scheduling information, and schedule the model deployment service to deploy the distilled lightweight large model onto the target device.

8. An automatic deployment device for large models, characterized in that, Including: An information receiving module, configured to obtain downstream task information input by a user terminal, and determine an original large model, a lightweight large model, a target device type of the target device, and a target model architecture that can run on the target device based on the downstream task information; A scheduling confirmation module, configured to generate scheduling information for a pre-registered model conversion service when at least one of the following conditions is met: the distillation device type in the process of distilling the original large model into the lightweight large model is inconsistent with the target device type, and the original model architecture of the lightweight large model is inconsistent with the target model architecture; A directed acyclic graph construction module, configured to construct scheduling information for the model distillation service for distilling the original large model into the lightweight large model that is pre-registered, and construct scheduling information for the model deployment service for the pre-registered target device, and construct directed acyclic graph task information based on the order of all the scheduling information; An automatic deployment module, configured to, when the directed acyclic graph task information runs, first schedule the model distillation service to distill the original large model onto the lightweight large model based on each of the scheduling information, then schedule the model conversion service to convert the distilled lightweight large model to the target device type or the target model architecture, and schedule the model deployment service to deploy the converted lightweight large model onto the target device.

9. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the large model automatic deployment method according to any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the large model automatic deployment method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Text classification method based on graph path knowledge extraction

    CN113515632A

  • Model acquisition method and device, model deployment method and device, equipment and medium

    CN116204321A

  • Cloud edge-end collaborative large-model lightweight deployment platform system and method

    CN117608591A

  • Method and device for improving data through efficiency based on large model and medium

    CN119622206A

  • Active adaptation of networked compute devices using vetted reusable software components

    US20190171438A1