End-to-End Distillation Deployment Method, Device, Equipment and Medium for Large Models Aimed at Low-Computing-Power Devices

By deploying large models end-to-end distillation on low-computing devices, the information security risks and service response timeout problems when users access large language models in the cloud are solved, and privacy protection and access efficiency are improved.

CN120066803BActive Publication Date: 2025-07-29SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510542975.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-07-29
Estimated Expiration
2045-04-28

AI Technical Summary

Technical Problem

When accessing the cloud-based large language model, users face information security risks and computing cluster service response timeout problems, which affects the user experience.

Method used

Large-scale model end-to-end distillation deployment is carried out on low-computing equipment. By determining the target data set and student models in the computing cluster in the target cloud, distillation is performed using the teacher model, and the distillation model is deployed to the target equipment in combination with preset model quantization algorithms and inference frameworks.

Benefits of technology

Improve user privacy protection, improve the efficiency of access to large language models, and enhance user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120066803B_ABST
    Figure CN120066803B_ABST
Patent Text Reader

Abstract

The present application discloses an end-to-end distillation deployment method, device, equipment and medium for large models for low-computing-power devices, which relates to the field of artificial intelligence and includes: determining a first target data set and a target student model in a computing cluster, and deploying a distillation model training framework; deploying a preset large language model to the computing cluster by using a first preset large model inference framework, and determining the deployed preset large language model as a teacher model; if the distillation model training framework is a black-box knowledge distillation framework, determining a second target data set based on the teacher model and the first target data set, and distilling the target student model by using the second target data set to obtain a distillation model; if the distillation model training framework is a white-box knowledge distillation framework, distilling the target student model based on the teacher model to obtain a distillation model; deploying the distillation model to a target device based on a second preset large model inference framework. Therefore, the efficiency of accessing the large language model can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence, and particularly to an end-to-end distillation deployment method, device, equipment and medium for large models for low-computing-power devices. Background Art

[0002] In today's digital age, the outstanding capabilities of large language models in the field of human-computer dialogue have been continuously highlighted, and the demand for processing various text consultation tasks has shown a rapid growth trend. Currently, users generally use mobile devices to access cloud services to utilize large language models to perform daily tasks.

[0003] However, this operation mode has many drawbacks. From the perspective of information security, when users use mobile devices to access cloud large language model services, information security faces severe challenges. Due to the complexity of the network environment and security protection vulnerabilities, user information is extremely easy to be leaked during the access process, which may lead to the acquisition of important information such as personal privacy and business secrets by lawbreakers, causing serious losses to users. In terms of service response, when a large number of service requests flood into the computing cluster deploying the large language model at the same time, the computing cluster will be overwhelmed and frequently experience service response timeouts, seriously affecting the user experience.

[0004] Therefore, how to strengthen user privacy protection during the process of users accessing large language models and improve the efficiency of accessing large language models is a technical problem that needs to be solved urgently at present. Summary of the Invention

[0005] In view of this, the purpose of the present invention is to provide an end-to-end distillation deployment method, device, equipment and medium for large models for low-computing-power devices, which can strengthen user privacy protection during the process of users accessing large language models and improve the efficiency of accessing large language models. The specific solutions are as follows:

[0006] In a first aspect, the present application provides an end-to-end distillation deployment method for large models for low-computing-power devices, including:

[0007] In the computing cluster of the target cloud, based on the obtained problem domain, determine the first target data set from the preset database, and determine the target student model from the preset student model library by using the configuration information of the obtained target device, and deploy the distillation model training framework; the target device is a device whose own computing power meets the preset low-computing-power determination condition;

[0008] Use the first preset large model inference framework and the graphics card of the computing cluster to deploy the preset large language model to the computing cluster, and determine the deployed preset large language model as the teacher model;

[0009] If the distillation model training framework is a black-box knowledge distillation framework, determine a second target dataset based on the teacher model and the first target dataset, and use the second target dataset to distill the target student model to obtain a distillation model;

[0010] If the distillation model training framework is a white-box knowledge distillation framework, distill the target student model based on the output probability distribution of the teacher model to obtain the distillation model;

[0011] Deploy the distillation model to the target device based on a preset model quantization algorithm and a second preset large model inference framework.

[0012] Optionally, the deploying the preset large language model to the computing cluster by using the first preset large model inference framework and the graphics cards of the computing cluster includes:

[0013] Determine the vLLM framework as the first preset large model inference framework, and perform a first model deployment operation on the preset large language model by using the paged attention mechanism and the preset scheduler of the first preset large model inference framework;

[0014] Perform a second model deployment operation on the preset large language model by using the flash attention mechanism of the first preset large model inference framework;

[0015] Deploy the preset large prediction model to the computing cluster based on the first model deployment operation and the second model deployment operation.

[0016] Optionally, the deploying the preset large language model to the computing cluster by using the first preset large model inference framework and the graphics cards of the computing cluster further includes:

[0017] When receiving a model conversation request corresponding to the preset large language model sent by the target client, determine a corresponding model access path, the IP address and port number of the server where the preset large language model is located, the tensor parallel size of the preset large language model, and the graphics card utilization rate of the server to determine model conversation parameters;

[0018] Determine model generation information based on the model conversation parameters and the preset large language model, and output the model generation information to the target client.

[0019] Optionally, the determining the second target dataset based on the teacher model and the first target dataset, and using the second target dataset to distill the target student model to obtain a distillation model includes:

[0020] The teacher model generates data classification probabilities, confidence levels, and intermediate layer output results of the teacher model based on the first target data set, and determines the data classification probabilities, the confidence levels, and the intermediate layer output results as the second target data set;

[0021] Determine a soft label distillation loss function and a hard label distillation loss function, and determine a target distillation loss function based on a preset dynamic loss weight algorithm, the soft label distillation loss function, and the hard label distillation loss function;

[0022] Use the second target data set and the target distillation loss function to distill the target student model and obtain a distilled model.

[0023] Optionally, distilling the target student model based on the output probability distribution of the teacher model to obtain the distilled model includes:

[0024] Determine attention weights, hidden states of the intermediate layer, and prediction results of the output layer based on the teacher model, and determine the attention weights, the hidden states, and the prediction results as the output probability distribution of the teacher model;

[0025] Determine the output probability distribution of the target student model;

[0026] Distill the target student model based on the KL divergence method, a preset temperature scaling technique, the output probability distribution of the teacher model, and the output probability distribution of the target student model to obtain the distilled model.

[0027] Optionally, deploying the distilled model to the target device based on a preset model quantization algorithm and a second preset large model inference framework includes:

[0028] Take The algorithm as the preset model quantization algorithm, and determine the llama.cpp framework as the second preset large model inference framework;

[0029] Deploy the distilled model to the target device based on the preset model quantization algorithm, the second preset large model inference framework, a preset dynamic memory allocation mechanism, and a preset resource adaptation mechanism.

[0030] Optionally, deploying the distilled model to the target device based on a preset model quantization algorithm and a second preset large model inference framework includes:

[0031] After receiving the distillation model quantization requirements sent by the current target client, determine corresponding quantization schemes, and use the preset model quantization algorithm and each quantization scheme to quantize the distillation model to obtain each quantized model;

[0032] Based on each of the quantized models, determine the model confusion metric, the first-character latency metric, and the character generation speed metric of the model inference service to determine the visualization metric, so that the target client can determine the model to be deployed from each of the quantized models based on the visualization metric;

[0033] Deploy the model to be deployed to the target device.

[0034] In a second aspect, the present application provides an end-to-end distillation deployment device for large models for low-computing-power devices, including:

[0035] A framework deployment module, configured to determine a first target data set from a preset database in a computing cluster of a target cloud based on the obtained problem domain, determine a target student model in a preset student model library using the configuration information of the obtained target device, and deploy a distillation model training framework; the target device is a device whose own computing power meets the preset low-computing-power determination condition;

[0036] A first model deployment module, configured to deploy a preset large language model to the computing cluster using a first preset large model inference framework and a graphics card of the computing cluster, and determine the deployed preset large language model as a teacher model;

[0037] A first model distillation module, configured to, if the distillation model training framework is a black-box knowledge distillation framework, determine a second target data set based on the teacher model and the first target data set, and use the second target data set to distill the target student model to obtain a distillation model;

[0038] A second model distillation module, configured to, if the distillation model training framework is a white-box knowledge distillation framework, distill the target student model based on the output probability distribution of the teacher model to obtain the distillation model;

[0039] A second model deployment module, configured to deploy the distillation model to the target device based on a preset model quantization algorithm and a second preset large model inference framework.

[0040] In a third aspect, the present application provides an electronic device, including:

[0041] A memory, configured to store a computer program;

[0042] A processor, configured to execute the computer program to implement the foregoing end-to-end distillation deployment method for large models for low-computing-power devices.

[0043] Fourthly, the present application provides a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, the foregoing end-to-end distillation deployment method for large models on low-computing-power devices is implemented.

[0044] In the present application, in the computing cluster of the target cloud, a first target data set is determined from a preset database based on the obtained problem domain, a target student model is determined in a preset student model library by using the obtained configuration information of the target device, and a distillation model training framework is deployed; the target device is a device whose own computing power meets the preset low-computing-power determination condition; the preset large language model is deployed to the computing cluster by using a first preset large model inference framework and the graphics cards of the computing cluster, and the deployed preset large language model is determined as the teacher model; if the distillation model training framework is a black-box knowledge distillation framework, a second target data set is determined based on the teacher model and the first target data set, and the target student model is distilled by using the second target data set to obtain a distillation model; if the distillation model training framework is a white-box knowledge distillation framework, the target student model is distilled based on the output probability distribution of the teacher model to obtain the distillation model; the distillation model is deployed to the target device based on a preset model quantization algorithm and a second preset large model inference framework. As can be seen from the above, in the present application, in the computing cluster of the target cloud, first, according to the obtained problem domain, a first target data set is selected from the preset database. At the same time, with the help of the obtained configuration information of the target device, a target student model is selected in the preset student model library, and the deployment of the distillation model training framework is completed. The target device here refers to a device whose own computing power meets the preset low-computing-power determination condition. Then, the preset large language model is deployed to the computing cluster by using the first preset large model inference framework and the graphics cards of the computing cluster, and the deployed preset large language model is set as the teacher model. Subsequently, different treatments are performed according to the type of the distillation model training framework. If the distillation model training framework belongs to the black-box knowledge distillation framework, then based on the teacher model and the first target data set, a second target data set is determined, and then the target student model is distilled by using the second target data set to obtain a distillation model. If the distillation model training framework is a white-box knowledge distillation framework, the target student model is distilled based on the output probability distribution of the teacher model to obtain a distillation model. Finally, the obtained distillation model is deployed to the target device through a preset model quantization algorithm and a second preset large model inference framework. In this way, the present application can strengthen the privacy protection of users during the process of accessing the large language model, improve the efficiency of accessing the large language model, and thus enhance the user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on the provided drawings.

[0046] Figure 1 It is a flowchart of an end-to-end distillation deployment method for large models targeting low-computing-power devices disclosed in this application;

[0047] Figure 2 It is a flowchart of a specific end-to-end distillation deployment method for large models targeting low-computing-power devices disclosed in this application;

[0048] Figure 3 It is a schematic structural diagram of an end-to-end distillation deployment device for large models targeting low-computing-power devices disclosed in this application;

[0049] Figure 4 It is a structural diagram of an electronic device disclosed in this application. Specific embodiments

[0050] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0051] Currently, users generally access cloud services through mobile devices to use large language models to perform daily tasks. However, this operation mode has many drawbacks. From the perspective of information security, when users use mobile devices to access cloud large language model services, information security faces severe challenges. Due to the complexity of the network environment and security protection vulnerabilities, user information is extremely easy to be leaked during the access process, which may lead to important information such as personal privacy and business secrets being obtained by lawbreakers, causing serious losses to users. In terms of service response, when a large number of service requests flood into the computing cluster deploying the large language model at the same time, the computing cluster will be overwhelmed and frequently experience service response timeouts, seriously affecting the user experience. Therefore, this application provides an end-to-end distillation deployment method, device, equipment, and medium for large models targeting low-computing-power devices, which can strengthen user privacy protection during the process of users accessing large language models and improve the efficiency of accessing large language models.

[0052] See Figure 1 As shown, the embodiments of the present invention disclose an end-to-end distillation deployment method for large models targeting low-computing-power devices, including:

[0053] Step S11: In the computing cluster of the target cloud, determine the first target data set from a preset database based on the obtained problem domain, determine the target student model in the preset student model library using the configuration information of the obtained target device, and deploy a distillation model training framework; the target device is a device whose own computing power meets the preset low-computing-power determination condition.

[0054] In this embodiment, it should be noted first that for the computing cluster of the target cloud, the computing cluster is composed of multiple nodes, and each node is configured with multiple GPU (i.e., Graphics Processing Unit) cards and single or multiple high-performance CPUs. Before starting the distillation deployment task, the system will automatically collect the computing resource information of the cluster, including key parameters such as the CPU usage rate of each node, the GPU video memory usage, and the network bandwidth. These information provide certain information for the subsequent model distillation process.

[0055] When the target cloud receives an end-to-end distillation deployment instruction for a low-computing-power device, it first determines the problem domain for this distillation task. The determination of the problem domain can be achieved in various ways. For example, when the user initiates a distillation deployment request, the user directly enters a specific problem domain label in the user interface. When the label is received, it is matched with the domain classification in the preset database to locate the first target data set.

[0056] At the same time, it is necessary to obtain the configuration information of the target device. The target device refers to a device whose own computing power meets the preset low-computing-power determination condition. The collection of device configuration information can be achieved by installing a lightweight information collection tool on the device side. This information collection tool can automatically collect hardware parameters such as the processor model, memory size, and storage capacity of the target device and upload this information to the cloud in real time. After the target cloud receives the configuration information, it is matched with the model configuration requirements in the preset student model library to determine the target student model. Among them, the preset student model library can store multiple small-sized large models with different sizes and complexities, and these small-sized large models are optimized for different low-computing-power device configurations to ensure that they can run on the target device.

[0057] After determining the first target data set and the target student model, a distillation model training framework will be deployed on the computing cluster in the target cloud. Among them, this distillation model training framework is the core component for realizing the knowledge distillation of large models, and it supports multiple distillation methods, including black-box knowledge distillation and white-box knowledge distillation. During the framework deployment process, the computing resources can be reasonably allocated to the distillation model training framework according to the computing resource status of the current cluster. For example, for nodes with relatively tight computing resources, the system will preferentially deploy lightweight distillation model training modules to avoid training task failures due to insufficient resources. At the same time, the distillation model training framework can be initialized and configured, including setting training parameters, loading necessary dependency libraries, etc., to ensure that the framework can run normally.

[0058] Step S12: Use the first preset large model inference framework and the graphics cards of the computing cluster to deploy the preset large language model to the computing cluster, and determine the deployed preset large language model as the teacher model.

[0059] In this embodiment, the vLLM framework (an open-source large model inference acceleration framework) is determined as the first preset large model inference framework because the vLLM framework has the advantage of high-concurrency inference and can effectively improve the inference performance of large models in the computing cluster. Moreover, the first model deployment operation is performed on the preset large language model by using the paged attention mechanism and the preset scheduler of the first preset large model inference framework. Among them, the paged attention mechanism (i.e., Page Attention) and the preset scheduler (i.e., Scheduler) are important components in the vLLM framework, and they can effectively manage the paging of GPU video memory within the server and nodes.

[0060] It should be noted that the preset large language model in this embodiment has been widely applied in multiple fields and can effectively provide intelligent support for various industries. In a specific implementation, for the medical industry, the demand for patient health consultation and diagnostic support is huge. The large language model can be used as an intelligent medical assistant to help patients and medical workers. Specifically, when a patient inputs a description of symptoms, the model can give a preliminary analysis of the condition based on a vast amount of medical knowledge and clinical experience, and recommend corresponding measures, such as suggesting rest, dietary adjustments, or recommending further examination items. Medical workers can use the large language model to query information on complex diseases and obtain references for diagnostic ideas and treatment plans. Especially in remote areas or regions with scarce medical resources, low-computing-power medical devices equipped with this model, such as portable diagnostic instruments and mobile medical terminals, can enable patients to quickly obtain professional medical advice without relying on powerful computing power, making up for the shortage of medical resources and improving the accessibility of medical services. In another specific implementation, for the field of intelligent customer service, a large number of customer consultations need to be processed every day, and it is difficult for traditional human customer service to meet the efficiency requirements. The large language model empowers intelligent customer service to quickly understand customer demands. On e-commerce platforms, when customers consult product information and logistics delivery progress, the model can quickly give accurate answers. In online travel services, when customers consult travel routes and hotel reservations, the model can also provide reasonable suggestions. Through multi-round conversations, the model can solve complex problems and improve customer satisfaction. For the model deployed on low-computing-power devices such as enterprise local servers, enterprises do not need to invest a large amount of funds to upgrade hardware, but can optimize the customer service process with the help of the model, reduce labor costs, achieve 24 / 7 uninterrupted service, respond to customer needs in a timely manner, and enhance the competitiveness of enterprises.

[0061] Furthermore, in this embodiment, the second model deployment operation is performed on the preset large language model by using the flash attention mechanism of the first preset large model inference framework. Among them, the flash attention mechanism is a core technology in the vLLM framework, which can effectively handle the Attention (i.e., attention) calculation process in the inference process of the large language model. It should be noted that when the large model processes long sequence data, the Attention calculation often becomes a performance bottleneck. Especially the problem of excessive memory access frequency of Softmax (a mathematical function widely used in large language models) will lead to an increase in processing latency. The flash attention mechanism optimizes the Softmax memory access operation, reduces the number of memory accesses, and thus effectively reduces the processing latency of the large language model in the case of processing long sequences.

[0062] Finally, based on the first model deployment operation and the second model deployment operation, the preset large language model is successfully deployed to the computing cluster. Through the synergistic effect of these two operations, the advantages of the vLLM framework are fully utilized, improving the inference performance and processing efficiency of the model. Additionally, the running status of the preset large language model deployed to the computing cluster can be monitored in real time to ensure its stable operation in the computing cluster. If any abnormalities occur during the operation of the preset large language model, such as out-of-memory errors or calculation errors, adjustments and processing can be carried out in a timely manner to ensure the normal operation of the preset large language model.

[0063] In this embodiment, the deployed preset large language model is determined as the teacher model. It should be noted that the teacher model plays a crucial role in the knowledge distillation process and will provide supervision and guidance for the subsequent distilled model. The teacher model has strong language understanding and generation capabilities and can accurately reason and predict the input corpus data.

[0064] When receiving a model conversation request corresponding to the preset large language model sent by the target client, the model conversation request needs to be processed to determine the corresponding model conversation parameters. Specifically, based on the model conversation request, the corresponding model access path, the IP address and port number of the server where the preset large language model is located, the tensor parallel size of the preset large language model, and the GPU utilization rate of the server will be determined. Among them, the model access path specifies how to access the preset large language model, which can be a local file path or a network address. The IP address and port number of the server where the preset large language model is located are used to determine the location of the server to ensure that the system can communicate with the server accurately. The tensor parallel size determines the parallel computing method of the model on multiple GPUs, and a reasonable tensor parallel size can improve the computing efficiency of the model. The GPU utilization rate of the server reflects the usage of the GPU, and the system can dynamically adjust the computing resource allocation of the model according to the GPU utilization rate.

[0065] Next, model generation information is generated based on the model conversation parameters and the preset large language model. The model generation information is the response of the preset large language model to the user input and can be in the form of text, images, etc. It can be understood that the generated model generation information is output to the target client. During the output process, the model generation information can be formatted to ensure its readability and usability. At the same time, the transmission process of the model generation information is encrypted to ensure the security of the information.

[0066] Step S13: If the distillation model training framework is a black-box knowledge distillation framework, then based on the teacher model and the first target dataset, a second target dataset is determined, and the target student model is distilled using the second target dataset to obtain a distilled model.

[0067] In this embodiment, under the black-box knowledge distillation framework, first, a second target dataset needs to be generated with the help of a teacher model and a first target dataset. The first target dataset is determined from a preset database based on the obtained problem domain in step S11. The teacher model is a preset large language model deployed to a computing cluster through a first preset large model inference framework in step S12.

[0068] Specifically, the teacher model will generate data classification probabilities, confidence levels, and intermediate layer output results based on the first target dataset, and determine the data classification probabilities, confidence levels, and intermediate layer output results as the second target dataset. Among them, the data classification probability reflects the judgment probability of the teacher model for the category to which the input data belongs, the confidence level indicates the degree of certainty of the teacher model for this judgment, and the intermediate layer output result contains information such as the activation values of the intermediate layer neurons during the process of the model processing the input data. These information combined provide a certain number of supervision signals for the distillation of the target student model.

[0069] To achieve effective model distillation, it is necessary to determine an appropriate target distillation loss function. First, a soft-label distillation loss function and a hard-label distillation loss function are determined. Specifically, the soft-label distillation loss function usually uses the probability distribution output by the teacher model as the supervision signal, which can enable the target student model to learn the generalization ability of the teacher model. The hard-label distillation loss function is based on the true label information and usually uses the cross-entropy loss function. It can enable the target student model to directly learn the correct classification result.

[0070] Then, based on a preset dynamic loss weight algorithm, the soft-label distillation loss function, and the hard-label distillation loss function, the target distillation loss function is determined. The role of the preset dynamic loss weight algorithm is to dynamically adjust the weights of the soft-label distillation loss and the hard-label distillation loss during the training process. At the beginning of the training, the weight of the soft-label distillation loss can be set relatively large, allowing the target student model to learn more about the generalization ability of the teacher model. As the training progresses, the weight of the hard-label distillation loss gradually increases, enabling the target student model to pay more attention to the correct classification result.

[0071] Finally, the target student model is distilled using the second target dataset and the target distillation loss function. In each round of training during the distillation process, the samples in the second target dataset are input into the target student model, and the output results of the target student model are calculated. Then, according to the target distillation loss function, the loss value of the target student model is calculated. Through the backpropagation algorithm, the gradient of the loss value with respect to the parameters of the target student model is calculated, and the parameters of the target student model are updated using an optimization algorithm. After multiple rounds of training and parameter updates, the target student model gradually learns the knowledge and capabilities of the teacher model, and finally obtains the distilled model. This distilled model retains some of the capabilities of the teacher model while having a smaller size and lower computational complexity, making it suitable for deployment and operation on low-computing-power devices.

[0072] Step S14: If the distillation model training framework is a white-box knowledge distillation framework, then the target student model is distilled based on the output probability distribution of the teacher model to obtain the distilled model.

[0073] In this embodiment, under the white-box knowledge distillation framework, the output probability distribution of the teacher model is first obtained. That is, based on the teacher model, the attention weights, the hidden states of the intermediate layers, and the prediction results of the output layer are determined, and the attention weights, hidden states, and prediction results are determined as the output probability distribution of the teacher model. Among them, the attention weights reflect the degree of attention of the teacher model to different parts when processing the input data, and it is crucial for capturing the semantic information of the text in natural language processing. The hidden states of the intermediate layers contain the intermediate feature representations of the teacher model during the processing. These hidden states record the information after the input data has undergone a series of neural network layer transformations and are the internal representations learned by the teacher model. Different intermediate layer hidden states may capture different levels of semantic information, gradually abstracting from local features to global features. The prediction results of the output layer are the final judgments of the teacher model on the input data, usually presented in the form of a probability distribution.

[0074] It should be noted that the target student model is selected from a preset student model library according to the configuration information of the target device, aiming to be able to operate efficiently on low-computing-power devices. Therefore, during the distillation process, it is necessary to determine the output probability distribution of the target student model.

[0075] When the same data is input into the target student model, the target student model performs calculations and inferences through its own neural network structure. During the calculation process, the target student model generates attention weights, intermediate layer hidden states, and output layer prediction results similar to those of the teacher model. These results together constitute the output probability distribution of the target student model. Since the structure and parameters of the target student model are set during initialization, there may be a large difference between its output probability distribution in the initial state and that of the teacher model. The distillation of the model is to make the output probability distribution of the target student model gradually approach that of the teacher model.

[0076] To more accurately achieve the knowledge transfer from the target student model to the teacher model, the KL divergence method and the preset temperature scaling technique are used for distillation. KL divergence (i.e., Kullback-Leibler divergence) is a method for measuring the difference between two probability distributions. In knowledge distillation, by calculating the KL divergence between the output probability distribution of the teacher model and that of the target student model, the degree of difference between the two can be obtained.

[0077] Specifically, the output probability distributions of the corresponding layers of the teacher model and the target student model are input into the KL divergence calculation unit, and the KL divergence between the two is calculated layer by layer. For example, for the attention weights, intermediate layer hidden states, and output layer prediction results, their KL divergences are calculated respectively. By minimizing these KL divergences, the output probability distribution of the target student model can gradually approach that of the teacher model. During the calculation of the KL divergence, the preset temperature scaling technique can be used to smooth the probability distribution. The temperature scaling technique adjusts the output probability of the model by introducing a temperature parameter.

[0078] During the distillation process, by continuously adjusting the parameters of the target student model, the KL divergence between its output probability distribution and that of the teacher model gradually decreases. In each round of training, the input data is input into both the teacher model and the target student model simultaneously, and their output probability distributions are obtained respectively. The KL divergence between them is calculated as the loss function, and then the parameters of the target student model are updated according to the gradient of the loss function. After multiple rounds of training and parameter updates, the output probability distribution of the target student model gradually approaches that of the teacher model, and finally the distilled model is obtained.

[0079] Step S15: Deploy the distilled model to the target device based on the preset model quantization algorithm and the second preset large model inference framework.

[0080] In this embodiment, the SmoothQuant algorithm is used as the preset model quantization algorithm, and the llama.cpp framework is used as the second preset large model inference framework. Among them, the SmoothQuant algorithm is an effective model quantization algorithm, which can retain the performance of the model as much as possible while reducing the model storage space and computational complexity. By quantizing the weights and activation values of the model, the original high-precision floating-point representation is converted into a low-precision integer representation, thereby significantly reducing the storage requirements and computational complexity of the model. The llama.cpp framework is a lightweight large model inference framework with efficient inference performance and low resource usage, which is very suitable for running on low-computing power devices.

[0081] Furthermore, a pre-set dynamic memory allocation mechanism plays a key role in the distillation model deployment process. Target devices often have limited memory resources, especially those with low computing power. This dynamic memory allocation mechanism dynamically allocates and releases memory based on the model's operating status and actual needs. During model inference, when large amounts of data need to be processed, the system automatically allocates more memory resources. Once data processing is complete, unused memory is promptly released to avoid memory waste.

[0082] The pre-set resource adaptation mechanism automatically adjusts the distillation model's operating parameters and calculation methods based on the target device's hardware resources and real-time load. If the target device's CPU performance is low, the distillation model's computational complexity can be automatically reduced, using simpler calculation methods to ensure inference speed. If the target device's memory is insufficient, the distillation model's cached data can be reduced to avoid memory overflow. This adaptive mechanism ensures stable operation of the distillation model in diverse hardware environments.

[0083] After receiving the distillation model quantization requirements from the target user, the corresponding quantization schemes need to be determined. The choice of quantization scheme directly affects the performance and resource usage of the quantized model. Different quantization schemes can be developed based on different quantization accuracies and quantization ranges. For scenarios with high performance requirements, higher quantization accuracies can be selected, while for scenarios with strict resource requirements, lower quantization accuracies can be selected.

[0084] It is understood that the distillation model is quantized using the preset model quantization algorithm and various quantization schemes to obtain the quantized models. During the quantization process, it is important to control the quantization error. Although quantization can reduce the model's storage space and computational complexity, it also introduces a certain amount of quantization error, which may affect the model's performance. The SmoothQuant algorithm can effectively reduce quantization error and improve the performance of the quantized model by smoothing the model's weights and activation values.

[0085] Next, based on each quantized model, determine the model confusion degree index, the first-character delay index of the model inference service, and the character generation speed index. The model confusion degree index reflects the accuracy and stability of the model in classification or prediction tasks. If the confusion degree of the model is high, it means that the model is prone to confusion when processing data of different categories and has poor performance. The first-character delay index of the model inference service represents the time interval from the user's input request to the model's output of the first character, and this index is very important for real-time interaction scenarios. The character generation speed index represents the number of characters generated by the model per unit time, reflecting the inference efficiency of the model.

[0086] Determine the above-mentioned model confusion degree index, the first-character delay index of the model inference service, and the character generation speed index as visualization indexes, so that the target user side can determine the model to be deployed from each quantized model based on the visualization indexes. Through visualization, users can intuitively compare the performance and resource occupancy of different quantized models, and thus select the model that best suits their needs.

[0087] Finally, deploy the model to be deployed to the target device. During the deployment process, it is necessary to ensure that the model is compatible with the hardware and software environment of the target device. The llama.cpp framework provides good cross-platform support and can run on different operating systems and hardware architectures. At the same time, combined with the preset dynamic memory allocation mechanism and the preset resource adaptive mechanism, ensure that the model can run stably and efficiently on the target device and provide high-quality services for users.

[0088] As can be seen from the above, in the computing cluster of the target cloud in this application, first, according to the obtained problem domain, the first target data set is screened out from the preset database. At the same time, with the help of the obtained configuration information of the target device, the target student model is selected in the preset student model library, and the deployment of the distillation model training framework is completed. Here, the target device refers to a device whose computing power meets the preset low-computing-power determination condition. Then, using the first preset large model inference framework and the graphics card of the computing cluster, the preset large language model is deployed into the computing cluster, and the deployed preset large language model is set as the teacher model. Subsequently, different processes are carried out according to the type of the distillation model training framework. If the distillation model training framework belongs to the black-box knowledge distillation framework, then based on the teacher model and the first target data set, the second target data set is determined, and then the second target data set is used to perform distillation operations on the target student model to obtain the distillation model. If the distillation model training framework is a white-box knowledge distillation framework, then the target student model is distilled based on the output probability distribution of the teacher model to obtain the distillation model. Finally, through the preset model quantization algorithm and the second preset large model inference framework, the obtained distillation model is deployed to the target device. In this way, this application can strengthen the privacy protection of users during the process of accessing the large language model, improve the efficiency of accessing the large language model, and thus enhance the user experience.

[0089] Next, in combination with the medical Q&A scenario and Figure 2 the schematic diagram shown below, the technical solution of the embodiment of this application will be specifically described.

[0090] Specifically, first, preliminary preparation work is carried out in the computing cluster of the target cloud. Based on the specific scenario of medical Q&A, the first target data set is screened out from the preset medical database. This database can cover a vast amount of medical literature, and these literatures can include cutting-edge medical research results and classic case analysis reports. The database can also contain various clinical cases, which can be common diseases or rare diseases. And the database can also contain common disease Q&A data, providing certain materials for subsequent model training. At the same time, through the information collection module, the configuration information of grass-roots medical devices is collected. These configuration information can include key parameters such as the processor performance, memory capacity, and storage capacity of grass-roots medical devices. According to these configuration information, a match is made in the preset student model library to determine the target student model suitable for grass-roots medical devices, and then the distillation model training framework is deployed immediately to lay the foundation for subsequent model training.

[0091] Next, the teacher model is deployed. That is, using the first preset large model inference framework vLLM, leveraging its high-concurrency inference advantage and efficient utilization characteristics of GPU resources, and combined with the graphics cards of the computing cluster, the preset medical Q&A large language model is deployed to the computing cluster. After the deployment is completed, this preset medical Q&A large language model serves as the teacher model to provide certain guidance for knowledge distillation.

[0092] In the knowledge distillation process, if the black-box knowledge distillation framework is adopted, the teacher model will analyze the first target data set. Taking the symptom description of a patient as an example, the teacher model will generate data classification probabilities, confidence levels, and intermediate layer output results based on its own computing power and knowledge reserve, and then determine the second target data set. For example, in the face of the symptom description of a patient with "cough, fever, and fatigue", the teacher model can give the classification probabilities of possible diseases such as cold, flu, and pneumonia, and give the corresponding diagnostic confidence levels. Then, the second target data set is used to conduct distillation training on the target student model to guide the student model to learn the diagnostic ideas and knowledge of the teacher model. If the white-box knowledge distillation framework is adopted, the teacher model will output attention weights, intermediate layer hidden states, and output layer prediction results, and conduct distillation on the target student model based on these output probability distributions to help the student model understand the reasoning process of the teacher model.

[0093] Finally, based on the SmoothQuant algorithm and the llama.cpp framework, the distilled model is compressed and optimized and deployed to grass-roots medical devices. When a doctor is making a diagnosis, he only needs to input the symptom information of the patient into the device, and with the help of the model in the device, he can quickly obtain professional diagnostic suggestions, greatly improving the diagnostic efficiency and enabling grass-roots patients to enjoy better medical services.

[0094] Correspondingly, as shown in Figure 3 The embodiment of the present application provides an end-to-end distillation deployment device for large models for low-computing-power devices, including:

[0095] A framework deployment module 11, configured to determine a first target data set from a preset database in a computing cluster of a target cloud based on the obtained problem domain, determine a target student model in a preset student model library using the configuration information of the obtained target device, and deploy a distillation model training framework; the target device is a device whose own computing power meets the preset low-computing-power determination condition;

[0096] A first model deployment module 12, configured to deploy a preset large language model to the computing cluster using a first preset large model inference framework and the graphics cards of the computing cluster, and determine the deployed preset large language model as the teacher model;

[0097] The first model distillation module 13 is used to, if the distillation model training framework is a black-box knowledge distillation framework, determine a second target dataset based on the teacher model and the first target dataset, and use the second target dataset to distill the target student model to obtain a distillation model;

[0098] The second model distillation module 14 is used to, if the distillation model training framework is a white-box knowledge distillation framework, distill the target student model based on the output probability distribution of the teacher model to obtain the distillation model;

[0099] The second model deployment module 15 is used to deploy the distillation model to the target device based on a preset model quantization algorithm and a second preset large model inference framework.

[0100] As can be seen from the above, in this application, in the computing cluster of the target cloud, first, according to the obtained problem domain, a first target dataset is screened out from a preset database. At the same time, with the help of the obtained configuration information of the target device, a target student model is selected from a preset student model library, and the deployment of the distillation model training framework is completed. Here, the target device refers to a device whose own computing power meets the preset low computing power determination condition. Then, using the first preset large model inference framework and the graphics card of the computing cluster, the preset large language model is deployed into the computing cluster, and the deployed preset large language model is set as the teacher model. Subsequently, different processing is performed according to the type of the distillation model training framework. If the distillation model training framework belongs to the black-box knowledge distillation framework, then based on the teacher model and the first target dataset, a second target dataset is determined, and then the second target dataset is used to perform a distillation operation on the target student model to obtain a distillation model. If the distillation model training framework is a white-box knowledge distillation framework, then the target student model is distilled based on the output probability distribution of the teacher model to obtain a distillation model. Finally, through a preset model quantization algorithm and a second preset large model inference framework, the obtained distillation model is deployed to the target device. In this way, this application can strengthen the privacy protection of users during the process of accessing the large language model, improve the efficiency of accessing the large language model, and thus enhance the user experience.

[0101] In some specific embodiments, the first model deployment module 12 specifically includes:

[0102] The first operation execution unit is used to determine the vLLM framework as the first preset large model inference framework, and use the page attention mechanism and the preset scheduler of the first preset large model inference framework to perform a first model deployment operation on the preset large language model;

[0103] A second operation execution unit, configured to perform a second model deployment operation on the preset large language model by using the flash attention mechanism of the first preset large model inference framework;

[0104] A first model deployment unit, configured to deploy the preset large prediction model to the computing cluster based on the first model deployment operation and the second model deployment operation.

[0105] In some specific embodiments, the first model deployment module 12 further specifically includes:

[0106] A parameter determination unit, configured to, when receiving a model conversation request corresponding to the preset large language model sent by a target client, determine a corresponding model access path, the IP address and port number of the server where the preset large language model is located, the tensor parallelism size of the preset large language model, and the graphics card utilization rate of the server, so as to determine model conversation parameters;

[0107] An information output unit, configured to determine model generation information based on the model conversation parameters and the preset large language model, and output the model generation information to the target client.

[0108] In some specific embodiments, the first model distillation module 13 specifically includes:

[0109] A dataset determination unit, configured to generate data classification probabilities, confidence levels, and intermediate layer output results of the teacher model based on the first target dataset, and determine the data classification probabilities, the confidence levels, and the intermediate layer output results as a second target dataset;

[0110] A function determination unit, configured to determine a soft label distillation loss function and a hard label distillation loss function, and determine a target distillation loss function based on a preset dynamic loss weight algorithm, the soft label distillation loss function, and the hard label distillation loss function;

[0111] A first model distillation unit, configured to distill the target student model by using the second target dataset and the target distillation loss function, and obtain a distilled model.

[0112] In some specific embodiments, the second model distillation module 14 specifically includes:

[0113] A first probability distribution determination unit, configured to determine attention weights, hidden states of the intermediate layer, and prediction results of the output layer based on the teacher model, and determine the attention weights, the hidden states, and the prediction results as the output probability distribution of the teacher model;

[0114] A second probability distribution determination unit for determining the output probability distribution of the target student model;

[0115] A second model distillation unit for distilling the target student model based on the KL divergence method, a preset temperature scaling technique, the output probability distribution of the teacher model, and the output probability distribution of the target student model to obtain the distilled model.

[0116] In some specific embodiments, the second model deployment module 15 specifically includes:

[0117] A framework determination unit for taking the algorithm as a preset model quantization algorithm and determining the llama.cpp framework as the second preset large model inference framework;

[0118] A second model deployment unit for deploying the distilled model to the target device based on the preset model quantization algorithm, the second preset large model inference framework, a preset dynamic memory allocation mechanism, and a preset resource adaptation mechanism.

[0119] In some specific embodiments, the second model deployment module 15 specifically includes:

[0120] A model quantization unit for determining corresponding quantization schemes after receiving the distilled model quantization requirement sent by the current target client, and quantizing the distilled model using the preset model quantization algorithm and each quantization scheme to obtain each quantized model;

[0121] A model determination unit for determining a model confusion degree index, a first character delay index of the model inference service, and a character generation speed index based on each quantized model to determine a visualization index, so that the target client determines the model to be deployed from each quantized model based on the visualization index;

[0122] A third model deployment unit for deploying the model to be deployed to the target device.

[0123] Furthermore, an embodiment of the present application also discloses an electronic device, Figure 4It is a structural diagram of an electronic device 20 shown according to an exemplary embodiment. The content in the figure should not be regarded as any limitation on the scope of use of this application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. Among them, the memory 22 is used to store a computer program, and the computer program is loaded and executed by the processor 21 to implement the relevant steps in the end-to-end distillation deployment method of the large model for low-computing-power devices disclosed in any of the foregoing embodiments. Additionally, the electronic device 20 in this embodiment may specifically be an electronic computer.

[0124] In this embodiment, the power supply 23 is used to provide working voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows can be any communication protocol applicable to the technical solution of this application, and no specific limitation is imposed on it here; the input / output interface 25 is used to obtain external input data or output data to the outside, and its specific interface type can be selected according to specific application needs, and no specific limitation is made here.

[0125] In addition, as a carrier for resource storage, the memory 22 can be a read-only memory, a random access memory, a magnetic disk, or an optical disc, etc. The resources stored thereon can include an operating system 221, a computer program 222, etc., and the storage method can be temporary storage or permanent storage.

[0126] Among them, the operating system 221 is used to manage and control each hardware device and the computer program 222 on the electronic device 20, and it can be Windows Server, Netware, Unix, Linux, etc. In addition to the computer program that can be used to complete the end-to-end distillation deployment method of the large model for low-computing-power devices executed by the electronic device 20 disclosed in any of the foregoing embodiments, the computer program 222 may further include computer programs that can be used to complete other specific tasks.

[0127] Furthermore, this application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the end-to-end distillation deployment method of the large model for low-computing-power devices disclosed above. For the specific steps of this method, reference can be made to the corresponding content disclosed in the foregoing embodiments, and no further elaboration will be provided here.

[0128] In the present specification, the various embodiments are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the embodiments, reference can be made to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple. For the relevant parts, reference can be made to the description in the method section.

[0129] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0130] The steps of the methods or algorithms described in combination with the embodiments disclosed herein can be directly implemented by hardware, software modules executed by a processor, or a combination of both. The software modules can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the technical field.

[0131] Finally, it should also be noted that in this document, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising a..." does not exclude the existence of additional identical elements in the process, method, article or device comprising the element.

[0132] The above has introduced the technical solutions provided by this application in detail. Specific examples are used herein to elaborate on the principles and implementation manners of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application; at the same time, for those of ordinary skill in the art, according to the idea of this application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to this application.

Claims

1. An end-to-end distillation deployment method for large models on low-computing-power devices, characterized in that Including: In the computing cluster of the target cloud, determine the first target data set from the preset database based on the obtained problem domain, determine the target student model in the preset student model library using the obtained configuration information of the target device, and deploy the distillation model training framework; the target device is a device whose own computing power meets the preset low computing power determination condition; Deploy the preset large language model to the computing cluster using the first preset large model inference framework and the graphics card of the computing cluster, and determine the deployed preset large language model as the teacher model; If the distillation model training framework is a black box knowledge distillation framework, determine the second target data set based on the teacher model and the first target data set, and use the second target data set to distill the target student model to obtain the distillation model; If the distillation model training framework is a white box knowledge distillation framework, distill the target student model based on the output probability distribution of the teacher model to obtain the distillation model; Deploy the distillation model to the target device based on the preset model quantization algorithm and the second preset large model inference framework; Among them, the deploying the preset large language model to the computing cluster using the first preset large model inference framework and the graphics card of the computing cluster includes: Determine the vLLM framework as the first preset large model inference framework, and perform the first model deployment operation on the preset large language model using the page attention mechanism and the preset scheduler of the first preset large model inference framework; Perform the second model deployment operation on the preset large language model using the flash attention mechanism of the first preset large model inference framework; Deploy the preset large language model to the computing cluster based on the first model deployment operation and the second model deployment operation.

2. The end-to-end distillation deployment method for large models targeting low-computing-power devices according to claim 1, wherein, The deploying the preset large language model to the computing cluster using the first preset large model inference framework and the graphics card of the computing cluster further includes: When receiving a model conversation request corresponding to the preset large language model sent by the target client, determine the corresponding model access path, the IP address and port number of the server where the preset large language model is located, the tensor parallel size of the preset large language model, and the graphics card utilization rate of the server to determine the model conversation parameters; Determine the model generation information based on the model conversation parameters and the preset large language model, and output the model generation information to the target client.

3. The end-to-end distillation deployment method for large models targeting low-computing-power devices according to claim 1, wherein The determining the second target data set based on the teacher model and the first target data set, and using the second target data set to distill the target student model to obtain the distillation model includes: Generate the data classification probability, confidence, and the intermediate layer output result of the teacher model based on the first target data set through the teacher model, and determine the data classification probability, the confidence, and the intermediate layer output result as the second target data set; Determine the soft-label distillation loss function and the hard-label distillation loss function, and determine the target distillation loss function based on a preset dynamic loss weight algorithm, the soft-label distillation loss function, and the hard-label distillation loss function; Use the second target data set and the target distillation loss function to distill the target student model and obtain a distilled model.

4. The end-to-end distillation deployment method for large models targeting low-computing-power devices according to any one of claims 1 to 3, characterized in that The distilling the target student model based on the output probability distribution of the teacher model to obtain the distilled model includes: Determine the attention weights, the hidden states of the intermediate layer, and the prediction results of the output layer based on the teacher model, and determine the attention weights, the hidden states, and the prediction results as the output probability distribution of the teacher model; Determine the output probability distribution of the target student model; Distill the target student model based on the KL divergence method, a preset temperature scaling technique, the output probability distribution of the teacher model, and the output probability distribution of the target student model to obtain the distilled model.

5. The end-to-end distillation deployment method for large models targeting low-computing-power devices according to claim 1, wherein, The deploying the distilled model to the target device based on a preset model quantization algorithm and a second preset large model inference framework includes: Determine the SmoothQuant algorithm as the preset model quantization algorithm and determine the llama.cpp framework as the second preset large model inference framework; Deploy the distilled model to the target device based on the preset model quantization algorithm, the second preset large model inference framework, a preset dynamic memory allocation mechanism, and a preset resource adaptation mechanism.

6. The end-to-end distillation deployment method for large models targeting low-computing-power devices according to claim 1, wherein The deploying the distilled model to the target device based on a preset model quantization algorithm and a second preset large model inference framework includes: After receiving the distillation model quantization requirement sent by the current target client, determine corresponding quantization schemes, and use the preset model quantization algorithm and each quantization scheme to quantize the distilled model to obtain each quantized model; Determine the model confusion degree index, the first-character delay index of the model inference service, and the character generation speed index based on each quantized model to determine the visualization index, so that the target client can determine the model to be deployed from each quantized model based on the visualization index; Deploy the model to be deployed to the target device.

7. An end-to-end distillation deployment device for large models targeting low-computing-power devices, characterized in that, Includes: A framework deployment module, configured to determine a first target data set from a preset database in a computing cluster of a target cloud based on an obtained problem domain, determine a target student model in a preset student model library using the configuration information of the obtained target device, and deploy a distillation model training framework; the target device is a device whose own computing power meets a preset low computing power determination condition; A first model deployment module, configured to deploy a preset large language model to the computing cluster using a first preset large model inference framework and a graphics card of the computing cluster, and determine the deployed preset large language model as a teacher model; The first model distillation module is used to, if the distillation model training framework is a black-box knowledge distillation framework, determine a second target dataset based on the teacher model and the first target dataset, and use the second target dataset to distill the target student model to obtain a distillation model; The second model distillation module is used to, if the distillation model training framework is a white-box knowledge distillation framework, distill the target student model based on the output probability distribution of the teacher model to obtain the distillation model; The second model deployment module is used to deploy the distillation model to the target device based on a preset model quantization algorithm and a second preset large model inference framework; Among them, the first model deployment module includes: The first operation execution unit is used to determine the vLLM framework as the first preset large model inference framework, and use the page attention mechanism and the preset scheduler of the first preset large model inference framework to perform the first model deployment operation on the preset large language model; The second operation execution unit is used to perform the second model deployment operation on the preset large language model by using the flash attention mechanism of the first preset large model inference framework; The first model deployment unit is used to deploy the preset large language model to the computing cluster based on the first model deployment operation and the second model deployment operation.

8. An electronic device, characterized in that, Includes: A memory for storing a computer program; A processor for executing the computer program to implement the end-to-end distillation deployment method of the large model for low-computing-power devices according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, For storing a computer program; wherein, when the computer program is executed by the processor, the end-to-end distillation deployment method of the large model for low-computing-power devices according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Neural network black box aggressive defense method based on knowledge distillation

    CN111027060A

  • Large language model distillation method, device and equipment and storage medium

    CN119089975A