End-cloud collaborative reasoning method and system based on large language model, electronic equipment and product

Through the large language model combined with the end-cloud collaborative inference method, the problem of inference efficiency of edge intelligence technology in diversified application scenarios is solved, accurate identification and rapid response to user needs is achieved, and the model's inference efficiency in latency-sensitive or bandwidth-constrained environments is improved.

CN120373455APending Publication Date: 2025-07-25SOUTHWEST UNIVERSITY FOR NATIONALITIES
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510427511.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

Existing edge intelligence technologies are difficult to quickly and accurately confirm user needs, resulting in inaccurate inference efficiency in diversified application scenarios, and the computing power and storage resources of edge devices are limited, making it difficult to meet real-time needs.

Method used

A large language model is used for task analysis, and the data sets and candidate models matching the task category are obtained. The target model is divided into edge devices and cloud server deployment through model slicing points, and the end-cloud collaborative reasoning is performed, and the model scale is optimized through pruning processing to improve efficiency.

Benefits of technology

It realizes accurate identification and rapid response to user needs, improves the model's inference efficiency in latency-sensitive or bandwidth-constrained environments, and is suitable for diversified application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120373455A_ABST
    Figure CN120373455A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of computers, and aims to provide an end-cloud collaborative reasoning method and system based on a large language model, electronic equipment and a product. According to the method, the large language model and the edge intelligence technology are combined, accurate and rapid identification of the user demand information can be achieved, the target model meeting the user demand can be further obtained, the target model is segmented, and the user experience is improved. The target model is subjected to end-cloud cooperative reasoning by adopting an end-cloud cooperative reasoning mode, so that the reasoning efficiency of the model can be improved, the edge intelligence technology is enabled to respond more quickly in a delay-sensitive or bandwidth-limited environment, and the edge intelligence technology is further efficiently applied to diversified application scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer technology, and particularly relates to an edge-cloud collaborative inference method, system, electronic device, and product based on a large language model. Background Art

[0002] With the rapid development of the Internet of Things (IoT), 5G technology, and edge computing technology, Edge AI, as an innovative technology that deeply integrates artificial intelligence (AI) and edge computing, has become an important technology to promote the transformation of the data processing mode. Edge AI realizes fast response, local data processing, and privacy protection by sinking data analysis and processing capabilities to locations close to the data source. This distributed computing mode effectively solves problems such as high latency, bandwidth limitation, and privacy leakage existing in traditional cloud computing architectures.

[0003] However, in the process of using the existing technology, the inventor found that there are at least the following problems in the existing technology: The application scenarios of Edge AI are becoming increasingly complex, and the user requirements vary greatly. It is difficult for the existing technology to quickly and accurately confirm user requirements, making it difficult for Edge AI technology to efficiently handle diversified application scenarios. In addition, for different application scenarios, it is usually necessary to conduct targeted model training and deploy and infer on edge devices. However, existing edge devices are usually limited by computing power and storage resources, resulting in the inference efficiency of the model being difficult to meet real-time requirements. Summary of the Invention

[0004] The present invention aims to solve the above technical problems to at least a certain extent, and provides an edge-cloud collaborative inference method, system, electronic device, and product based on a large language model.

[0005] To achieve the above object, the present invention adopts the following technical solutions: In a first aspect, the present invention provides an edge-cloud collaborative inference method based on a large language model, which is executed by an edge device. The method includes: After receiving user requirement information, call a preset large language model, and perform task analysis and processing on the user requirement information based on the large language model to obtain a task category and task metrics; Obtain a data set and multiple candidate models that match the task category, and perform pre-training on the multiple candidate models based on the data set to obtain multiple pre-trained models; Extract a target model that matches the task metrics from the multiple pre-trained models based on the large language model; Obtain the model splitting point of the target model, and use the model splitting point to split the target model into a first sub-model and a second sub-model. Then deploy the first sub-model on the edge device and deploy the second sub-model on the cloud server corresponding to the edge device, so that the edge device and the cloud server perform end-cloud collaborative inference on the target model.

[0006] In a possible design, the task categories include image classification tasks, crop pest and disease identification tasks, and weed identification tasks. The dataset matching the image classification task is the CIFAR10 dataset, the dataset matching the crop pest and disease identification task is the AgriculturalDisease dataset, the dataset matching the weed identification task is the Kaggle archive dataset, and the multiple candidate models matching the image classification task, the crop pest and disease identification task, and the weed identification task all include the AlexNet model, the ResNet50 model, and the VGG16 model.

[0007] In a possible design, obtaining the model splitting point of the target model includes: Determine a candidate splitting point in the target model, and use the current candidate splitting point to split the target model into a first candidate sub-model and a second candidate sub-model; Deploy the first candidate sub-model on the edge device and deploy the second candidate sub-model on the cloud server corresponding to the edge device, so that the edge device and the cloud server perform end-cloud collaborative inference on the target model; Calculate the computing delay of the edge device, the data transmission delay from the edge device to the cloud server, and the computing delay of the cloud server through timestamps, and merge the computing delay of the edge device, the data transmission delay from the edge device to the cloud server, and the computing delay of the cloud server to obtain the total end-cloud collaborative inference delay of the current candidate splitting point; Traverse each candidate splitting point of the target model, obtain the candidate splitting point with the minimum total end-cloud collaborative inference delay among all candidate splitting points, and use the candidate splitting point with the minimum total end-cloud collaborative inference delay as the model splitting point of the target model, and use the total end-cloud collaborative inference delay of the model splitting point as the optimal total delay.

[0008] In a possible design, after using the model splitting point to split the target model into a first sub-model and a second sub-model, the method further includes: Respectively take the first sub-model and the second sub-model as the target sub-model, obtain the optimal pruning sparsity rate combination of the target sub-model, and perform pruning processing on the target sub-model based on the optimal pruning sparsity rate combination to obtain the pruned target sub-model, so as to deploy the pruned target sub-model corresponding to the first sub-model on the edge device and deploy the pruned target sub-model corresponding to the second sub-model on the cloud server.

[0009] In a possible design, a genetic algorithm is used to obtain the optimal pruning sparsity rate combination of the target sub-model.

[0010] In a possible design, when performing pruning processing on the target sub-model based on the optimal pruning sparsity rate combination, the torch.nn.utils.prune module in PyTorch is called to execute.

[0011] In a possible design, after obtaining the model splitting point of the target model, the method further includes: Input the target model and the model splitting point into the large language model, so that the large language model generates a task plan matching the user demand information according to the target model and the model splitting point.

[0012] In a second aspect, the present invention provides an edge-cloud collaborative inference method based on a large language model, including: A task analysis module, configured to, after receiving user demand information, call a preset large language model, and perform task analysis processing on the user demand information based on the large language model to obtain a task category and task metrics; A model training module, communicatively connected to the task analysis module, configured to obtain a data set and multiple candidate models matching the task category, and perform pre-training on the multiple candidate models respectively based on the data set to obtain multiple pre-trained models; A target model extraction module, communicatively connected to the model training module, configured to extract a target model matching the task metrics from the multiple pre-trained models based on the large language model; A collaborative inference module, communicatively connected to the target model extraction module, configured to obtain the model splitting point of the target model, and use the model splitting point to split the target model into a first sub-model and a second sub-model, and then deploy the first sub-model on an edge device and deploy the second sub-model on a cloud server corresponding to the edge device, so that the edge device and the cloud server perform edge-cloud collaborative inference on the target model.

[0013] In a third aspect, the present invention provides an electronic device, including: A memory for storing computer program instructions; and, A processor for executing the computer program instructions to complete the operations of a method for end - cloud collaborative inference based on a large language model as described in any one of the above.

[0014] In a fourth aspect, the present invention provides a computer program product including a computer program or instructions, where the computer program or the instructions, when executed by a computer, implement a method for end - cloud collaborative inference based on a large language model as described in any one of the above.

[0015] The beneficial effects of the present invention are as follows: The present invention discloses a method, system, electronic device, and product for end - cloud collaborative inference based on a large language model, which can be applied to diversified application scenarios and improve the model inference efficiency. Specifically, by combining the large language model with edge intelligence technology, the present invention can accurately and quickly identify user demand information, and further obtain a target model that meets user needs. In the present invention, the target model is also segmented, and the target model is subjected to end - cloud collaborative inference in a manner of end - cloud collaborative inference, which can improve the inference efficiency of the model, enabling edge intelligence technology to respond more quickly in environments sensitive to latency or bandwidth - limited, and then being efficiently applied to diversified application scenarios.

[0016] Other beneficial effects of the present invention will be further described in the specific implementation manners. Description of the Drawings

[0017] Figure 1 is a flowchart of a method for end - cloud collaborative inference based on a large language model in an embodiment; Figure 2 is a block diagram of a module of a method for end - cloud collaborative inference based on a large language model in an embodiment; Figure 3 is a block diagram of a module of an electronic device in an embodiment. Specific Implementation Manners

[0018] To more clearly illustrate the technical solutions in the embodiments of the present invention or in the prior art, the following will briefly introduce the present invention in combination with the drawings and the descriptions of the embodiments or the prior art. Obviously, the following descriptions of the structures of the drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts. It should be noted here that the descriptions of these embodiments are used to help understand the present invention, but do not constitute a limitation to the present invention.

[0019] Embodiment 1: This embodiment discloses an edge-cloud collaborative inference method based on a large language model, which is executed by an edge device. The edge device can be, but is not limited to, a computer device or a virtual machine with certain computing resources, such as an electronic device like a personal computer, a smartphone, a personal digital assistant, or a wearable device, or a virtual machine.

[0020] As Figure 1 shown, an edge-cloud collaborative inference method based on a large language model can, but is not limited to, include the following steps: S1. After receiving the user requirement information, call a preset large language model (Large Language Models, LLM), and perform task analysis and processing on the user requirement information based on the large language model to obtain a task category and task metrics. Specifically, the task category can be, for example, an image classification task, a crop pest and disease identification task, or a weed identification task, etc., and the task metrics can include model metrics requirement information such as model accuracy and inference latency, which are not limited here.

[0021] It should be noted that the structure of the large language model has achieved extensive context awareness through pre-training on large-scale unlabeled text, and has multiple functions such as user semantic understanding, task planning and scheduling, and code generation. It can adapt to various tasks with minimal adjustment, and is particularly suitable for the edge device environment with limited computing resources. In this embodiment, the large language model can accurately capture the nuances in natural language through deep and bidirectional context representations, and then perform task analysis and processing on the user requirement information, providing the edge device with strong adaptability and intelligent capabilities. As an example, in this embodiment, the large language model can be, but is not limited to, Tongyi Qianwen large model, Zidong Taichu large model, Doubao large model, or ChatGPT series large models, which are not limited here.

[0022] As an example, the user requirement information is: "I have a machine that can perform image acquisition. Now I need to complete an image classification task. Please give a suitable model to ensure an accuracy of 80% and minimize the model inference latency." Correspondingly, the large language model performs task analysis and processing on the user requirement information, and the obtained task category is: image classification task, and the task metrics are: the model accuracy is greater than 80% and the inference latency is minimized.

[0023] S2. Obtain a data set and multiple candidate models that match the task category, and perform pre-training on the multiple candidate models based on the data set to obtain multiple pre-trained models. It should be noted that in this embodiment, a knowledge base including various tasks, candidate models, and data sets respectively matching various tasks can be preset in advance to facilitate the extraction of candidate models and data sets according to the task category in a timely manner.

[0024] Specifically, in this embodiment, the initial model is, for example, an AlexNet model, a ResNet50 model, or a VGG16 model, etc. The sample dataset includes, for example, the CIFAR10 dataset (from the datasets in the torchvision library of the open-source deep learning framework PyTorch), the AgriculturalDisease dataset (from the Github platform), or the Kagglearchive dataset (from the Kaggle platform), etc., which is not limited here.

[0025] As an example, in this embodiment, when the task categories are an image classification task, a crop pest and disease identification task, and a weed identification task respectively, the datasets matching the above three task categories are the CIFAR10 dataset, the AgriculturalDisease dataset, and the Kaggle archive dataset respectively. The multiple candidate models matching the above three task categories all include the AlexNet model, the ResNet50 model, and the VGG16 model. At this time, multiple candidate models can be pre-trained based on the three groups of datasets respectively, and multiple pre-trained models can be obtained for subsequent calling of the pre-trained models.

[0026] In this embodiment, after pre-training the candidate models for the three task categories in the example, the accuracies of the multiple pre-trained models are shown in Table 1 below:

[0027] Table 1 S3. Extract a target model matching the task metrics from multiple pre-trained models based on the large language model. Specifically, in this embodiment, the decision-making and output of the target model are performed through the large language model.

[0028] S4. Obtain the model splitting point of the target model, and use the model splitting point to split the target model into a first sub-model and a second sub-model. Then, deploy the first sub-model on the edge device and deploy the second sub-model on the cloud server corresponding to the edge device, so that the edge device and the cloud server perform end-cloud collaborative inference on the target model. Specifically, in this embodiment, in the target model, the inference process of the first sub-model before the model splitting point is completed on the edge device, and the inference process of the second sub-model after the model splitting point is completed on the cloud server (i.e., the cloud).

[0029] In this embodiment, the model splitting point is obtained with the total end-cloud collaborative inference latency as the optimization goal. Specifically, in step S4, obtaining the model splitting point of the target model includes: S401. Determine a candidate splitting point a in the target model, and use the current candidate splitting point a to split the target model into a first candidate sub-model and a second candidate sub-model; specifically, the current candidate splitting point a can divide the target model into two parts. Among them, the part of the target model from the first layer to the (a - 1)th layer forms the first candidate sub-model, and the part from the a-th layer to the last layer forms the second candidate sub-model.

[0030] S402. Deploy the first candidate sub-model on the edge device, and deploy the second candidate sub-model on the cloud server corresponding to the edge device, so that the edge device and the cloud server perform end-cloud collaborative inference on the target model; S403. Calculate the computing delay of the edge device, the data transmission delay from the edge device to the cloud server, and the computing delay of the cloud server through timestamps, and merge the computing delay of the edge device, the data transmission delay from the edge device to the cloud server, and the computing delay of the cloud server to obtain the total end-cloud collaborative inference delay of the current candidate splitting point a; specifically, in this embodiment, the time module from the edge device to the cloud server (which can provide time-related functions for processing time access and conversion) can be used to take the time of model inference as an index to measure the model delay, and calculate by reading the system timestamps before and after loading and before and after inference, so as to obtain the delay data.

[0031] S404. Traverse each candidate splitting point of the target model, obtain the candidate splitting point with the minimum total end-cloud collaborative inference delay among each candidate splitting point, and use the candidate splitting point with the minimum total end-cloud collaborative inference delay as the model splitting point of the target model, and use the total end-cloud collaborative inference delay of the model splitting point as the optimal total delay.

[0032] Based on the above steps S401 to S404, the model splitting point with the minimum total end-cloud collaborative inference delay can be obtained, which is beneficial to improving the overall efficiency of end-cloud collaborative inference.

[0033] Specifically, for a neural network model containing a nested block structure, the first layer of each nested block structure can be used as a splitting point.

[0034] As an example, taking the task category as the image classification task and the target model as the ResNet50 model, if the splitting point is the a-th layer, after inputting the picture data, the edge device calculates layer by layer from the first layer in the ResNet50 model. When calculating to the (a - 1)-th layer, the intermediate feature data of the output data of this layer, data, is obtained. Subsequently, the edge device transmits the intermediate feature data to the cloud server, and the cloud server completes the inference of the part of the model after the splitting point, that is, starting from the a-th layer, inputting the intermediate feature data and calculating layer by layer until the last layer is calculated and the result is output. Then, by reading the system timestamp at key positions during the running process and performing calculations, the computing latency of the edge device (obtained by calculating the timestamps before and after model inference on the edge device), the data transmission latency from the edge device to the cloud server (obtained by calculating the timestamps before and after data transmission and reception), and the computing latency of the cloud server (obtained by calculating the timestamps before and after model inference on the cloud server) can be obtained.

[0035] In this embodiment, the total end-cloud collaborative inference latency at each candidate splitting point in the example is shown in Table 2 below:

[0036] Table 2 According to Table 2, it can be seen that selecting the layer4 layer as the model splitting point can minimize the total end-cloud collaborative inference latency of the target model. At this time, the model splitting point can be determined as the layer4 layer.

[0037] Similarly, other models can be traversed and selected for splitting points, and the optimal splitting point selections for each model are shown in Table 3 below:

[0038] Table 3 In step S4, after splitting the target model into the first sub-model and the second sub-model using the model splitting point, the method further includes: A. Respectively taking the first sub-model and the second sub-model as the target sub-models, obtaining the optimal pruning sparsity rate combination of the target sub-model, and performing pruning processing on the target sub-model based on the optimal pruning sparsity rate combination to obtain the pruned target sub-model, so as to deploy the pruned target sub-model corresponding to the first sub-model on the edge device and deploy the pruned target sub-model corresponding to the second sub-model on the cloud server. It should be noted that performing pruning processing on the target sub-model can greatly compress the model scale, facilitate model transmission, and reduce the computational complexity of the model, so that a model with a complex structure can run efficiently on the edge device and the cloud server, improving the model inference efficiency.

[0039] Specifically, when performing edge collaborative inference on the target sub-model, since data transmission is required and model loading needs to be performed both on the local edge device side and the cloud server side, the latency will increase. Therefore, in this embodiment, the target sub-model is pre-pruned, and a certain sparsity rate is set for the convolutional layer (Conv2d) and linear layer (Linear) of the target sub-model to reduce the model size and optimize the accuracy loss. Thereby, the computational complexity in the subsequent edge-cloud collaborative inference process can be reduced, the inference speed can be improved, and energy consumption can be reduced.

[0040] Specifically, in this embodiment, a genetic algorithm is used to obtain the optimal pruning sparsity rate combination of the target sub-model. Specifically, in this embodiment, after obtaining the optimal pruning sparsity rate combination of the target sub-model, the time required for the model to process a single test sample can be calculated as the latency. If the accuracy loss is too large or even lower than the user's accuracy requirement (such as 80%), the latency of the model will be penalized, and its latency will be increased to 100 times the original to eliminate this selection.

[0041] As an example, taking the task category as an image classification task and the target sub-model as the ResNet50 model, the population size per generation is set to 10, and 5 rounds of iterative updates are performed. Each generation is generated by crossbreeding and mutation of the high-quality (low-latency) parent generations of the previous generation. After 5 generations of crossbreeding and mutation processes, the optimal pruning sparsity rate combination can finally be obtained as [0.17, 0.29, 0.11, ……, 0.11, 0.33, 0.09, 0.18] (a total of 53 convolutional layers plus linear layers in the ResNet50 model).

[0042] In this embodiment, when pruning the target sub-model based on the optimal pruning sparsity rate combination, the torch.nn.utils.prune module in PyTorch can be called to execute. In this embodiment, the total latency of the pruned target sub-model obtained after pruning is 5.7732, and the accuracy is 0.859, indicating that the pruning process does not cause a large loss to the accuracy of the target sub-model.

[0043] As an example, the latency and accuracy after slicing and pruning of each model are shown in Table 4 below:

[0044] Table 4 In step S4, after obtaining the model splitting point of the target model, the method further includes: S5. Input the target model and the model splitting point into the large language model, so that the large language model can generate a task plan that matches the user demand information according to the target model and the model splitting point.

[0045] It should be noted that the large language model has semantic understanding and task planning capabilities, can automatically analyze the user demand tasks and perform task planning processing, and generate a task plan that matches the user demand information. In this embodiment, by introducing a large language model for user demand task analysis and task planning, and automatically generating a task plan adapted to the user demand information, manual intervention can be greatly reduced, the computing power burden of the edge device can be reduced, the system adaptability can be improved, and it is applicable to complex and diverse application scenarios.

[0046] As an example, when the large language model uses Tongyi Qianwen large model, the generated task plan is as follows: "Considering that your requirement is to perform an image classification task, and the required accuracy is not less than 80%, while also considering factors such as cloud collaborative inference, the latency of edge computing, and the model size, I will recommend the following solution: 1. Model selection: Based on the accuracy and latency during local inference, ResNet-10 CIFAR10.pth is a good choice because it has a relatively high accuracy (0.8659) and moderate loading and inference latency. Although AlexNet also has a high accuracy, its inference latency is slightly higher.

[0047] 2. Model splitting point: For ResNet-10, the optimal splitting point is after 'layer4'. This divides the model into a part for local processing (the first 4 layers) and a part for cloud processing (the subsequent layers).

[0048] 3. Sparsity rate setting: The sparsity rate combination optimized by the genetic algorithm for ResNet-10 is [0.1732..., 0.6316..., 0.8543...], which will reduce the computational amount of the local part and improve the edge computing efficiency.

[0049] 4. Collaborative inference: Run the first 4 layers (including the sparsified part) of ResNet-10 on an edge device (such as a Raspberry Pi or Jetson TK1), and then send the processed features to the cloud (such as NVIDIA Tesla P100), where subsequent deep learning calculations are performed. This can utilize the powerful computing power of P100 while maintaining a relatively high recognition accuracy.

[0050] 5. Performance Evaluation: After pruning, the accuracy of ResNet-10 slightly decreases (0.859), but still meets the requirement of 80%. The local inference latency increases, but the overall latency may decrease due to reduced cloud interaction, depending on the network transmission time.

[0051] Please note that during actual deployment, you may need to further adjust the splitting point and sparsity rate according to network conditions and real-time requirements to find the best balance. As an example, when the large language model adopts the Purple East Taichu large model, the generated task plan is as follows: "According to your needs, it is recommended to select the ResNet model for edge collaborative inference. The accuracy of this model on the CIFAR-10 dataset reaches 0.859, meeting the accuracy requirement of not less than 80%.

[0052] The model splitting point is selected at 'layer4', and the part before this is run locally. For the sparsity rate combination, through genetic algorithm optimization, multiple groups of parameters are obtained. The specific values are relatively complex, but can be generally summarized as pruning each layer of the model to different degrees.

[0053] After splitting and pruning, the performance of the ResNet model has significant advantages in the following aspects: 1. The accuracy remains relatively high at 0.859, meeting the task requirements.

[0054] 2. The inference latency is reduced. The specific value is not given, but pruning usually can reduce model parameters and improve the inference speed.

[0055] A brief description of the edge collaborative solution is as follows: - Model Selection: ResNet - Splitting Point: layer4 - Sparsity Rate Combination: Prune each layer of the model to different degrees, and the specific parameters need to be determined according to the optimization results of the genetic algorithm.

[0056] - Model Performance: The accuracy after pruning is 0.859, and the inference latency is reduced. In addition, in this embodiment, after the large language model generates the task plan, the large language model can also be evaluated through preset model evaluation metrics. Specifically, the model evaluation metrics include, for example, model size, inference speed, model perplexity, accuracy, and F1-Score, etc. Since answering user needs and giving reasonable solutions are answers to open-ended questions without a fixed generation template, in this embodiment, three types of metrics, namely BLEU metric, ROUGE score, and BERTScore metric, are selected to evaluate the answer level of the large language model.

[0057] Among them, the BLEU metric can be used to evaluate the similarity between the generated answer and the standard answer. Its basic principle is to quantify the quality of the generated text by comparing the candidate text generated by the large language model with one or more groups of standard answer reference texts.

[0058] The ROUGE score is a metric commonly used to evaluate the quality of natural language generation tasks such as text summarization and machine translation. It evaluates the quality of the text generated by the large language model by comparing the overlap between the generated text and the reference text. The ROUGE score includes ROUGE-1, ROUGE-2, and ROUGE-L. Among them, ROUGE-1 represents the F-measure score based on the out-of-vocabulary word level; ROUGE-2 represents the F-measure score based on the bigram level; ROUGE-L represents the F-measure score based on the longest common subsequence. The higher these three scores are, the higher the similarity between the generated text and the standard answer, and the better the quality.

[0059] The BERTScore metric includes three main metrics: F1 score, recall, and precision. Among them, recall is used to measure the proportion of information from the reference text contained in the generated text; precision evaluates how much of the information identified as relevant in the generated text is truly relevant, that is, the accuracy of the information in the generated text; the F1 score is the harmonic mean of precision and recall, aiming to comprehensively consider precision and recall.

[0060] This embodiment can be applied to diversified application scenarios and can improve the model inference efficiency. Specifically, during the implementation of this embodiment, after receiving user requirement information, a preset large language model is called, and based on the large language model, the user requirement information is analyzed and processed for tasks to obtain a task category and task metrics; subsequently, a data set and multiple candidate models matching the task category are obtained, and multiple candidate models are respectively pre-trained based on the data set to obtain multiple pre-trained models; then, a target model matching the task metrics is extracted from the multiple pre-trained models based on the large language model; finally, a model splitting point of the target model is obtained, and the target model is split into a first sub-model and a second sub-model by using the model splitting point, and then the first sub-model is deployed on the edge device, and the second sub-model is deployed on the cloud server corresponding to the edge device, so that the edge device and the cloud server perform end-cloud collaborative inference on the target model. By combining the large language model with edge intelligence technology, this embodiment can accurately and quickly identify user requirement information, and can further obtain a target model that meets user requirements. In this embodiment, by splitting the target model and adopting the method of end-cloud collaborative inference to perform end-cloud collaborative inference on the target model, the inference efficiency of the model can be improved, so that the edge intelligence technology can respond more quickly in an environment sensitive to latency or limited in bandwidth, and thus can be efficiently applied to diversified application scenarios.

[0061] Embodiment 2: This embodiment discloses a large language model-based end-cloud collaborative inference method for implementing the large language model-based end-cloud collaborative inference method in Embodiment 1; as Figure 2 shown, the end-cloud collaborative inference method includes: A task analysis module, configured to, after receiving user requirement information, call a preset large language model, and perform task analysis and processing on the user requirement information based on the large language model to obtain a task category and task metrics; A model training module, communicatively connected to the task analysis module, configured to obtain a data set and multiple candidate models matching the task category, and respectively pre-train the multiple candidate models based on the data set to obtain multiple pre-trained models; A target model extraction module, communicatively connected to the model training module, configured to extract a target model matching the task metrics from the multiple pre-trained models based on the large language model; A collaborative inference module, communicatively connected to the target model extraction module, is configured to obtain the model segmentation points of the target model, and use the model segmentation points to split the target model into a first sub-model and a second sub-model, and then deploy the first sub-model on an edge device and the second sub-model on a cloud server corresponding to the edge device, so that the edge device and the cloud server perform end-cloud collaborative inference on the target model.

[0062] It should be noted that for the working process, working details and technical effects of the end-cloud collaborative inference method based on a large language model provided in this Embodiment 2, reference can be made to Embodiment 1 and will not be elaborated here.

[0063] Embodiment 3: Based on Embodiment 1 or 2, this embodiment discloses an electronic device, which may be a smart phone, a tablet computer, a laptop computer, a desktop computer, etc. The electronic device may be referred to as a user terminal, a portable terminal, a desktop terminal, etc., such as Figure 3 As shown, the electronic device includes: A memory for storing computer program instructions; and, A processor for executing the computer program instructions to complete the operations of a method for end-cloud collaborative inference based on a large language model as described in any one of Embodiment 1.

[0064] Specifically, the processor 301 may include one or more processing cores, such as a 4-core processor or an 8-core processor. The processor 301 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), or PLA (Programmable Logic Array). The processor 301 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 301 may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen.

[0065] The memory 302 may include one or more computer-readable storage media, which may be non-transitory. The memory 302 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices. In some embodiments, the non-transitory computer-readable storage media in the memory 302 is used to store at least one instruction for being executed by the processor 301 to implement the end-cloud collaborative inference method based on a large language model provided in Embodiment 1 of this application.

[0066] In some embodiments, the terminal may further optionally include: a communication interface 303 and at least one peripheral device. The processor 301, the memory 302, and the communication interface 303 may be connected through a bus or signal lines. Each peripheral device may be connected to the communication interface 303 through a bus, signal lines, or a circuit board. Specifically, the peripheral device includes at least one of a radio frequency circuit 304, a display screen 305, and a power supply 306.

[0067] The communication interface 303 can be used to connect at least one peripheral device related to I / O (Input / Output) to the processor 301 and the memory 302. In some embodiments, the processor 301, the memory 302, and the communication interface 303 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 301, the memory 302, and the communication interface 303 can be implemented on a separate chip or circuit board, and this embodiment does not limit this.

[0068] The radio frequency circuit 304 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency circuit 304 communicates with a communication network and other communication devices through electromagnetic signals.

[0069] The display screen 305 is used to display a UI (User Interface). The UI may include any combination of graphics, text, icons, and videos.

[0070] The power supply 306 is used to supply power to each component in the electronic device.

[0071] Embodiment 4: Based on any one of Embodiments 1 to 3, this embodiment discloses a computer program product, including a computer program or instruction, and the computer program or the instruction, when executed by a computer, implements a method for end-cloud collaborative inference based on a large language model as described in any one of Embodiments 1. Wherein, the computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices.

[0072] Obviously, those skilled in the art should understand that the above-mentioned modules or steps of the present invention can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. Optionally, they can be implemented by program codes executable by the computing device. Thus, they can be stored in a storage device and executed by the computing device, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module for implementation. In this way, the present invention is not limited to any specific combination of hardware and software.

[0073] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. An edge-cloud collaborative inference method based on large language models, characterized in that, Executed by an edge device, the method includes: After receiving user requirement information, calling a preset large language model, and performing task analysis and processing on the user requirement information based on the large language model to obtain a task category and task metrics; Obtaining a data set and multiple candidate models that match the task category, and respectively pre-training the multiple candidate models based on the data set to obtain multiple pre-trained models; Extracting a target model that matches the task metrics from the multiple pre-trained models based on the large language model; Obtaining a model splitting point of the target model, and using the model splitting point to split the target model into a first sub-model and a second sub-model, and then deploying the first sub-model on the edge device and deploying the second sub-model on a cloud server corresponding to the edge device, so that the edge device and the cloud server perform end-cloud collaborative inference on the target model.

2. The end-cloud collaborative inference method based on a large language model according to claim 1, wherein The task category includes an image classification task, a crop pest and disease identification task, and a weed identification task. The data set matching the image classification task is the CIFAR10 data set, the data set matching the crop pest and disease identification task is the AgriculturalDisease data set, the data set matching the weed identification task is the Kaggle archive data set, and the multiple candidate models matching the image classification task, the crop pest and disease identification task, and the weed identification task all include the AlexNet model, the ResNet50 model, and the VGG16 model.

3. A method for end-cloud collaborative inference based on a large language model according to claim 1, characterized in that Obtaining the model splitting point of the target model includes: Determining a candidate splitting point in the target model, and using the current candidate splitting point to split the target model into a first candidate sub-model and a second candidate sub-model; Deploying the first candidate sub-model on the edge device and deploying the second candidate sub-model on a cloud server corresponding to the edge device, so that the edge device and the cloud server perform end-cloud collaborative inference on the target model; Calculating the computing delay of the edge device, the data transmission delay from the edge device to the cloud server, and the computing delay of the cloud server through timestamps, and combining the computing delay of the edge device, the data transmission delay from the edge device to the cloud server, and the computing delay of the cloud server to obtain the total end-cloud collaborative inference delay of the current candidate splitting point; Traversing each candidate splitting point of the target model, obtaining the candidate splitting point with the minimum total end-cloud collaborative inference delay among each candidate splitting point, and using the candidate splitting point with the minimum total end-cloud collaborative inference delay as the model splitting point of the target model, and using the total end-cloud collaborative inference delay of the model splitting point as the optimal total delay.

4. A method for end-cloud collaborative inference based on a large language model according to claim 1, characterized in that, After using the model splitting point to split the target model into a first sub-model and a second sub-model, the method further includes: Respectively take the first sub-model and the second sub-model as the target sub-model, obtain the optimal pruning sparsity rate combination of the target sub-model, and perform pruning processing on the target sub-model based on the optimal pruning sparsity rate combination to obtain the pruned target sub-model, so as to deploy the pruned target sub-model corresponding to the first sub-model on the edge device and deploy the pruned target sub-model corresponding to the second sub-model on the cloud server.

5. The end-cloud collaborative inference method based on a large language model according to claim 4, wherein Use the genetic algorithm to obtain the optimal pruning sparsity rate combination of the target sub-model.

6. The end-cloud collaborative inference method based on a large language model according to claim 4, characterized in that, When performing pruning processing on the target sub-model based on the optimal pruning sparsity rate combination, call the torch.nn.utils.prune module in PyTorch to execute.

7. A method for end-cloud collaborative inference based on a large language model according to claim 1, characterized in that, After obtaining the model splitting point of the target model, the method further includes: Input the target model and the model splitting point into the large language model, so that the large language model generates a task plan matching the user requirement information according to the target model and the model splitting point.

8. An end-cloud collaborative inference method based on a large language model, characterized in that, Include: A task analysis module, configured to, after receiving user requirement information, call a preset large language model, and perform task analysis processing on the user requirement information based on the large language model to obtain a task category and task metrics; A model training module, communicatively connected to the task analysis module, configured to obtain a data set and a plurality of candidate models matching the task category, and perform pre-training on the plurality of candidate models respectively based on the data set to obtain a plurality of pre-trained models; A target model extraction module, communicatively connected to the model training module, configured to extract a target model matching the task metrics from the plurality of pre-trained models based on the large language model; A collaborative inference module, communicatively connected to the target model extraction module, configured to obtain the model splitting point of the target model, and use the model splitting point to split the target model into a first sub-model and a second sub-model, and then deploy the first sub-model on the edge device and deploy the second sub-model on the cloud server corresponding to the edge device, so that the edge device and the cloud server perform end-cloud collaborative inference on the target model.

9. An electronic device, characterized in that, Include: A memory, configured to store computer program instructions; And, A processor, configured to execute the computer program instructions to complete the operations of a method for end-cloud collaborative inference based on a large language model according to any one of claims 1 to 7.

10. A computer program product, comprising a computer program or instructions, characterized in that, The computer program or the instructions, when executed by a computer, implement a method for end-cloud collaborative inference based on a large language model according to any one of claims 1 to 7.