Base model-oriented scene self-adaption edge cloud collaborative inference system and method

Through the scene-adaptive edge-cloud collaborative inference system, superclass routing and model distillation are used to fine-tune the customized model on the edge device, which solves the problems of insufficient computing power and transmission delay of the basic model on the edge device, realizes efficient and accurate inference and protects data privacy.

CN119721238BActive Publication Date: 2025-10-10BEIJING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411765028.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-04
Publication Date
2025-10-10
Estimated Expiration
2044-12-04

AI Technical Summary

Technical Problem

Existing technologies make it difficult to effectively deploy basic models on edge devices, resulting in insufficient computing power and transmission delays. Existing methods may cause loss of inference accuracy or be unsuitable for the Transformer architecture.

Method used

A scenario-adaptive edge-cloud collaborative reasoning system is adopted. Hybrid expert models are extracted through superclass routing and model distillation. Custom models are fine-tuned on edge devices by combining model distillation and contrastive learning. Data transmission is optimized by uploading deciders and data block selectors, and collaborative reasoning strategies are dynamically adjusted.

Benefits of technology

It achieves efficient and accurate basic model inference on edge devices, reduces transmission delays and protects data privacy, and is suitable for a variety of application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119721238B_ABST
    Figure CN119721238B_ABST
Patent Text Reader

Abstract

The present application relates to a kind of scene self-adapting edge cloud collaborative reasoning system and method for basic model, including scene self-adapting edge side model customization and adaptive reasoning two components;Scene self-adapting edge side model customization component extracts mixed expert model from the basic model deployed in cloud side by the method of superclass routing and model distillation, then, using a small amount of unlabeled original data on edge device, select appropriate expert module to compress and fine-tune edge side customization model.Adaptive reasoning component cooperates with edge side customization model and the basic model deployed in cloud side to process reasoning task.Edge side customization model reasoning result is uploaded to decision module, calculate confidence score, and decide whether to need to upload data block selector, if need to upload, upload local and important data block to the basic model deployed in cloud side, reduce transmission overhead and protect data privacy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to artificial intelligence and edge computing technologies, and in particular to a scene-adaptive edge-cloud collaborative reasoning system and method for a basic model. Background Art

[0002] As an emerging artificial intelligence technology, the basic model has received widespread attention. The basic model can effectively process multiple data modalities, such as text, images, video, audio, etc. The basic model obtains a rich knowledge reserve through pre-training on large-scale data sets, supports a variety of downstream tasks, and greatly reduces the data requirements and training costs for specific tasks.

[0003] Based on the above characteristics, the basic model has some typical applications on edge devices; for example:

[0004] ① Intelligent monitoring, using smart cameras and other sensors to identify unusual activities or sounds;

[0005] ② Intelligent assistants use multimodal information such as data records, text, and images on mobile phones to analyze user profiles and provide intelligent services;

[0006] ③Virtual / augmented reality, allowing users in different regions to share and modify virtual / augmented reality content in real time.

[0007] Although the basic model has excellent performance, with the continuous improvement of people's production and living standards, the demand for various intelligent technologies is also gradually increasing, which makes the deployment solutions on existing edge devices face many challenges.

[0008] Due to the limited computing power and storage resources of edge devices, they cannot meet the memory overhead of the basic model, making it difficult for them to fully realize their potential in existing practical applications.

[0009] At present, most research uses technologies such as data compression, model lightweighting, and model partitioning to deploy basic models on edge devices; some work directly compresses the original data and uploads it to the cloud server for processing. Although this reduces transmission costs, simple data compression will lead to the loss of important information, greatly affecting the inference accuracy.

[0010] There are also some works that use model lightweight techniques (such as pruning, quantization, and knowledge distillation) to compress the basic model and perform inference on edge devices, but this approach will cause loss of accuracy and generalization of the basic model.

[0011] Some other works divide the model in a cross-layer or hierarchical manner and perform edge-cloud collaborative inference, but are more suitable for models with convolutional neural network (CNN) architecture; the base model is mostly of Transformer architecture, which cannot be divided due to its self-attention mechanism, making the cross-layer method unsuitable.

[0012] In addition, there is a lack of dimension reduction operation inside the base model, and the input and output shapes remain consistent, resulting in poor optimization effect of the hierarchical method. SUMMARY

[0013] To solve the above technical problems, the application provides a scene-adaptive edge-cloud collaborative inference system and method for a base model, which utilizes a single powerful base model, customizes a lightweight model based on different scene data types and task requirements, performs edge-cloud collaborative inference, is suitable for various applications such as intelligent monitoring, intelligent assistants, virtual / augmented reality, effectively solves the problems of high computational overhead and transmission delay caused by insufficient computing power of edge devices, and uploads local data blocks of original data to protect data privacy.

[0014] To achieve the above purpose, the technical solution adopted by the application is as follows: first, a scene-adaptive edge-cloud collaborative inference system for a base model is provided, which includes a scene-adaptive edge-side model customization component and an adaptive inference component.

[0015] The scene-adaptive edge-side model customization component includes:

[0016] A hybrid expert model extraction module is used to extract a hybrid expert model from the base model deployed on the cloud side by the method of superclass routing and model distillation, and to divide the knowledge of different expert modules in the hybrid expert model in the form of superclasses, each expert module being responsible for a superclass derived from the categories of a public data set, the hybrid expert model serving as a pre-training model for the edge-side customization model;

[0017] An edge-side customization model fine-tuning module is used to fine-tune the edge-side customization model after the extraction of the hybrid expert model, select an expert module with the highest importance from each hybrid expert layer according to the data type and task requirement of the edge device, compress the hybrid expert model, and fine-tune the edge-side customization model using the unlabeled original data of the edge device, the fine-tuning method combining model distillation and contrastive learning, using the base model deployed on the cloud side as the teacher model for model distillation, and selecting text features required for contrastive learning from the text feature library T on the cloud side;

[0018] The adaptive inference component is used to cooperate with the edge-side customization model and the base model deployed on the cloud side to process inference tasks, and includes:

[0019] Upload decider: This is used to evaluate the confidence score of the edge-customized model inference result. If the confidence score is higher than the set confidence threshold, the edge-customized model inference result is adopted. Otherwise, the data sample is input into the upload data block selector.

[0020] Upload data block selector: used to evaluate the importance of data blocks of input data samples, select and upload local and important data sample blocks to the basic model deployed on the cloud side;

[0021] Collaborative Strategy Adjuster: This is used to estimate real-time network bandwidth to update transmission latency, query the threshold search table, calculate the predicted end-to-end inference latency, and dynamically adjust the confidence threshold of the upload decider and the number of data sample blocks of the upload data block selector based on different data types and task requirements to balance the latency and accuracy of the inference task.

[0022] Preferably, the hybrid expert module includes several backbone network layers, a superclass routing, several hybrid expert layers and a feature mapping layer, and the hybrid expert layer includes several expert modules responsible for different superclasses.

[0023] Preferably, the importance calculation process of the edge customization model fine-tuning module is as follows:

[0024] The importance is evaluated by calculating the probability of each expert module being selected using the data samples of the edge device. The data samples are sent to the expert module e by the superclass routing. j The importance is expressed as I j , the calculation formula is:

[0025] P i =softmax(W r ),

[0026]

[0027] Among them, P i ={P i,1 ,P i,2 ,...,P i,N , represents the probability that data sample i is assigned to N expert modules, W r is the N-dimensional probability distribution vector output by the superclass routing, P i,j The data sample i is sent to the expert module e by the superclass routing j The probability of , i is the data sample number, j is the expert module number.

[0028] Preferably, the fine-tuning process of the edge customized model fine-tuning module is as follows:

[0029] Using unlabeled raw data from edge devices, we fine-tune a customized edge model. This fine-tuning approach combines model distillation and contrastive learning. The base model deployed on the cloud side serves as the teacher model for model distillation. Furthermore, the text feature library on the cloud side is used to select the text features required for contrastive learning.

[0030] First, the data samples are input into the basic model deployed on the cloud side and the customized model on the edge side respectively, and the sample features extracted by the basic model are calculated. and sample features extracted by the side-customized model The mean squared error loss (MSE loss) is used as the loss function L for model distillation. dist , the calculation formula is:

[0031]

[0032] At the same time, in order to adapt to the unlabeled data environment of edge devices, the loss function L of contrastive learning is set con , query the sample features extracted from the text feature library and the basic model Most similar text features Closely define the sample features extracted by the side-customized model and text features The distance, the loss function L of contrastive learning con The calculation formula is:

[0033] L con =L I,T +L T,I ,

[0034]

[0035] Where sim(·) is the cosine similarity of the features, and τ is the temperature coefficient;

[0036] The loss function L for side-by-side custom model fine-tuning is the loss function L for model distillation. dist And the loss function L of contrastive learning con Combination of:

[0037]

[0038] in, is the weight of each data sample.

[0039] Preferably, the confidence score calculation process of the upload decision maker module is as follows:

[0040] The data sample is input into the edge-side customization model, and the sample feature is output, the cosine similarity between the sample feature and each text feature in the text feature library is calculated, the entropy value is calculated according to the cosine similarity distribution, and the entropy value is normalized to obtain the confidence score of the input data sample.

[0041] Preferably, the upload data block selector module evaluates the data block importance of the input data sample, and the process is as follows:

[0042] The cosine similarity between the class token of the last layer of the edge-side customization model and the data token is taken as the category attention score A cls , which represents the importance of the data block of the input data sample, and the calculation formula is:

[0043]

[0044] Wherein, q cls is the class token, K T is the key matrix, that is, the set of all input data tokens, and d is the dimension of the class token and the data token.

[0045] Preferably, the end-to-end inference delay of the collaborative strategy adjuster is calculated as follows:

[0046] Periodically collect raw data samples from edge devices, calculate the proportion r(thre) of the edge-side customization model and the cloud-side basic model processing data samples, the overall accuracy acc(thre), the edge-side processing delay l e , the transmission delay l tr (thre), the cloud-side processing delay l c , and the end-to-end inference delay , and store them in the threshold search table, and the calculation formula is as follows:

[0047]

[0048] The application also provides a scene-adaptive edge-cloud collaborative inference method for a basic model, which is performed in the scene-adaptive edge-cloud collaborative inference system for the basic model and includes the following steps:

[0049] S1 extracts a mixed expert model from the cloud-side deployed basic model by the method of superclass routing and model distillation, divides the knowledge of a specific field for different expert modules in the mixed expert model in the form of a superclass, each expert module is responsible for a superclass derived from the category of a public data set, and the mixed expert model serves as a pre-training model of the edge-side customization model.

[0050] After S2 extracts the hybrid expert model, it selects the most important expert module from each hybrid expert layer based on the edge device's data type and task requirements. It then compresses the hybrid expert model and fine-tunes it using the edge device's unlabeled raw data to produce an edge-customized model. This fine-tuning approach combines model distillation and contrastive learning, using the base model deployed on the cloud as the teacher model for model distillation. Furthermore, the cloud-side text feature library T is used to select the text features required for contrastive learning.

[0051] The side-customized model in S3 divides the input data sample into multiple data blocks and calculates the preliminary inference results; then, the preliminary inference results are passed to the upload decision maker;

[0052] S4 evaluates the confidence score of the edge-customized model inference result. If the confidence score is higher than the set confidence threshold, the edge-customized model inference result is adopted. Otherwise, the data sample is input into the upload data block selector.

[0053] S5 evaluates the importance of data blocks of input data samples, selects and uploads local and important data sample blocks to the basic model deployed on the cloud side, and uses the basic model deployed on the cloud side to infer the results;

[0054] S6 estimates the real-time network bandwidth to update the transmission delay, queries the threshold search table, calculates the predicted end-to-end inference delay, and dynamically adjusts the confidence threshold of the upload decision maker and the number of data sample data blocks of the upload data block selector according to different data types and task requirements to balance the latency and accuracy of the inference task.

[0055] The above technical solution has the following advantages or beneficial effects:

[0056] This paper designs a scenario-adaptive edge-cloud collaborative inference system and method for basic models, which enables efficient and accurate edge-cloud collaborative inference of basic models on edge devices in different scenarios. It has the following innovations:

[0057] ① This invention solves the deployment problem of the basic model on edge devices and achieves efficient and accurate reasoning.

[0058] ② The present invention realizes the use of a single basic model to adaptively customize the side lightweight model based on the data type and task requirements of different scenarios.

[0059] ③ The present invention realizes the importance assessment of the original data on the side, and only uploads local and important data blocks, reducing transmission delay and protecting data privacy.

[0060] ④ The present invention realizes the dynamic adjustment of collaborative reasoning strategy to balance the requirements of reasoning latency and accuracy.

[0061] The above summary is intended to illustrate only and is not intended to be limiting in any way. Further aspects, implementations, and features of the present application will become apparent from the following detailed description, taken in conjunction with the accompanying drawings and the description of illustrative aspects, implementations and features described above. BRIEF DESCRIPTION OF DRAWINGS

[0062] In the drawings, like numerals refer to like elements throughout the various drawings. The drawings are not necessarily to scale, the emphasis instead being placed on illustrating principles of the application. It should be understood that the drawings are merely illustrative of certain embodiments of the application and that they, therefore, do not limit the scope of the application.

[0063] Figure 1 A scene-adaptive edge cloud collaborative reasoning system for a basic model provided by the present application is shown in the overall architecture diagram of the system;

[0064] Figure 2 A hybrid expert model architecture provided by the present application is shown in the architecture diagram of the hybrid expert model;

[0065] Figure 3 An edge side customized model fine-tuning flowchart provided by the present application is shown in the flowchart of the edge side customized model fine-tuning; DETAILED DESCRIPTION

[0066] In the following, only certain example embodiments are simply described. As those skilled in the art can recognize, the described embodiments can be modified in various different ways without departing from the spirit or scope of the present application. Therefore, the drawings and the description are considered to be essentially illustrative rather than limiting.

[0067] The present application first provides a scene-adaptive edge cloud collaborative reasoning system for a basic model, the overall architecture of the system is shown in Figure 1 The system includes a scene-adaptive edge side model customization component and an adaptive reasoning component.

[0068] The scene-adaptive edge side model customization component realizes the output of edge side customized models for edge devices of different data types and task requirements by using a single powerful basic model and a small amount of unlabeled edge device raw data, including a hybrid expert model extraction module and an edge side customized model fine-tuning module.

[0069] The existing method can obtain a lightweight edge side customized model through model compression and fine-tuning technology, but this method has high training cost and requires a large amount of training data. Therefore, the present application extracts a hybrid expert model from the basic model deployed on the cloud side, which is used as a pre-training model to compress and fine-tune a lightweight edge side customized model.

[0070] The architecture of the hybrid expert model extraction module is shown in Figure 2As shown, the system comprises several backbone network layers, a superclass routing, several hybrid expert layers, and a feature mapping layer. The hybrid expert layer includes several expert modules, which are activated according to different data types and task requirements. Each expert module has expertise in a specific field. The backbone network layer receives data samples as input and extracts low-level sample features. The superclass routing selects the most appropriate expert module based on these low-level sample features. The expert module extracts high-level sample features. The feature mapping layer maps the high-level sample features to the same number of dimensions as the base model, which are then used as the final output sample features.

[0071] This module is used for training through superclass routing and model distillation methods. Hybrid expert model The categories of public datasets are grouped into superclasses, the backbone network layers are frozen, and supervised learning is used to train superclass routing and expert modules based on the base model. This ensures that each expert module handles a specific superclass while preserving the performance of the base model to the greatest extent possible. For datasets without superclass classification, such as the ImageNet dataset, the base model is used to calculate the confusion matrix on the dataset's validation set, from which a confusion graph is constructed. A graph clustering algorithm is then applied to determine the superclass groupings of the dataset.

[0072] The present invention extracts a hybrid expert model from a large-scale public dataset, and selects the image dataset ImageNet (1000 categories, 10 superclasses) and the audio dataset Audioset (1000 categories, 12 superclasses) to implement two types of tasks: image and audio recognition.

[0073] The present invention selects CLIP (Contrastive Language-Image Pre-Training), a basic model multimodal pre-training neural network, and ImageBind, a method for learning joint feature embedding, to extract a hybrid expert model. Among them, CLIP has an image encoder and a text encoder, and can perform image recognition tasks; ImageBind has an image encoder, an audio encoder, and a text encoder, and can perform image and audio recognition tasks. Specifically, taking the image recognition task as an example, the basic model uses a text encoder to calculate text features of a series of categories, and uses an image encoder to calculate image features of image data. Finally, it calculates and selects the text feature with the highest cosine similarity with the image feature, and uses the category corresponding to the text feature as the recognized image category.

[0074] The extracted mixture of experts model consists of 12 Transformer layers, with the last three layers being mixture of experts. The hidden embedding dimension within the Transformer layer is 192, the number of heads in the multi-head self-attention mechanism is 3, and the dimension of the multilayer perceptron is four times the hidden embedding dimension. For the image recognition task, each mixture of experts layer has 10 expert modules, the input image size is 224×224×3, and the data block size is 16×16. For the audio recognition task, each mixture of experts layer has 12 expert modules, the input data size is 128×204×3, and the data block size is 16×16.

[0075] The edge-customized model fine-tuning module is used to complete the extraction of the hybrid expert model, select the most important expert module from each hybrid expert layer according to the data type and task requirements of the edge device, compress the hybrid expert model, and use the unlabeled raw data of the edge device to fine-tune the edge-customized model. At the same time, the size of the edge-customized model can be adjusted according to the scenario and computing power of the edge device.

[0076] The importance calculation process is as follows:

[0077] The importance is evaluated by calculating the probability of each expert module being selected using the data samples of the edge device. The data samples are sent to the expert module e by the superclass routing. j The importance is expressed as I j , the calculation formula is:

[0078] P i =softmax(W r ),

[0079]

[0080] Among them, P i ={P i,1 ,P i,2 ,...,P i,N}, represents the probability that data sample i is assigned to N expert modules, W r is the N-dimensional probability distribution vector output by the superclass routing, P i,j The data sample i is sent to the expert module e by the superclass routing j The probability of , i is the data sample number, j is the expert module number.

[0081] This invention collects multiple categories from multiple edge devices and public datasets. The base model's text encoder calculates text features for these categories to form a text feature library T. It also supports the flexible addition of new categories. During inference, the edge-customized model extracts image or audio features and performs image or audio recognition by selecting the text features from the text feature library T that have the highest cosine similarity with the image or audio features.

[0082] like Figure 3 As shown, the present invention uses the unlabeled raw data of the edge device to fine-tune the edge-customized model. This fine-tuning method combines model distillation and contrastive learning, and uses the basic model deployed on the cloud side as the teacher model for model distillation. At the same time, the text feature library T on the cloud side is used to select the text features required for contrastive learning.

[0083] The fine-tuning process is as follows: First, the data samples are input into the basic model and the side customized model respectively, and the sample features extracted by the basic model are calculated. and sample features extracted by the side-customized model The mean squared error loss (MSE) is used as the loss function L for model distillation. dist , the calculation formula is:

[0084]

[0085] At the same time, in order to effectively utilize the unlabeled raw data of edge devices, the loss function L of contrastive learning is set con , query the sample features extracted from the text feature library T and the basic model Most similar text features Closely define the sample features extracted by the side-customized model and text features The distance, the loss function L of contrastive learning con The calculation formula is:

[0086] L con =L I,T +L T,I ,

[0087]

[0088] Where sim(·) is the cosine similarity of the features, and τ is the temperature coefficient;

[0089] The loss function L for side-by-side custom model fine-tuning is the loss function L for model distillation. dist And the loss function L of contrastive learning con Combination of:

[0090]

[0091] in, is the weight of each data sample.

[0092] The adaptive inference component dynamically balances the latency and accuracy of inference tasks by flexibly adjusting the collaborative inference strategy between the edge-side customized model and the basic model deployed on the cloud side. It mainly consists of three modules: upload decider, upload data block selector, and collaborative strategy adjuster:

[0093] Existing methods use the output of traditional models for evaluation, which is not applicable to the base models deployed on the cloud side. This is because traditional models take data samples as input and output a one-hot encoding, which uses an N-bit feature vector to encode N states, where each state represents the confidence score for each category in the recognition task. However, the base models deployed on the cloud side use image or audio encoders to calculate data samples as high-dimensional sample features, making it impossible to directly calculate the confidence score of the inference result from them.

[0094] The upload decider evaluates the confidence score of the edge-customized model's inference results. It first inputs the data sample into the edge-customized model, calculates the image or audio features, and then calculates the cosine similarity between them and each text feature in the text feature library T. It then calculates the entropy value based on the cosine similarity distribution and normalizes the entropy value to obtain the confidence score of the input data sample. If the confidence score exceeds the set confidence threshold, the edge-customized model's inference result is used. Otherwise, the data sample is input into the upload data block selector.

[0095] Upload Data Block Selector: This function evaluates the importance of data blocks of input data samples and selects and uploads data blocks of local and important data samples to the base model deployed on the cloud side, reducing transmission costs and protecting data privacy on edge devices. Specifically, when performing inference, the edge-based customized model divides the input data samples into multiple data blocks. For example, for image recognition tasks, the edge-based customized model divides a 224×224 image data sample into multiple 16×16 data blocks (patches). During inference, each data block is mapped to high-dimensional data tokens. These tokens, along with a specially introduced class token, are input into the model for computation. The class token aggregates the overall information of the data sample and outputs the final sample features. By calculating the cosine similarity between each data block and the class token, the similarity between the class token and the data block can be quantified, thereby determining which data blocks contribute most to the final result.

[0096] The process of evaluating the importance of a data block of an input data sample is as follows:

[0097] The cosine similarity between the class token and the data tokens of the last layer of the side custom model is used as the category attention score A cls , represents the importance of the data block of the input data sample, and the calculation formula is:

[0098]

[0099] Among them, q cls Mark the class token, K T is the key matrix, that is, the set of all input data labels, and d is the dimension of category labels and data labels.

[0100] Collaborative Strategy Adjuster: To adapt to different edge device scenarios and network transmission conditions, this paper designs a collaborative strategy adjuster. During inference execution, this paper estimates the real-time network bandwidth to update the transmission latency, queries the threshold search table, and calculates the predicted end-to-end inference latency. Based on the task requirements of different edge devices, including transmission latency constraints and accuracy constraints, this paper dynamically adjusts the confidence threshold of the upload decider and the number of data sample blocks in the upload data block selector, balancing transmission latency and inference accuracy to achieve efficient and accurate edge-cloud collaborative inference.

[0101] The end-to-end inference latency is calculated as follows:

[0102] Regularly collect raw data samples from edge devices, calculate the ratio r(thre) of data samples processed by the edge-side custom model and the basic model on the cloud side, the overall accuracy acc(thre), and the edge-side processing delay l e , transmission delay l tr (thre), cloud-side processing delay l c , calculate end-to-end inference latency And store it in the threshold search table, the calculation formula is as follows:

[0103]

[0104] The present invention also provides a scene-adaptive edge-cloud collaborative reasoning method for a basic model, which is performed in the scene-adaptive edge-cloud collaborative reasoning system for a basic model, and includes the following steps:

[0105] S1 extracts a hybrid expert model from the base model deployed on the cloud side through superclass routing and model distillation. It then divides domain-specific knowledge into superclasses for different expert modules in the hybrid expert model. The superclasses for each expert module are derived from the categories of the public dataset. The hybrid expert model serves as a pre-trained model for the side-customized model.

[0106] After S2 extracts the hybrid expert model, it selects the most important expert module from each hybrid expert layer based on the edge device's data type and task requirements. It then compresses the hybrid expert model and fine-tunes it using the edge device's unlabeled raw data to produce an edge-customized model. This fine-tuning approach combines model distillation and contrastive learning, using the base model deployed on the cloud as the teacher model for model distillation. Furthermore, the cloud-side text feature library is used to select the text features required for contrastive learning.

[0107] The side-customized model in S3 divides the input data sample into multiple data blocks and calculates the preliminary inference results; then, the preliminary inference results are passed to the upload decision maker;

[0108] S4 evaluates the confidence score of the edge-customized model inference result. If the confidence score is higher than the set confidence threshold, the edge-customized model inference result is adopted. Otherwise, the data sample is input into the upload data block selector.

[0109] S5 evaluates the importance of data blocks of input data samples, selects and uploads local and important data sample blocks to the basic model deployed on the cloud side, and uses the basic model deployed on the cloud side to infer the results;

[0110] S6 estimates the real-time network bandwidth to update the transmission delay, queries the threshold search table, calculates the predicted end-to-end inference delay, and dynamically adjusts the confidence threshold of the upload decision maker and the number of data sample data blocks of the upload data block selector according to different data types and task requirements to balance the latency and accuracy of the inference task.

[0111] Experimental results demonstrate that the present invention significantly reduces latency in image and audio recognition tasks across a wide range of data types and task requirements, while maintaining comparable accuracy to baseline models. Specifically, in terms of latency performance, the present invention achieves an average speedup of 2.19 times compared to cloud-side inference, 8.75 times compared to edge-side inference, and 4.47 times compared to edge-cloud collaborative inference methods based on early exit. Similarly, the present invention significantly reduces latency under varying network transmission conditions. Under low (5Mbps), medium (25Mbps), and high (55Mbps) network bandwidth conditions, the present invention achieves the lowest latency compared to cloud-side inference, edge-side inference, and traditional edge-cloud collaborative inference methods. Under low bandwidth conditions, the present invention achieves latency reductions of 1.85 times, 4.82 times, and 2.07 times, respectively, compared to other methods. In terms of accuracy, the present invention achieves an average improvement of 4.0% over edge-side inference and 3.1% over early exit edge-cloud collaborative inference methods.

[0112] The above technical solution has the following advantages or beneficial effects: Compared with the existing technology, the present invention is based on the basic model of the Transformer architecture, outputs a scene-adaptive edge-customized model, and thus realizes accurate and efficient edge-cloud collaborative reasoning.

[0113] The present invention designs a scenario-adaptive edge-cloud collaborative reasoning system for basic models to solve the current deployment problem of basic models on edge devices.

[0114] The present invention utilizes a single powerful basic model to adaptively output edge-customized models based on the data types and task requirements in different scenarios; uploads local and important data blocks of the original data to reduce transmission costs and protect data privacy; and adaptively adjusts the edge-cloud collaborative reasoning strategy to achieve efficient and accurate reasoning.

[0115] At the same time, the present invention supports further expansion and application of the basic model in applications such as intelligent monitoring, intelligent assistants, virtual / augmented reality, etc.

[0116] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various modifications and substitutions within the technical scope disclosed in the present invention, and such modifications and substitutions are intended to be within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be subject to the scope of protection of the claims.

Claims

1. A scenario-adaptive edge-cloud collaborative reasoning system for a basic model, comprising: Scenario-adaptive side model customization components and adaptive reasoning components; The scene-adaptive side model customization component includes: Hybrid Expert Model Extraction Module: This module extracts a hybrid expert model from the base model deployed on the cloud side through superclass routing and model distillation. It then divides domain-specific knowledge into superclasses for the different expert modules in the hybrid expert model. The superclasses for each expert module are derived from the categories in the public dataset. The hybrid expert model serves as a pre-trained model for the custom model on the edge side. Edge-customized model fine-tuning module: After extracting the hybrid expert model, it selects the most important expert module from each hybrid expert layer based on the edge device's data type and task requirements. It compresses the hybrid expert model and fine-tunes it using the edge device's unlabeled raw data to produce an edge-customized model. This fine-tuning approach combines model distillation and contrastive learning, using the base model deployed on the cloud as the teacher model for model distillation. Furthermore, the cloud-side text feature library T is used to select the text features required for contrastive learning. The adaptive reasoning component includes: Upload decider: This is used to evaluate the confidence score of the edge-customized model inference result. If the confidence score is higher than the set confidence threshold, the edge-customized model inference result is adopted. Otherwise, the data sample is input into the upload data block selector. Upload data block selector: used to evaluate the importance of data blocks of input data samples, select and upload local and important data sample blocks to the basic model deployed on the cloud side; Collaborative Strategy Adjuster: This is used to estimate real-time network bandwidth to update transmission latency, query the threshold search table, calculate the predicted end-to-end inference latency, and dynamically adjust the confidence threshold of the upload decider and the number of data sample blocks of the upload data block selector based on different data types and task requirements to balance the latency and accuracy of the inference task.

2. The scenario-adaptive edge-cloud collaborative reasoning system for a basic model according to claim 1, characterized in that: The hybrid expert model extraction module includes several backbone network layers, a superclass routing, several hybrid expert layers and a feature mapping layer. The hybrid expert layer includes several expert modules responsible for different superclasses.

3. The scenario-adaptive edge-cloud collaborative reasoning system for a basic model according to claim 1, characterized in that: The importance calculation process of the side-customized model fine-tuning module is as follows: The importance is evaluated by calculating the probability of each expert module being selected using the data samples of the edge device. The data samples are sent to the expert module e by the superclass routing. j The importance is expressed as I j , the calculation formula is: P i =softmax(W r ), Among them, P i ={P i,1 ,P i,2 ,...,P i,N }, represents the probability that data sample i is assigned to N expert modules, W r is the N-dimensional probability distribution vector output by the superclass routing, P i,j The data sample i is sent to the expert module e by the superclass routing j The probability of , i is the data sample number, j is the expert module number.

4. The scenario-adaptive edge-cloud collaborative reasoning system for a basic model according to claim 1, characterized in that: The fine-tuning process of the side-customized model fine-tuning module is as follows: Input the data samples into the basic model and the side customized model respectively, and calculate the sample features extracted by the basic model and sample features extracted by the side-customized model The mean square error loss is used as the loss function L for model distillation dist , the calculation formula is: At the same time, in order to adapt to the unlabeled data environment of edge devices, the loss function L of contrastive learning is set con , query the sample features extracted from the text feature library and the basic model Most similar text features Closely define the sample features extracted by the side-customized model and text features The distance, the loss function L of contrastive learning con The calculation formula is: L con =L I,T +L T,I , Where sim(·) is the cosine similarity of the features, and τ is the temperature coefficient; The loss function L for side-by-side custom model fine-tuning is the loss function L for model distillation. dist And the loss function L of contrastive learning con Combination of: in, is the weight of each data sample.

5. The scenario-adaptive edge-cloud collaborative reasoning system for a basic model according to claim 1, characterized in that: The confidence score calculation process of the upload decision maker module is as follows: The data sample is input into the side-customized model, the sample features are output, the cosine similarity between the sample and each text feature in the text feature library is calculated, the entropy value is calculated based on the cosine similarity distribution, and the entropy value is normalized to obtain the confidence score of the input data sample.

6. The scenario-adaptive edge-cloud collaborative reasoning system for a basic model according to claim 1, characterized in that: The upload block selector module evaluates the importance of the data blocks of the input data samples as described in the following process: The cosine similarity between the class token and the data tokens of the last layer of the side custom model is used as the category attention score A cls , represents the importance of the data block of the input data sample, and the calculation formula is: Among them, q cls Mark the class token, K T is the key matrix, that is, the set of all input data labels, and d is the dimension of category labels and data labels.

7. The scenario-adaptive edge-cloud collaborative reasoning system for a basic model according to claim 1, characterized in that: The end-to-end inference latency of the collaborative strategy adjuster is calculated as follows: Regularly collect raw data samples from edge devices, calculate the ratio r(thre) of data samples processed by the edge-side custom model and the basic model on the cloud side, the overall accuracy acc(thre), and the edge-side processing delay l e , transmission delay l tr (thre), cloud-side processing delay l c , calculate end-to-end inference latency And store it in the threshold search table, the calculation formula is as follows:

8. A scenario-adaptive edge-cloud collaborative reasoning method for a basic model, performed in a scenario-adaptive edge-cloud collaborative reasoning system for a basic model as described in any one of claims 1 to 7, comprising the following steps: S1 extracts a hybrid expert model from the base model deployed on the cloud side through superclass routing and model distillation. It then divides domain-specific knowledge into superclasses for different expert modules in the hybrid expert model. The superclasses for each expert module are derived from the categories of the public dataset. The hybrid expert model serves as a pre-trained model for the side-customized model. After S2 extracts the hybrid expert model, it selects the most important expert module from each hybrid expert layer based on the edge device's data type and task requirements. It then compresses the hybrid expert model and fine-tunes it using the edge device's unlabeled raw data to produce an edge-customized model. This fine-tuning approach combines model distillation and contrastive learning, using the base model deployed on the cloud as the teacher model for model distillation. Furthermore, the cloud-side text feature library is used to select the text features required for contrastive learning. The side-customized model in S3 divides the input data sample into multiple data blocks and calculates the preliminary inference results; then, the preliminary inference results are passed to the upload decision maker; S4 evaluates the confidence score of the edge-customized model inference result. If the confidence score is higher than the set confidence threshold, the edge-customized model inference result is adopted. Otherwise, the data sample is input into the upload data block selector. S5 evaluates the importance of data blocks of input data samples, selects and uploads local and important data sample blocks to the basic model deployed on the cloud side, and uses the basic model deployed on the cloud side to infer the results; S6 estimates the real-time network bandwidth to update the transmission delay, queries the threshold search table, calculates the predicted end-to-end inference delay, and dynamically adjusts the confidence threshold of the upload decision maker and the number of data sample data blocks of the upload data block selector according to different data types and task requirements to balance the latency and accuracy of the inference task.

Citation Information

Patent Citations

  • Customized deep neural network model compression method and system based on cloud edge cooperation

    CN112486686A

  • Cloud edge collaborative fine tuning control method and device, equipment and storage medium

    CN117478713A