Large language model end cloud collaborative inference system based on low-rank fine tuning
By deploying a low-rank fine-tuned lightweight language model on the end and combining it with intelligent task scheduling and cache optimization, the high cloud computing cost problem of large-scale language model inference services is solved, task accuracy and response speed are improved, and the efficiency and adaptability of end-cloud collaborative reasoning are achieved.
Patent Information
- Application Number
- CN202511253301.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-03
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2045-09-03
AI Technical Summary
Existing technologies make it difficult to reduce the cloud computing costs of large-scale language model inference services without sacrificing user experience. In addition, the computing resources and storage capabilities of end-side devices are limited, resulting in a decline in user experience.
Low-rank fine-tuning technology is used to deploy lightweight language models on the end side, and combined with intelligent task scheduling and cache optimization mechanisms, high-frequency tasks are executed on the end side through low-rank adapters and intelligent classifiers, and complex tasks are dynamically allocated to cloud computing to optimize end-cloud collaborative reasoning.
It significantly reduces cloud computing costs and storage overhead, improves task accuracy and response speed, and achieves high efficiency and adaptability of end-cloud collaborative reasoning.
Smart Images

Figure CN120806170A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of inference optimization of edge-cloud computing, and particularly relates to an edge-cloud collaborative inference system for large language models based on low-rank fine-tuning. BACKGROUND
[0002] In recent years, inference services of large language models have been widely used globally by users to perform various general machine learning tasks, including but not limited to text translation, content summarization, knowledge retrieval and logical reasoning. With the rapid growth of the user base, in order to ensure service quality, service providers have to continuously expand cloud computing resources. However, the direct consequence of this expansion is a sharp rise in the cost of inference services, and in some cases, the inference cost has even exceeded the pre-training stage of large language models. Among them, the high usage cost of GPUs in cloud servers has become the main component of the total cost of inference services. In the face of this challenge, a new research direction has gradually attracted attention in recent years: taking advantage of the increasing computing power and storage efficiency of terminal devices to reduce the dependence of large language model inference services on cloud computing resources. In particular, deploying lightweight small language models on the edge side for inference has become a promising solution. Although small language models for edge devices have been emerging in recent years, and their quality has been continuously improving, due to the limited computing resources and storage capacity of edge devices, the decline in user experience when using these small models for inference is still unavoidable.
[0003] Studies have shown that increasing the parameter size of large language models can significantly improve their accuracy and response capabilities. Generally speaking, it usually requires at least 1 billion parameters to make a language model have strong expressive power, and to achieve multi-task understanding, the parameter quantity usually needs to exceed 3 billion. However, due to the memory capacity and power consumption constraints of edge devices, small language models based on the Transformer architecture are difficult to break through the scale of 8 billion parameters. This limitation in parameter quantity directly leads to a significant gap in the performance of edge-side models in handling complex tasks compared to large-scale language models. Therefore, how to reduce the cloud computing cost of large language model inference services without sacrificing user experience has become an important scientific problem that needs to be solved.
[0004] In response to this challenge, the academic and industrial communities have proposed various optimization schemes, mainly including two major strategies: model compression and system collaboration.
[0005] Model compression: Model compression techniques aim to reduce the computational and storage requirements of language models, enabling them to run more efficiently in constrained hardware environments. Current mainstream methods include knowledge distillation, pruning, low-rank decomposition, and quantization. Knowledge distillation trains a small "student model" to mimic the output of a large "teacher model," thereby preserving some performance with a smaller model size. Pruning techniques remove network weights that contribute less to inference, reducing computational complexity. Low-rank decomposition uses matrix decomposition methods to reduce model parameter quantity, while quantization reduces the bit precision of model parameters to reduce storage and computational overhead. However, although these methods can effectively reduce computational costs, they often sacrifice some inference accuracy, and the adaptability and universality of the model are also affected.
[0006] System synergy: At the system level, researchers explore various end-to-cloud collaborative inference strategies to seek the best balance between computational cost, accuracy, and inference delay. Among them, the more typical methods include model partitioning and multi-model collaborative inference. Model partitioning method splits large language models into multiple sub-models, part of the computation is performed on the edge side, and the remaining computation tasks are completed by the cloud. This method can reduce the computational burden on the edge side while retaining large model capabilities. However, due to the delay overhead of end-to-cloud data transmission, the model partitioning method still faces challenges in application scenarios with high real-time requirements. Multi-model collaborative inference (such as speculative decoding) attempts to run a lightweight model on the edge side to generate partial inference results, which are then corrected or completed by the cloud model. This strategy reduces cloud computing demand to some extent, but is still affected by the round-trip time of mobile networks, and may not meet user needs in high interactivity tasks.
[0007] Overall, reducing the cloud computing cost of large language model inference services is a comprehensive problem involving model architecture, system optimization, and intelligent scheduling. Although industry and academia have made extensive efforts to optimize video parameter configurations, large language model end-to-cloud collaborative inference technology still needs to explore more efficient computing architectures and collaborative mechanisms while ensuring user experience. SUMMARY
[0008] The present invention proposes a large language model end-to-cloud collaborative inference system based on low-rank fine-tuning, aiming to optimize the inference service of large language models, reducing cloud computing costs while improving task accuracy and reducing end-to-end inference delay.
[0009] The large language model end-to-cloud collaborative inference system based on low-rank fine-tuning has the following steps in the construction process:
[0010] Step one, establish an end-cloud collaborative reasoning architecture, deploy a lightweight language model on the end side, and introduce a low-rank adapter and an intelligent classifier; the cloud side stores the low-rank adapter set in the low-rank adapter library for distribution and calling of cloud side calculation after end side task classification;
[0011] Step two, in the offline stage, the cloud side fine-tunes the large language model based on the training data of different downstream tasks;
[0012] The fine-tuning process is as follows:
[0013] The weight matrix of the pre-trained model is decomposed into a direction component and an amplitude component, and the direction component is frozen and the amplitude component is retained; at the same time, the shared low-rank matrix A and the task-specific matrix B are initialized, and the matrix A is frozen and the B is trainable.
[0014] The pre-trained model weight is decomposed into a direction component and an amplitude component , and the mathematical expression is:
[0015]
[0016] wherein, is the column norm of the weight matrix, used to capture the amplitude component; is the direction component obtained by column normalization.
[0017] In the low-rank fine-tuning process, the shared frozen matrix is only updated with the amplitude component and the task-specific matrix of the low-rank fine-tuning. The specific weight update process is represented as:
[0018]
[0019] and represent the updated amplitude component and the specific task matrix respectively; represents the initial weight matrix; represents the column norm of the parameter after using LoRA.
[0020] The end-side device only needs to store a single matrix during the training process, and when needed, it updates the small language model by transmitting the low-rank adapter of the new task (i.e., the matrix and the fine-tuned amplitude parameter ).
[0021] Step three, in the online stage, when the user makes a request, the user request is classified by the "variational autoencoder-gaussian mixture model" clustering, and it is judged whether there is a low-rank adapter (B, m) in the end-side cache that matches the current task. If yes, load it to the end-side small language model and execute reasoning to get the task result; otherwise, forward the task to the cloud side for reasoning.
[0022] The process of task classification is modeled by the following formula:
[0023]
[0024]
[0025] wherein, and represent the input data (text) and the hidden variable, respectively; and are the outputs of the variational autoencoder, representing the mean and variance of the latent variable.
[0026] The process of task classification is as follows: the variational autoencoder maps the input data to a latent space conforming to a Gaussian distribution through the mean and variance of the latent variable. Then, the Gaussian mixture model models the task subcategories through multiple Gaussian components to obtain the task classification result. According to the result of task classification, the system can decide whether the task is executed on the end-side device or forwarded to the cloud side for processing.
[0027] Step four, when the architecture processes several user requests, the Mamba model is used to analyze user historical requests and cache states, and the end-side low-rank adapter library is dynamically updated, with high-frequency task parameters being preferentially retained.
[0028] The user historical access sequence and the current end-side low-rank adapter library storage state are used as the dual-modal data, and the optimal update decision is made by learning the decision of the algorithm, and the specific process is as follows:
[0029] In the training process of the model, the Mamba model based on the state space model architecture is used to extract the time features of user access behavior. The hidden state update equation of the Mamba model is as follows:
[0030]
[0031] wherein, is the selective update gate, represents the state space model, which supports parallel computing. is the residual coefficient. In the inference stage, this calculation process is converted to a serial mode.
[0032] To further improve the prediction ability, the system uses two independently initialized projection matrices to respectively embed the user access data and the cache information:
[0033]
[0034]
[0035] wherein, represents the embedded user access data, is the embedded cache information; the position encoding is generated by a sinusoidal function.
[0036] The user access data and the cache information are projected into a unified feature space, and the fused features are represented by the spliced features
[0037]
[0038] wherein, is the projection matrix, which projects into the hidden space together with
[0039] In the training process, a binary cross-entropy loss function is used to guide the model to learn the task classification and the cache replacement strategy:
[0040]
[0041] wherein, i=1,…,Q is the number of the edge-side cache, is the strategy prediction result, is LayerNorm ( ). refers to the actual result, is a vector with a length of Q, is the strategy of whether the i-th position of the edge-side cache should be replaced.
[0042] The cache replacement strategy is to release a certain adapter in the edge-side low-rank adapter library and replace it with an adapter suitable for a new task classification. When the cache replacement strategy makes the above loss function take the minimum value, the updated edge-side low-rank adapter library is obtained, and the dynamic nature of the adapter in the edge-side cache is maintained through this strategy.
[0043] Step five, real-time monitoring of the edge-cloud load and inference delay, triggering the cloud-side fallback when the edge-side confidence is insufficient, and incrementally issuing new adapters to the edge side according to the task repetition rate.
[0044] The trigger condition for the cloud side to issue a new adapter to the end side is to set a hyperparameter as a cloud access threshold, count the number of cloud accesses of a task in a unit time, and if the cloud side has a corresponding adapter, when the number of cloud accesses of the task exceeds the threshold, the cloud side deploys a new adapter to the end side.
[0045] Step six, periodically optimize the task classification threshold and the end side cache capacity to achieve the dynamic balance of cost, delay and accuracy.
[0046] The advantages and beneficial effects of the present application are that:
[0047] The end-to-cloud collaborative reasoning system based on low-rank fine-tuning of large language model proposed by the present application reduces the parameter transmission and storage requirements, optimizes task classification and distribution, and dynamically updates model parameters, while reducing the computing and storage overhead, ensuring the efficiency and adaptability of the system. BRIEF DESCRIPTION OF DRAWINGS
[0048] Figure 1 is the architecture diagram of the end-to-cloud collaborative reasoning system based on low-rank fine-tuning of large language model of the present application;
[0049] Figure 2 is the implementation flowchart of the end-to-cloud collaborative reasoning system based on low-rank fine-tuning of large language model of the present application. DETAILED DESCRIPTION
[0050] The present application will be further described in detail below with reference to the accompanying drawings and examples.
[0051] The existing end-to-cloud collaborative reasoning method has limited reasoning capability on the end side device, which is difficult to adapt to the demand of complex tasks. And these methods usually rely on cloud computing resources for main reasoning, while the end side only performs shallow calculation, which cannot fully exert the potential of the end side device, resulting in low utilization rate of the overall system computing resources. Therefore, based on the following key observations, the present application proposes a novel end-to-cloud collaborative reasoning system:
[0052] Firstly, although the end side device has limited computing power, it still has the potential to perform part of the reasoning task, especially for high-frequency and low-complexity tasks, more calculations can be performed on the end side through appropriate model optimization, thereby reducing the load of the cloud side. Secondly, most of the existing end-to-cloud collaborative methods adopt static segmentation strategy, which cannot adaptively adjust according to the dynamic changes of task complexity, resulting in limited reasoning efficiency. Finally, the low-rank fine-tuning technology has shown significant computing and storage efficiency advantages in model optimization, making it an ideal choice for end side model optimization.
[0053] Therefore, the application combines the low-rank fine-tuning technology to deploy an efficient lightweight model on the edge side, and realizes the optimization goal of end-to-cloud collaborative inference through an intelligent task scheduling and cache optimization mechanism, thereby improving inference efficiency and reducing overall computing overhead. The core idea is to introduce a small number of low-rank adapters into the edge-side model to enhance the inference ability of the edge-side model for high-frequency tasks, while intelligently classifying complex tasks and dynamically allocating them to cloud computing, thereby reducing cloud load while improving inference accuracy and reducing storage and transmission overhead of edge devices.
[0054] As shown in Figure 1 , the overall architecture of the system is composed of an offline stage and an online stage, which cooperatively optimize the allocation strategy of end-to-cloud computing resources.
[0055] 1. In the offline stage, the system performs decoupling of large model weights, optimization of low-rank adapters, and edge-side storage management to reduce end-to-cloud transmission overhead and fine-tuning computing cost.
[0056] Existing large language model fine-tuning schemes usually rely on complete parameter updates or adapter fine-tuning, resulting in high storage overhead and transmission cost, making it difficult to apply to end-to-cloud collaborative environments. Therefore, the present application proposes a resource-friendly low-rank fine-tuning method to reduce the fine-tuning computing cost and storage overhead of large language models.
[0057] Firstly, this method decomposes the weight matrix of the pre-trained model into a direction component and an amplitude component, so that the edge device only needs to store a low-rank matrix without storing all model parameters in full, thereby significantly reducing storage pressure and transmission cost. At the same time, this stage uses large-scale data for model training to ensure that the edge-side lightweight model still has high inference ability under limited computing resources. In addition, this method freezes part of the shared parameters to ensure that models for different tasks on the edge side can share optimized direction components, thereby reducing the computing demand on the edge side and improving the generalization ability of the low-rank adapter.
[0058] This algorithm significantly reduces the amount of parameters in the fine-tuning process by decomposing the weight matrix of the pre-trained model, while maintaining high efficiency after fine-tuning. Specifically, this algorithm decomposes the weight matrix of the pre-trained model into a direction component and an amplitude component , and the mathematical expression is as follows:
[0059]
[0060] wherein, is the column norm of the weight matrix, which is used to capture the amplitude component; is the direction component obtained by column normalization. This decomposition method can effectively reduce the computational complexity in the fine-tuning process and provide a more streamlined parameter structure for subsequent task optimization.
[0061] In the actual fine-tuning process, the algorithm updates only the amplitude component and the low-rank fine-tuned matrix by sharing the frozen matrix . This design allows the matrix to be shared and frozen during the fine-tuning process of different tasks, while the matrix is optimized according to the task requirements. The specific weight update process can be represented as:
[0062]
[0063] In this process, the matrix is fixed and shared across different tasks, while the matrix is only used to adjust for specific tasks. A key advantage of this design is that it reduces the amount of transmitted parameters by approximately 50%, and the effect of low-rank fine-tuning is focused on optimizing the direction component, ensuring the stability of the model's performance after fine-tuning.
[0064] In addition, the end-side device only needs to store a single matrix during the training process and update the model by transmitting the low-rank adapter (i.e., the matrix and the fine-tuned amplitude parameter ) of the new task when needed. This solution not only reduces the computational and storage overheads but also reduces the burden of data transmission, enabling more efficient inference on resource-constrained devices.
[0065] 2. In the online phase, the system implements intelligent scheduling of tasks based on task classification and cache management strategies.
[0066] During online inference, the system classifies tasks using the imitation learning method and combines dynamic updating of the end-side cache to improve task allocation efficiency and reduce inference delay. To further optimize end-side inference efficiency, the system uses a cache management method based on a state space model to dynamically update the end-side low-rank adapter, allowing it to adapt to users' inference needs at any time and improve the effective hit rate of the end-side model.
[0067] The key task of the online stage is to optimize the execution of inference tasks through efficient task classification and dynamic allocation. In this stage, a user-aware model scheduling algorithm based on imitation learning is proposed, which can dynamically adjust the execution mode of inference tasks according to the characteristics of user requests. First, task classification is performed by combining variational autoencoder and Gaussian mixture model. This process can efficiently identify the underlying structure between different tasks and allocate appropriate resources for each task. The task classification process can be modeled by the following formula:
[0068]
[0069]
[0070] where, and are the outputs of the variational autoencoder, representing the mean and variance of the latent variable. Through these outputs, the variational autoencoder can effectively map the input data to a latent space that conforms to the Gaussian distribution. The Gaussian mixture model models the task subcategories through multiple Gaussian components, effectively classifying the tasks. According to the results of task classification, the system can decide whether the task is executed on the edge device or forwarded to the cloud for processing.
[0071] To further improve the adaptability of the edge model, the system designs a method to dynamically update the edge low-rank fine-tuning parameter library. This method uses the user's historical access sequence and the current edge low-rank fine-tuning library storage state dual modal data, through learning algorithm decision to make the optimal update decision.
[0072] During the training process of the model, the Mamba model based on the state space model architecture is used to extract the time features of user access behavior. The hidden state update equation of the Mamba model is:
[0073]
[0074] where, is the selective update gate, represents the state space model, which supports parallel computing. In the inference stage, this computing process is converted to a serial mode.
[0075] To further improve the prediction ability, the system uses two independently initialized projection matrices and to embed the user access data and cache information respectively:
[0076]
[0077]
[0078] Among them, position encoding It is generated by the sine function. This process projects the user access data and cache information into a unified feature space. To represent the fused features:
[0079]
[0080] This fusion method can preserve the complete information of the original features and effectively improve the accuracy of feature extraction and prediction. During the training process, a binary cross entropy loss function is used to guide the model to learn task classification and cache replacement strategy:
[0081]
[0082] in, It is the strategy prediction result.
[0083] The cache replacement strategy involves releasing an adapter from the client-side low-rank adapter library and replacing it with an adapter suitable for the new task classification. When this cache replacement strategy minimizes the aforementioned loss function, the updated client-side low-rank adapter library is obtained. This strategy maintains the dynamic nature of the client-side cached adapters. In this way, the algorithm can effectively learn cache update strategies from historical data, thereby improving the client-side cache hit rate and optimizing task allocation efficiency.
[0084] like Figure 2 As shown in the figure, the specific implementation of the large language model end-cloud collaborative reasoning system based on low-rank fine-tuning includes the following main steps:
[0085] Step (1) In the offline phase, the pre-trained model weights are decomposed into direction components and amplitude components, the shared low-rank matrix A and the task-specific matrix B are initialized, matrix A is frozen and B is kept trainable.
[0086] In step (2), user requests are classified online through “Variational Autoencoder-Gaussian Mixture Model” clustering. If a matching low-rank adapter (B, m) exists in the client-side cache, it is loaded into the client-side model for inference. Otherwise, it is forwarded to the cloud.
[0087] Step (3) Analyze user historical requests and cache status based on the Mamba model, dynamically update the low-rank adapter library on the end side, and prioritize retaining high-frequency task parameters.
[0088] Step (4) monitors the end-cloud load and inference latency in real time. When the end-side confidence is insufficient, it triggers a cloud fallback and sends a new adapter to the end-side based on the task repetition rate increment.
[0089] The trigger condition for the cloud side to issue a new adapter to the end side is to set a hyperparameter as a cloud visit threshold, count the number of cloud visits of a task in a unit time, and if the cloud side has a corresponding adapter, when the number of cloud visits of the task exceeds the threshold, the cloud side deploys a new adapter to the end side.
[0090] Step (5) periodically optimizes the task classification threshold and the cache capacity to achieve the dynamic balance of cost, delay and accuracy.
[0091] In summary, the application proposes a large language model end-cloud collaborative reasoning system architecture based on low-rank fine-tuning, aiming to solve the key problems of high cloud computing resource cost, task accuracy and reasoning delay balance. The system uses a resource-efficient low-rank adaptation algorithm combined with a user perception model replacement algorithm based on imitation learning, which reduces the parameter transmission volume by 50% while maintaining the model reasoning accuracy, thereby significantly reducing the cloud computing cost. Through end-cloud task dynamic allocation and optimization, the application effectively balances the model accuracy (task accuracy), end-side response speed (delay) and resource efficiency.
Claims
1. A large language model end-cloud collaborative reasoning system based on low-rank fine-tuning, characterized by: The build process is as follows: In the first phase, we will establish a collaborative inference architecture between the device and the cloud. The architecture includes both the device and cloud sides. The device side deploys a lightweight language model and introduces low-rank adapters and intelligent classifiers. The cloud side stores the low-rank adapter set in a low-rank adapter library, which is used to classify device-side tasks and allocate them to cloud-side computations. The process of establishing this architecture includes offline and online stages, specifically: (1) In the offline phase, the cloud side fine-tunes the parameters of the large language model based on the training data of different downstream tasks; (2) In the online phase, when a user makes a request, the "Variational Autoencoder-Gaussian Mixture Model" clustering is used to classify the user request and determine whether there is a low-rank adapter (B, m) matching the current task in the client cache. If so, it is loaded into the client-side small language model, and reasoning is performed to obtain the task result; otherwise, the task is forwarded to the cloud side for reasoning. In the second phase, the end-cloud collaborative inference architecture is monitored and updated in real time based on the processing of user requests. S1, after the architecture has processed several user requests, it analyzes historical user requests and cache status based on the Mamba model, dynamically updates the low-rank adapter library on the client side, and prioritizes retaining high-frequency task parameters; Utilize user history access sequence and the current low-rank adapter library storage state on the client side Bimodal data, through learning The algorithm makes the best update decision. The specific process is: During the model training process, the Mamba model based on the state space model architecture is used to extract the temporal characteristics of user access behavior. The hidden state update equation of the Mamba model is: in, is the selective update gate, Represents state space models and supports parallel computing. is the residual coefficient; To further improve the prediction capability, the system uses two independently initialized projection matrices and Embed user access data and cache information separately: in, Indicates the embedded user access data. For embedded cache information; position encoding It is generated by the sine function; Project user access data and cache information into a unified feature space, and then use the concatenated features To represent the fused features: in, is the projection matrix, through Will Projection to and The hidden space together; During training, a binary cross entropy loss function is used to guide the model to learn task classification and cache replacement strategy: Where i=1,…,Q is the number of caches on the client side. is the strategy prediction result, is LayerNorm ( ); Refers to the actual results, is a vector of length Q, The policy for determining whether the client-side cache at position i should be replaced. The cache replacement strategy is to release an adapter from the low-rank adapter library on the end and replace it with an adapter suitable for the new task classification. When the cache replacement strategy minimizes the value of the above loss function, the updated low-rank adapter library on the end is obtained. This strategy maintains the dynamic nature of the adapters cached on the end. S2 monitors the end-cloud load and inference latency in real time. When the end-side confidence is insufficient, it triggers a cloud-side fallback and incrementally sends a new adapter to the end-side based on the task repetition rate. S3 periodically optimizes task classification thresholds and client-side cache capacity to achieve a dynamic balance between cost, latency, and accuracy.
2. The large language model end-cloud collaborative reasoning system based on low-rank fine-tuning according to claim 1 is characterized in that The fine-tuning process is as follows: Decompose the weight matrix of the pre-trained model into direction components and amplitude components, freeze the direction components and retain the amplitude components; at the same time, initialize the shared low-rank matrix A and the task-specific matrix B, freeze matrix A and keep B trainable; During low-rank fine-tuning, the frozen matrix is shared , only the amplitude component is updated and the task-specific matrix for low-rank fine-tuning ; The specific weight update process is expressed as: and denote the updated amplitude component and task-specific matrix, respectively; represents the initial weight matrix; Indicates the column norm of the parameters after using LoRA; The end device only needs to store a single matrix during training , and when needed, by transferring new tasks to the low-rank adapter ( ) to perform small language model updates.
3. The large language model end-cloud collaborative reasoning system based on low-rank fine-tuning according to claim 2 is characterized in that The weights of the pre-trained model Decompose into directional components and amplitude components , its mathematical expression is: in, is the column norm of the weight matrix, which is used to capture the amplitude component; is the directional component obtained by column normalization.
4. The large language model end-cloud collaborative reasoning system based on low-rank fine-tuning according to claim 1 is characterized in that User requests are classified through "Variational Autoencoder-Gaussian Mixture Model" clustering and modeled using the following formula: in, and Represent input data (text) and latent variables respectively; and is the output of the variational autoencoder, representing the mean and variance of the latent variable.
5. The large language model end-cloud collaborative reasoning system based on low-rank fine-tuning according to claim 1 or 4, characterized in that The process of classifying user requests is as follows: the variational autoencoder maps the input data to a latent space that conforms to a Gaussian distribution through the mean and variance of the latent variables; then, the Gaussian mixture model models the task subcategories through multiple Gaussian components to obtain the task classification results; finally, based on the task classification results, the system can decide whether to execute the task on the end device or forward it to the cloud for processing.
6. The large language model end-cloud collaborative reasoning system based on low-rank fine-tuning according to claim 1 is characterized in that The trigger condition for the cloud side to send a new adapter to the end side is: set a hyperparameter as the cloud access threshold, count the number of cloud accesses of a task per unit time, and if the cloud side has a corresponding adapter, then when the number of cloud accesses for the task exceeds the threshold, the cloud side will delegate the new adapter to the end side.
Citation Information
Patent Citations
Customized large model-oriented multi-model reasoning method and device
CN118195001A
Parameter efficient big language fine-tuning federal learning framework
CN118504526A
Cloud edge-end collaborative user behavior analysis system based on large language model
CN119415768A
Cloud road collaborative learning method of traffic multi-mode large language model
CN119692471A
Visual task processing method and device based on multi-modal large model of low-rank adaptive enhancement
CN120388272A
Cited By
Large language model low-delay reasoning method based on dynamic reasoning graph optimization
CN121072787A
A large language model low-latency inference method based on dynamic inference graph optimization
CN121072787B
Power grid field-oriented lightweight large language model fine tuning method
CN121094054A
Automatic driving model based on mixed low-rank experts and multi-domain adaptive fine tuning method
CN121523068A
Automatic driving model based on hybrid low-rank experts and multi-domain adaptation fine-tuning method
CN121523068B