A large model inference system and method based on an end-side cloud collaborative architecture

CN122819481APending Publication Date: 2026-09-25INST OF COMPUTING TECH CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611032063.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-13
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0004]综上所述,现有面向云边端架构的路由方案存在的问题是:路由计算的复杂度过高,并且忽略了网络状况、边缘与云端计算资源的异构性差异及不同节点的能耗成本等因素,导致难以在推理延迟、推理成本与推理质量之间取得最佳平衡

Benefits of technology

[0028]根据本发明的第二方面,提供一种基于本发明第一方面所述系统的推理方法,所述推理方法包括:T1、获取用户提出的问题。T2、将获取的问题输入基于端边云协同架构的大模型推理系统以获得该问题的答复。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122819481A_ABST
    Figure CN122819481A_ABST
Patent Text Reader

Abstract

The application provides a large model inference system based on an end-edge-cloud collaborative architecture, and the end side of the system is configured with a two-stage router, which comprises: a policy router configured to receive a question raised by a user and select an inference policy for the question; and a model router configured to receive the question raised by the user and the inference policy selected by the policy router, select a node under the constraint of the inference policy, and route the question and the selected inference policy to the selected node, so that the node invokes a large model thereon to perform inference according to the inference policy and obtains a reply to the question. The application decouples the strong coupling joint routing into two stages of policy selection and model selection, thereby significantly reducing the complexity of routing calculation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of interdisciplinary technology of artificial intelligence and edge computing, specifically to the field of large model inference optimization technology, and more specifically to a large model inference system and inference method based on an edge-cloud collaborative architecture. Background Technology

[0002] With the rapid development of Large Language Model (LLM) technology, it faces the challenge of heterogeneous deployment environments in practical applications. Typical deployment architectures include edge, cloud, and mobile devices. Small parameter models are usually deployed on the edge to reduce latency, while large parameter models are deployed on the cloud to ensure capabilities. Furthermore, various inference strategies can be used for different types of user questions, such as Direct Answer, Chain-of-Thought (CoT), and Tree-of-Thoughts (ToT). Different strategies exhibit significant differences in performance across different task types.

[0003] In the "multi-policy, multi-model" choice space of the aforementioned cloud-edge-device architecture, the optimal routing decision for each user's question determines inference efficiency and the speed at which the user obtains the answer. Existing routing solutions for cloud-edge-device architectures mainly include three categories: single-dimensional routing, joint routing methods, and static cloud-edge-device allocation. Traditional single-dimensional routing methods typically only route to the model or policy individually. These methods do not consider the joint optimization of inference policy and model, only routing to the best-performing inference policy or model, but failing to obtain the optimal policy-model pair, thus limiting the inference optimization effect. Recent work has begun to explore joint routing methods that train a router to simultaneously predict the optimal policy-model combination. This method strongly couples policy selection with model selection, requiring all policy-model combinations to be predicted as independent categories. When multiple models and multiple inference policies exist, the router needs to process a large number of output categories. Static cloud-edge-device allocation schemes are mostly based on static rules, lacking fine-grained awareness of the matching degree between inference policy and model capabilities, and cannot adapt to dynamic changes in real-time resource states such as network bandwidth and node load, resulting in limited routing policy quality.

[0004] In summary, the existing routing solutions for cloud-edge-device architectures have the following problems: the complexity of routing calculations is too high, and they ignore factors such as network conditions, heterogeneity differences between edge and cloud computing resources, and energy consumption costs of different nodes, making it difficult to achieve the best balance between inference latency, inference cost, and inference quality.

[0005] It should be noted that the background information presented here is only for illustrating relevant information about the present invention to aid in understanding the technical solution of the present invention, and does not imply that the relevant information is necessarily prior art. The relevant information was submitted and disclosed together with the present invention, and should not be considered prior art unless there is evidence that the relevant information was disclosed before the filing date of the present invention. Summary of the Invention

[0006] Therefore, the purpose of this invention is to overcome the shortcomings of the prior art and provide a large model inference system and inference method based on an edge-cloud collaborative architecture.

[0007] The objective of this invention is achieved through the following technical solution:

[0008] According to a first aspect of the present invention, a large model inference system based on an edge-cloud collaborative architecture is provided, wherein multiple nodes are deployed on both the edge and cloud sides of the system, and each node deploys a large model. The edge side of the system is used to receive a question raised by a user, select an inference strategy and a node for the question, and route the question and the inference strategy to the selected node so that the node calls the large model on it to perform inference according to the inference strategy to obtain an answer to the question. The edge side of the system is configured with a two-stage router, the two-stage router comprising: a policy router, which is obtained according to a predetermined policy router construction method, for receiving a question raised by a user and selecting an inference strategy for the question; and a model router, which is obtained according to a predetermined model router construction method, for receiving the question raised by the user and the inference strategy selected by the policy router, selecting a node under the constraints of the inference strategy, and routing the question and the selected inference strategy to the selected node so that the node calls the large model on it to perform inference according to the inference strategy to obtain an answer to the question.

[0009] This scheme can achieve at least the following beneficial technical effects: the system decouples the inference strategy selection and node selection into two independent stages. The first stage completes the selection of the inference strategy, and the second stage completes the selection of nodes under the constraints of the selected inference strategy. This realizes phased strategy-model joint optimization. Compared with the traditional coupled joint optimization scheme, the computational load can be significantly reduced from O(m×n) level to O(m+n) level.

[0010] Optionally, the policy router includes: a first encoding module for acquiring the question and extracting first question features; a first classification head for providing multiple prediction inference strategies and the probability corresponding to each prediction inference strategy based on the first question features, and taking the prediction inference strategy with the highest probability as the given inference strategy. The model router includes: a second encoding module for acquiring the question and extracting second question features; multiple second classification heads, each corresponding to a reasoning strategy, wherein when a reasoning strategy is given, the second classification head corresponding to the given reasoning strategy provides multiple prediction nodes and the probability corresponding to each prediction node based on the question; and a routing execution module for routing the question and the given reasoning strategy to the prediction node with the highest probability so that the node calls the large model on it to perform inference according to the given reasoning strategy to obtain the answer to the question.

[0011] Optionally, the policy router further includes a policy correction module, which is configured with a default inference policy and is used to compare the highest probability of the predicted inference policy output by the first classification head with a preset threshold. If the highest probability is greater than or equal to the preset threshold, the predicted inference policy corresponding to the highest probability is used as the given inference policy; otherwise, the default inference policy is used as the given inference policy.

[0012] This scheme can achieve at least the following beneficial technical effects: when determining a given inference strategy, the system also considers the confidence factor of its prediction results, making the inference strategy given by the system more accurate.

[0013] Optionally, the model router further includes: an environment detection module for acquiring real-time environment status, including the current load of each node and the network latency and network bandwidth from the end to each node; and a probability correction module for adjusting the probability corresponding to each predicted node output by the second classification head according to a preset calculation method based on the real-time environment status.

[0014] This scheme can achieve at least the following beneficial technical effects: when selecting nodes, the system also considers the real-time environmental status to adapt to the dynamic changes in the environment, and selects the node with the best overall conditions.

[0015] Optionally, the default calculation method in the probability correction module is:

[0016]

[0017] in, This represents the adjusted probability of the i-th prediction node; This represents the initial probability of the i-th prediction node; The real-time comprehensive resource cost penalty for the i-th prediction node is calculated by weighting the real-time environmental state. represents the sensitivity adjustment coefficient for real-time comprehensive resource cost penalty; j represents traversing all prediction nodes.

[0018] Optionally, the predetermined policy router construction method includes: S1, constructing an initial policy router, which includes: a first encoding module, used to acquire a question and extract a first question feature; and a first classification head, used to provide multiple prediction inference strategies and the probability corresponding to each prediction inference strategy based on the first question feature. S2, acquiring a policy router training set, which includes multiple training samples, each training sample including a question and a corresponding inference strategy label. S3, training the initial policy router to convergence using the policy router training set, wherein the training process takes the question as input, the prediction inference strategy and the probability corresponding to each prediction inference strategy as output, and updates the parameters of the first encoding module and the first classification head by minimizing the cross-entropy loss between the prediction result and the inference strategy label.

[0019] Optionally, the inference policy labels for the policy router training set are obtained as follows: S21. Collect a question set, an inference policy set, a model category set, and a model deployment node set, wherein the question set includes multiple questions, the inference policy set includes multiple inference policies, the model set includes multiple large models, and the model deployment location set includes different deployment nodes for each large model. S22. Traverse and combine each inference policy, each large model, and the nodes deployed by each large model to obtain multiple policy-model-node combinations. S23. For each question, perform inference one by one through each policy-model-node combination to obtain the answer to the question under each policy-model-node combination, and calculate the quality score of each answer according to a preset calculation method. S24. For each question, calculate its aggregate score under different inference policies, and take the inference policy with the highest aggregate score as the inference policy label for the question, wherein the aggregate score is calculated as follows:

[0020]

[0021] in, This represents the aggregate score corresponding to reasoning strategy j for a given question. This represents the calculation of the average, weighted average, or median. The quality score of the response given by the strategy-model-node combination JLP to question i is represented by this score.

[0022] This scheme achieves at least the following beneficial technical effects: it aggregates problem-strategy combinations across models to generate strategy-independent scores decoupled from model categories and deployment locations, thereby generating inference strategy labels. This approach solves the problem of inference strategy evaluation depending on specific models, making the selection of inference strategies by the system more generalizable and ensuring that the system can independently select the optimal inference strategy in the first stage.

[0023] Optionally, in S23, the preset calculation method is:

[0024]

[0025] in, The quality score of the response given by JLP for question i, representing the strategy-model-node combination. The correctness of the JLP's response to question i represents the combination of strategy, model, and node. The JLP strategy-model-node combination represents the number of lexical units consumed to provide a response to question i. The representative question is about the maximum word consumption across all strategy-model-node combinations. The transmission latency of the response to question i under the strategy-model-node combination JLP. The computation time of JLP, representing the strategy-model-node combination, to provide a response to question i. This represents the computing cost of a large model across different deployment nodes. , , , Each represents a different weight.

[0026] This solution can achieve at least the following beneficial technical effects: the calculation of quality score also takes into account factors such as resource consumption, network latency, and computing time, so that the system's routing decision not only focuses on the quality of the answer, but also comprehensively considers the actual time consumption and cost in the edge-cloud environment, making it suitable for complex edge-cloud architectures.

[0027] Optionally, the predetermined model router construction method includes: M1, constructing an initial model router, which includes: a second encoding module for acquiring a question and extracting second question features; multiple second classification heads, each corresponding to a reasoning strategy, wherein, given a reasoning strategy, the second classification head corresponding to the given reasoning strategy provides multiple prediction nodes and the probability corresponding to each prediction node based on the question; a routing execution module for routing the question and the given reasoning strategy to the prediction node with the highest probability so that the node calls the large model on it to perform reasoning according to the given reasoning strategy to obtain the answer to the question. M2, acquiring a model router training set, which includes multiple training samples, each training sample including a question, a given reasoning strategy, and corresponding node labels. M3, training the initial model router to convergence using the model router training set, wherein the training process takes the question and the given reasoning strategy as input, multiple prediction nodes and the probability corresponding to each prediction node as output, and updates the parameters of the second encoding module and all second classification heads by minimizing the cross-entropy loss between the prediction results and the node labels.

[0028] According to a second aspect of the present invention, a reasoning method based on the system described in the first aspect of the present invention is provided, the reasoning method comprising: T1, obtaining a question raised by a user; T2, inputting the obtained question into a large-scale model reasoning system based on an edge-cloud collaborative architecture to obtain an answer to the question.

[0029] Compared with existing technologies, the advantages of this invention are as follows: By decoupling the strongly coupled joint routing into two stages, policy selection and model selection, this invention significantly reduces the complexity of routing calculations; it also considers factors such as network conditions, heterogeneity differences between edge and cloud computing resources, and energy consumption costs of different nodes, so as to dynamically adapt to changes in network and computing resources. While ensuring inference quality, it significantly reduces inference latency and inference costs, effectively improving the overall efficiency and stability of the edge-cloud collaborative inference system, and has broad application prospects. Attached Figure Description

[0030] The embodiments of the present invention will be further described below with reference to the accompanying drawings, wherein:

[0031] Figure 1 This is a schematic diagram of a large model inference system based on an edge-cloud collaborative architecture according to an embodiment of the present invention.

[0032] Figure 2 This is a schematic diagram illustrating the inference policy tag construction process of the policy router construction method for a large-model inference system based on an edge-cloud collaborative architecture according to an embodiment of the present invention. Detailed Implementation

[0033] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are merely illustrative and are not intended to limit the invention.

[0034] As mentioned in the background technology section, the existing routing solutions for cloud-edge-device architecture have the following problems: the complexity of routing calculation is too high, and factors such as network conditions, heterogeneity differences between edge and cloud computing resources, and energy consumption costs of different nodes are ignored, making it difficult to achieve the best balance between inference latency, inference cost and inference quality.

[0035] The inventors' analysis revealed that existing joint routing methods tightly couple policy selection with model selection, leading to a product-like increase in routing complexity with the number of models and policies. This method requires predicting each policy-model pair to form a routing table, resulting in a huge exploration space and long processing time for joint routing, making it difficult to meet the real-time requirements of the edge. Existing joint routing methods do not consider decoupling policy selection and model selection into two independent stages because they assume joint optimization can obtain a globally optimal solution. However, in reality, inference policies are essentially properties of the problem itself and are relatively independent of specific models. A two-stage decomposition can significantly reduce computational complexity while maintaining routing quality. Furthermore, existing solutions suffer from insufficient resource awareness: focusing only on model accuracy or token consumption, ignoring network transmission latency in edge-cloud architectures, differences in computing resources between the edge and cloud, and energy costs. This results in an inability to achieve an optimal balance between latency, cost, and accuracy in actual deployments, leading to poor overall optimization performance.

[0036] While researching the optimization of inference cost and latency in large-scale models within an edge-cloud collaborative environment, the inventors discovered that inference strategies are essentially attributes of the problem itself. For example, mathematical problems are better suited to thought processes, while factual problems are better suited to direct answers. Therefore, inference strategies can be decoupled from specific models. Furthermore, existing methods, when predicting strategy-model combinations, often neglect network transmission latency and heterogeneity of computing resources in edge-cloud architectures. Routing complexity increases exponentially with the number of both, making it difficult to meet the real-time requirements of the edge. Therefore, the selection of model nodes should comprehensively consider computing resources and transmission latency. However, achieving independent evaluation of strategies is a technical challenge, as the same strategy performs differently in different models and at the edge or cloud. Moreover, decomposing joint routing into two stages means that an incorrect strategy selection in the first stage cannot be corrected in the second stage, leading to error propagation. Through in-depth research, the inventors discovered that by aggregating the performance of the same inference strategy across all available model nodes through a cross-model and cross-location aggregation voting mechanism, the strategy selection can be generalized, reducing errors. Regarding training data labeling, two-stage routing requires generating training labels for policy selection and model selection respectively. The inventors discovered that by using a multi-dimensional weighted inference score calculation method, all policy-model combinations of nodes are traversed. A comprehensive score is obtained by weighting the number of tokens consumed, correctness, network transmission latency, and computational resource cost. Then, cross-model aggregation is used to obtain independent scores for the policies, which can generate inference policy labels decoupled from specific models. In addition, to adapt to the real-time nature of edge routing and dynamic resource changes, lightweight routers can be deployed on the edge, and an uncertainty handling mechanism can be introduced to trigger a default robust inference policy when the policy confidence is low, while dynamically adjusting the model node selection probability according to the real-time network status.

[0037] Based on the above research, this invention proposes a large-model inference system based on an edge-cloud collaborative architecture. This system decouples policy selection from model selection, building upon existing joint routing schemes. In the first stage, the inference policy is selected independently; in the second stage, the optimal model node is selected based on a comprehensive assessment of resource status given the policy. This significantly reduces complexity, enhances resource awareness, and accelerates routing speed.

[0038] The purpose of this invention is to address the limitations of existing routing schemes for cloud-edge-device architectures, such as the inability to jointly optimize policies and models, resulting in limited inference performance; high routing complexity leading to high edge-side routing latency; and the uncontrollable overall latency due to neglecting the heterogeneity of edge-cloud resources. This invention first uses a multi-dimensional weighted inference score calculation module to traverse all policy-model combinations of all nodes and evaluate the comprehensive score of each combination. This score data is then used to train the model router for the second stage. Since the router used in the first stage needs to select the optimal inference policy without fixed models and nodes, the score data for each combination needs to be processed. Specifically, a cross-model policy generalization evaluation module is used to aggregate the performance of the same inference policy across all available models and nodes, generating inference policy labels decoupled from specific models, which are then used to train the first-stage inference policy router. Then, in the first stage, the optimal inference policy is independently selected based on the input problem. Next, in the second stage, based on the input problem and the selected inference policy, the model router selects the large model node with the best compatibility with that policy. Ultimately, the two-stage decomposition significantly reduces routing complexity, while simultaneously optimizing the inference strategy and model, thereby significantly reducing routing latency and improving the resource utilization efficiency and task completion quality of the entire large-scale model inference system.

[0039] The present invention will now be described in detail with reference to specific embodiments.

[0040] According to one embodiment of the present invention, referring to Figure 1The large-model inference system based on an edge-cloud collaborative architecture provided by this invention has multiple nodes deployed on both the edge and cloud sides. Each node deploys a large model. The edge side of the system receives user-submitted questions, selects an inference strategy and node for the question, and routes the question and inference strategy to the selected node so that the node can invoke its large model to perform inference and obtain an answer to the question. The edge side of the system is configured with a two-stage router, which includes: a policy router, obtained according to a predetermined policy router construction method, used to receive user-submitted questions and select an inference strategy for the question; and a model router, obtained according to a predetermined model router construction method, used to receive user-submitted questions and the inference strategy selected by the policy router, select a node under the constraints of the inference strategy, and route the question and the selected inference strategy to the selected node so that the node can invoke its large model to perform inference and obtain an answer to the question. After obtaining the answer, routing information, actual time consumption, and results can be recorded for subsequent incremental updates and system optimization.

[0041] The key point of this invention is to decouple policy selection from model selection based on existing joint routing methods. In the first stage, the inference policy is selected independently, and in the second stage, the optimal node is selected under the given inference policy. This realizes phased policy-model joint optimization. Compared with the traditional coupled joint optimization, the computational load can be significantly reduced from O(m×n) to O(m+n), thereby speeding up routing and reducing inference latency.

[0042] According to one embodiment of the present invention, the policy router includes: a first encoding module, configured to acquire a question and extract a first question feature; a first classification head, configured to provide multiple predictive inference strategies and the probability corresponding to each predictive inference strategy based on the first question feature, and to use the predictive inference strategy with the highest probability as a given inference strategy. The model router includes: a second encoding module, configured to acquire a question and extract a second question feature; multiple second classification heads, each second classification head corresponding to a different inference strategy, wherein, when a given inference strategy is provided, the second classification head corresponding to the given inference strategy provides multiple predictive nodes and the probability corresponding to each predictive node based on the question; and a routing execution module, configured to route the question and the given inference strategy to the predictive node with the highest probability, so that the node calls the large model on it to perform inference according to the given inference strategy to obtain the answer to the question.

[0043] According to one embodiment of the present invention, the policy router further includes a policy correction module. The policy correction module is configured with a default inference policy and is used to compare the highest probability of the predicted inference policy output by the first classification head with a preset threshold. If the highest probability is greater than or equal to the preset threshold, the predicted inference policy corresponding to the highest probability is used as the given inference policy; otherwise, the default inference policy is used as the given inference policy. The present invention, by introducing a policy correction module, provides the system with an uncertainty handling mechanism. This mechanism can trigger the system to select a robust default inference policy when the confidence of the predicted inference policy is too low, making the inference policy provided by the system more accurate.

[0044] According to one embodiment of the present invention, the model router further includes: an environment detection module for acquiring real-time environment status, including the current load of each node and the network latency and bandwidth from the end side to each node; and a probability correction module for adjusting the probability corresponding to each predicted node output by the second classification head according to a preset calculation method based on the real-time environment status. The key point of this scheme is that a conditional model router is constructed by introducing the environment detection module and the probability correction module, which is used to ensure compatibility between explicit modeling strategies and model nodes. It generates a probability distribution for routing to each node based on task attributes and the inference strategy selected in the first stage, and then dynamically adjusts the probability distribution of node selection according to the real-time network resource status to adapt to the dynamic changes in the network environment, thereby achieving the selection of the optimal model node under a given inference strategy.

[0045] According to one embodiment of the present invention, the preset calculation method in the probability correction module is as follows:

[0046]

[0047] in, This represents the adjusted probability of the i-th prediction node; This represents the initial probability of the i-th prediction node; The real-time comprehensive resource cost penalty for the i-th prediction node is calculated by weighting the real-time environmental state. represents the sensitivity adjustment coefficient for real-time comprehensive resource cost penalty; j represents traversing all prediction nodes. This invention increases the probability weight of selecting low-latency, low-load nodes by introducing a resource penalty factor. For example, if the cloud network latency exceeds a preset threshold, the probability weight of selecting cloud nodes is reduced, and the probability weight of selecting edge nodes is increased accordingly.

[0048] According to one embodiment of the present invention, a predetermined policy router construction method includes:

[0049] S1. Construct an initial policy router, which includes: a first encoding module for acquiring a question and extracting a first question feature; and a first classification head for providing multiple prediction and inference strategies and the probability corresponding to each prediction and inference strategy based on the first question feature.

[0050] S2. Obtain the policy router training set, which includes multiple training samples. Each training sample includes a question and a corresponding inference policy label.

[0051] S3. Train the initial policy router to convergence using the policy router training set. The training process takes the question as input and the predicted inference policy and the probability corresponding to each policy as output. The parameters of the first encoding module and the first classification head are updated by minimizing the cross-entropy loss between the prediction result and the inference policy label. This cross-entropy loss is expressed as:

[0052]

[0053] in, Let represent the cross-entropy loss, n represent the total number of available nodes in the edge-cloud architecture, and K represent the total number of inference strategies. The label represents the inference strategy label corresponding to question i. The predictive inference strategy of the policy router for problem i The corresponding probabilities. It should be noted that the first encoding module can use encoders such as BERT or RoBERTa; the first classification head can use a linear classification head or an MLP classification head, and its output can be represented as... ,in, It represents the predicted probability corresponding to the i-th inference strategy, and softmax is the normalized exponential activation function. It is the weight matrix of the first classifier head. It is the feature vector of the i-th reasoning strategy in the first classifier head hidden layer. It is the bias vector of the first classifier head.

[0054] According to an embodiment of the present invention, in S2, the inference policy label of the policy router training set is obtained as follows: S21, collecting a question set, an inference policy set, a model category set, and a model deployment node set, wherein the question set includes multiple questions, the inference policy set includes multiple inference policies, the model set includes multiple large models, and the model deployment location set includes different deployment nodes for each large model. S22, traversing and combining each inference policy, each large model, and the nodes deployed by each large model to obtain multiple policy-model-node combinations. S23, for each question, performing inference one by one through each policy-model-node combination to obtain the answer to the question under each policy-model-node combination, and calculating the quality score of each answer according to a preset calculation method. S24, for each question, calculating its aggregate score under different inference policies, and taking the inference policy with the highest aggregate score as the inference policy label corresponding to the question, wherein the aggregate score is calculated as follows:

[0055]

[0056] in, This represents the aggregate score corresponding to reasoning strategy j for a given question. This represents the calculation of the average, weighted average, or median. The quality score of the response given by the strategy-model-node combination JLP to question i is represented by this score.

[0057] According to one embodiment of the present invention, in S23, the preset calculation method is as follows:

[0058]

[0059] in, The quality score of the response given by JLP for question i, representing the strategy-model-node combination. The correctness of the JLP's response to question i represents the combination of strategy, model, and node. The JLP strategy-model-node combination represents the number of lexical units consumed to provide a response to question i. The representative question is about the maximum word consumption across all strategy-model-node combinations. The transmission latency of the response to question i under the strategy-model-node combination JLP. The computation time of JLP, representing the strategy-model-node combination, to provide a response to question i. This represents the computing cost of a large model across different deployment nodes. , , , These represent different weights. In simple terms, the default calculation method is to traverse all policy-model combinations of all nodes to obtain the answer for each combination, and then perform a weighted calculation combining the number of tokens consumed, correctness, network transmission latency, and computational resource cost to obtain the quality score of the answer. This quality score is used to obtain an independent score for the inference strategy through cross-model aggregation, so as to generate inference strategy labels decoupled from specific models. The key point of this scheme is to traverse all policy-model combinations of each node, obtain the accuracy of each combination's answer, token consumption, and node resource status data, and then construct a composite score function that includes network transmission latency, computation time, and resource cost. Next, the comprehensive score of all policy-model combinations is evaluated based on this score function, so that the system's routing decision not only focuses on answer quality but also comprehensively considers the actual time and cost in the edge-cloud environment, making this invention applicable to complex edge-cloud architectures.

[0060] To better understand how the inference strategy labels are obtained, the following example illustrates steps S21-S24. Figure 2 :

[0061] First, collect a representative set of questions. and define a set of reasoning strategies. (like To answer directly For the chain of thought, (Thinking in a tree structure) and the set of available model categories and model node set .

[0062] Then, the nodes deployed by each inference strategy, each large model, and each large model are traversed and combined to obtain multiple strategy-model-node combinations. .

[0063] Next, iterate through all strategy-model-node combinations. For each question Perform inference and obtain the answer text Correctness label and token consumption Simultaneously record transmission delay (depending on the node) and bandwidth ) and computation time (Depending on the model category) and nodes (Computing power); Based on this data, calculate a fine-grained quality score for each strategy-model-node combination regarding the answer text. And store the score data for all combinations (in the format of...) This can be used for subsequent cross-model aggregation.

[0064] Finally, for each question The models are grouped according to their inference strategies to collect scores for each strategy across all model categories and nodes (this process can be represented as...). ), and calculate each problem Aggregate scores (using methods such as average, weighted average, or median) for each inference strategy to eliminate the impact of performance fluctuations in a single model category or node on the evaluation of the inference strategy. The inference strategy with the highest aggregate score is selected as the inference strategy label for that problem.

[0065] Therefore, the key point of steps S21-S24 is that, since the policy router in the first stage needs to select an inference policy without a specified node, this invention generates a policy-independent score decoupled from the model deployment node by performing cross-model aggregation on the problem-policy combination, thereby generating an inference policy label. This method solves the problem that the evaluation of inference policies in traditional joint routing schemes depends on specific models, enabling the system trained based on this inference policy label to have generalization ability when selecting inference policies, ensuring that it can independently select a high-quality inference policy in the first stage.

[0066] According to one embodiment of the present invention, a predetermined model router construction method includes:

[0067] M1. Construct the initial model router, which includes: a second encoding module (which can also be BERT, RoBERTa, etc.), used to acquire the question and extract second question features; multiple second classification heads (which can also be linear classification heads or MLP classification heads), each second classification head corresponding to an inference strategy. When a given inference strategy is given, the second classification head corresponding to the given inference strategy provides multiple prediction nodes and the probability of each prediction node based on the question; and a routing execution module, used to route the question and the given inference strategy to the prediction node with the highest probability so that the node can call the large model on it to perform inference according to the given inference strategy to obtain the answer to the question.

[0068] M2. Obtain the model router training set, which includes multiple training samples. Each training sample includes a question, a given inference policy, and a corresponding node label. The node label can be obtained by recombining all policy-model-node combinations according to the given inference policy to obtain multiple nodes corresponding to the given inference policy. For each question and its corresponding given inference policy, the node with the highest aggregation score is used as the node label.

[0069] M3. The initial model router is trained to convergence using the model router training set. The training process takes the problem and the given inference policy as inputs and multiple prediction nodes and the probability corresponding to each prediction node as outputs. The parameters of the second encoding module and all second classification heads are updated by minimizing the cross-entropy loss between the prediction results and the node labels.

[0070] According to one embodiment of the present invention, the present invention also provides a reasoning method based on the system described herein, the reasoning method comprising: T1, obtaining a question raised by a user; T2, inputting the obtained question into a large-scale model reasoning system based on an edge-cloud collaborative architecture to obtain an answer to the question.

[0071] In summary, the two-stage decomposition joint routing scheme proposed in this invention addresses the problem that in edge-cloud collaborative large-model inference scenarios, when faced with the selection of multiple inference strategies and multiple model nodes, individual model routing or policy routing methods cannot achieve the optimal policy-model combination, resulting in limited inference optimization performance. Existing joint routing methods tightly couple policies and model nodes, causing routing complexity to increase exponentially with the number of both, resulting in an excessively large output space and ignoring network transmission latency and heterogeneity of computing resources, leading to uncontrollable overall latency. The core of this invention lies in decoupling policy selection from model node selection in edge-cloud collaborative scenarios through two-stage decomposition. First, training data is generated through a multi-dimensional weighted inference score calculation module, and the first-stage router is trained using a cross-model aggregation policy router training method. Then, during inference, the edge independently selects the optimal inference strategy first, and then selects the large model node with the best compatibility based on real-time resource status and the selected inference strategy. Compared with existing technologies, this invention significantly reduces routing computational complexity while achieving joint optimization of policies and models, comprehensively considers transmission and computational latency, and improves the response speed and task completion quality of the entire large-model inference system.

[0072] In summary, this invention significantly reduces the complexity of routing computation by decoupling the tightly coupled joint routing into two stages: policy selection and model selection. Furthermore, it considers factors such as network conditions, heterogeneity differences between edge and cloud computing resources, and energy consumption costs of different nodes to dynamically adapt to changes in network and computing resources. While ensuring inference quality, it significantly reduces inference latency and inference costs, effectively improving the overall efficiency and stability of the edge-cloud collaborative inference system, and has broad application prospects.

[0073] It should be noted that although the steps are described in a specific order above, it does not mean that the steps must be executed in the above specific order. In fact, some of these steps can be executed concurrently, or even in a different order, as long as the required function can be achieved.

[0074] This invention can be a computer program product, which mainly refers to a software product that implements this solution through a computer program.

[0075] The various embodiments of the present invention have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical application, or technical improvements to the embodiments in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.

Claims

1. A large-model inference system based on an edge-cloud collaborative architecture, wherein, The system has multiple nodes deployed on both the edge and cloud sides. Each node deploys a large model. The edge side of the system receives user questions, selects inference strategies and nodes for those questions, and routes the questions and inference strategies to the selected nodes so that the nodes can invoke the large model on them to perform inference according to the inference strategy to obtain the answer to the question. The system is characterized by having a two-stage router configured on the edge side, the two-stage router comprising: A policy router, which is obtained according to a predetermined policy router construction method, is used to receive questions raised by users and select an inference policy for those questions; A model router, obtained according to a predetermined model router construction method, is used to receive a user-submitted question and a reasoning strategy selected by the policy router, select a node under the constraints of the reasoning strategy, and route the question and the selected reasoning strategy to the selected node so that the node can invoke the large model on it to perform reasoning according to the reasoning strategy to obtain an answer to the question.

2. The large model inference system based on edge-cloud collaborative architecture according to claim 1, characterized in that: The policy router includes: a first encoding module, used to acquire a question and extract a first question feature; and a first classification head, used to provide multiple prediction and inference strategies and the probability of each prediction and inference strategy based on the first question feature, and to take the prediction and inference strategy with the highest probability as the given inference strategy. The model router includes: a second encoding module for acquiring the question and extracting second question features; multiple second classification heads, each corresponding to a reasoning strategy, wherein, when a reasoning strategy is given, the second classification head corresponding to the given reasoning strategy provides multiple prediction nodes and the probability corresponding to each prediction node based on the question; and a routing execution module for routing the question and the given reasoning strategy to the prediction node with the highest probability, so that the node calls the large model on it to perform reasoning according to the given reasoning strategy to obtain the answer to the question.

3. The large-model inference system based on an edge-cloud collaborative architecture according to claim 2, characterized in that, The policy router further includes a policy correction module, which is configured with a default inference policy and is used to compare the highest probability of the predicted inference policy output by the first classification head with a preset threshold. If the highest probability is greater than or equal to the preset threshold, the predicted inference policy corresponding to the highest probability is used as the given inference policy; otherwise, the default inference policy is used as the given inference policy.

4. The large-model inference system based on an edge-cloud collaborative architecture according to claim 2, characterized in that, The model router also includes: An environment detection module is used to acquire real-time environmental status, which includes the current load of each node and the network latency and network bandwidth from the end side to each node. The probability correction module is used to adjust the probability of each prediction node output by the second classification head according to a preset calculation method based on the real-time environmental conditions.

5. The large-model inference system based on an edge-cloud collaborative architecture according to claim 4, characterized in that, The default calculation method in the probability correction module is: in, This represents the adjusted probability of the i-th prediction node; This represents the initial probability of the i-th prediction node; The real-time comprehensive resource cost penalty for the i-th prediction node is calculated by weighting the real-time environmental state. represents the sensitivity adjustment coefficient for real-time comprehensive resource cost penalty; j represents traversing all prediction nodes.

6. The large-model inference system based on an edge-cloud collaborative architecture according to claim 1, characterized in that, The predetermined policy router construction method includes: S1. Construct an initial policy router, which includes: a first encoding module, used to acquire a question and extract a first question feature; and a first classification head, used to provide multiple prediction and inference policies and the probability corresponding to each prediction and inference policy based on the first question feature. S2. Obtain the policy router training set, which includes multiple training samples. Each training sample includes a question and a corresponding inference policy label. S3. The initial policy router is trained to convergence using the policy router training set. The training process takes the problem as input, the prediction inference policy and the probability corresponding to each prediction inference policy as output, and updates the parameters of the first encoding module and the first classification head by minimizing the cross-entropy loss between the prediction result and the inference policy label.

7. The large model inference system based on an edge-cloud collaborative architecture according to claim 6, characterized in that, In S2, the inference policy labels of the policy router training set are obtained as follows: S21. Collect a problem set, an inference strategy set, a model category set, and a model deployment node set. The problem set includes multiple problems, the inference strategy set includes multiple inference strategies, the model set includes multiple large models, and the model deployment location set includes different deployment nodes for each large model. S22. Traverse and combine each inference strategy, each large model, and each node deployed by the large model to obtain multiple strategy-model-node combinations; S23. For each question, reasoning is performed one by one through each strategy-model-node combination to obtain the answer to the question under each strategy-model-node combination, and the quality score of each answer is calculated according to the preset calculation method. S24. For each question, calculate its aggregate score under different reasoning strategies, and take the reasoning strategy with the highest aggregate score as the reasoning strategy label for that question. The aggregate score is calculated as follows: in, This represents the aggregate score corresponding to reasoning strategy j for a given question. This represents the calculation of the average, weighted average, or median. The quality score of the response given by the strategy-model-node combination JLP to question i is represented by this score.

8. The large-model inference system based on an edge-cloud collaborative architecture according to claim 7, characterized in that, In S23, the preset calculation method is as follows: in, The quality score of the response given by JLP for question i, representing the strategy-model-node combination. The correctness of the JLP's response to question i represents the combination of strategy, model, and node. The JLP strategy-model-node combination represents the number of lexical units consumed to provide a response to question i. The representative question is about the maximum word consumption across all strategy-model-node combinations. The transmission latency of the response to question i under the strategy-model-node combination JLP. The computation time of JLP, representing the strategy-model-node combination, to provide a response to question i. This represents the computing cost of a large model across different deployment nodes. , , , Each represents a different weight.

9. The large-model inference system based on an edge-cloud collaborative architecture according to claim 1, characterized in that, The predetermined model router construction method includes: M1. Construct an initial model router, which includes: a second encoding module for acquiring the question and extracting second question features; multiple second classification heads, each corresponding to a reasoning strategy, wherein, given a reasoning strategy, the second classification head corresponding to the given reasoning strategy provides multiple prediction nodes and the probability of each prediction node based on the question; and a routing execution module for routing the question and the given reasoning strategy to the prediction node with the highest probability, so that the node calls the large model on it to perform reasoning according to the given reasoning strategy to obtain the answer to the question. M2. Obtain the training set of the model router, which includes multiple training samples. Each training sample includes a question, a given inference policy, and the corresponding node label. M3. The initial model router is trained to convergence using the model router training set. The training process takes the problem and the given inference policy as inputs and multiple prediction nodes and the probability corresponding to each prediction node as outputs. The parameters of the second encoding module and all second classification heads are updated by minimizing the cross-entropy loss between the prediction results and the node labels.

10. A reasoning method based on the system according to any one of claims 1-9, characterized in that, The reasoning method includes: T1. Obtain the questions raised by users; T2. Input the obtained question into the large model inference system based on the edge-cloud collaborative architecture to obtain the answer to the question.

11. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method of claim 10.