An ai model routing system and method for handling stateful sessions
Patent Information
- Application Number
- CN202611232205.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-14
- Publication Date
- 2026-09-22
AI Technical Summary
[0004]本申请的目的在于,提供一种用于处理有状态会话的AI模型路由系统及方法,旨在解决现有技术中AI模型路由策略仅依赖技术指标而无法实现成本效益最优化,以及在故障切换时因无法处理会话状态而导致用户体验中断的技术问题
[0016]与现有技术相比,本申请提供的技术方案具有如下有益效果:首先,通过建立模型调用的技术成本与业务结果指标之间的闭环反馈,计算成本效益指标并据此进行路由决策,能够将请求动态地分配给“单位业务结果成本”最低的模型,实现了成本效益驱动的智能路由,从而在商业层面实现价值最大化。其次,通过在模型故障切换时利用上下文适配器转换会话历史记录,能够保持多轮对话等有状态应用的连续性,避免了因模型切换导致的会话中断和重置,极大地提升了用户体验和应用的健壮性。此外,本申请通过统一的多模型接入层,将新模型的接入简化为配置工作,降低了开发和长期维护成本。最后,统一的监控和计量模块为深度的成本效益分析提供了精确的数据基础,使运营决策更加科学、精准,提升了系统的可观测性与管理效率。
Smart Images

Figure CN122802596A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of artificial intelligence platforms and large model management technology, and in particular to an AI model routing system and method for handling stateful sessions. Background Technology
[0002] With the widespread application of large-scale artificial intelligence (AI) models, enterprises often need to access and manage various heterogeneous underlying models when building AI capabilities. To address the high development and maintenance costs resulting from differences in interfaces, parameters, and authentication methods among different models, existing technologies have developed AI model management platforms or gateways. These platforms typically employ an adapter pattern to uniformly encapsulate various underlying heterogeneous models, providing a unified calling interface to upper-layer applications. Simultaneously, these platforms also possess basic monitoring and fault tolerance capabilities; for example, they can track model call volume and automatically switch requests to a pre-configured backup model when a model service failure is detected.
[0003] However, existing model routing and failover mechanisms still have shortcomings. On the one hand, routing decisions mainly rely on the platform's own technical status (such as model load and availability) or static business tags carried in the request. This routing strategy is relatively rigid and cannot directly link the cost of calling the model with the actual business results generated by the call for comprehensive consideration, thus making it difficult to achieve true cost-effectiveness optimization. On the other hand, for stateful application scenarios such as multi-turn dialogues and programming assistants that require maintaining session history, existing platforms are usually stateless when performing failover, that is, simply sending new requests to the backup model. This leads to the loss of the context history of previous dialogues, causing session interruption and seriously affecting the user experience. Summary of the Invention
[0004] The purpose of this application is to provide an AI model routing system and method for handling stateful sessions, aiming to solve the technical problems in the prior art where AI model routing strategies rely solely on technical indicators and cannot achieve cost-effectiveness optimization, and where user experience is interrupted due to the inability to handle session state during failover.
[0005] To achieve the above objectives, this application provides an AI model routing system for processing stateful sessions, comprising: a model selection strategy engine, configured to acquire computational resource consumption and terminal task execution status signals associated with historical model calls, calculate and update the resource conversion efficiency index of each base model by dividing the computational resource consumption by the number of successfully executed execution status signals extracted, and select a target base model and initiate a call to it based on the resource conversion efficiency index when a call request associated with a stateful session is received; a context manager, configured to update the call request and the response received from the target base model to the context history of the stateful session after a successful call to the target base model; and an automatic failover mechanism, configured to record cost information associated with the failed call and trigger a switch to a backup base model when a call to the target base model fails; wherein the automatic failover mechanism includes a context adapter, configured to read the context history of the stateful session when the switch occurs and convert the context history from the format of the target base model to a format compatible with the backup base model.
[0006] The resource conversion efficiency index is a cost-effectiveness indicator used to characterize the computational resources or technical costs consumed by the basic model to obtain a unit of successful business result. Specifically, the resource conversion efficiency index can be expressed as: Resource Conversion Efficiency Index = Computational Resource Consumption in the Statistical Period / Number of Successful Execution Status Signals in the Statistical Period. The smaller the value of this index, the stronger the basic model's ability to obtain a successful business result with fewer resources. The unit business result cost is a monetized representation of the resource conversion efficiency index, which can be obtained by dividing the technical cost in the statistical period by the business result index.
[0007] Optionally, the model selection strategy engine is further configured to: when selecting the target base model based on the cost-effectiveness index, further combine the semantic features of the call request and the real-time status information of each base model, wherein the semantic features include intent tags or keywords extracted by natural language processing of the call request, and the real-time status information includes at least one of model latency, load, or error rate, and the combination refers to using forward normalization for high-value priority indicators such as semantic matching degree and model capability level during numerical normalization processing; and using reverse normalization for low-value priority indicators such as model latency, model load, error rate, unit business result cost, or resource conversion efficiency, so that the lower the original value of the low-value priority indicator, the higher its normalization score. The model selection strategy engine calculates the comprehensive score as follows: Comprehensive Score = w1 × Semantic Matching Score + w2 × Resource Conversion Efficiency Reverse Normalization Score + w3 × Latency Reverse Normalization Score + w4 × Load Reverse Normalization Score + w5 × Error Rate Reverse Normalization Score, where w1 to w5 are pre-configured weight coefficients, and the sum of each weight coefficient is 1. The model selection strategy engine selects the base model with the highest comprehensive score that is in an available state as the target base model.
[0008] Optionally, the cost-benefit indicator is the cost per unit of business outcome; the model selection strategy engine is specifically used to calculate the cost per unit of business outcome by dividing the technology cost by the business outcome indicator.
[0009] Optionally, the computational resource consumption includes the number of tokens consumed for model calls; the terminal task execution status signals include positive status signals and negative status signals. The positive status signals include at least one of the following: a code adoption operation trigger signal returned by the terminal application, a task completion confirmation signal, and a positive review signal; the negative status signals include at least one of the following: a code abandonment signal, a negative review signal, a continuous interaction round exceeding the limit disconnection signal, and a user-initiated session reset signal. The number of successful execution status signals only counts the number of positive status signals; the negative status signals are used to correct business result indicators, trigger demotion processing, or serve as model quality evaluation data.
[0010] Optionally, the automatic fault switching mechanism is used to: trigger a switch to the backup basic model when the number of consecutive timeouts for the target basic model reaches a preset threshold or the error rate per unit time exceeds a preset threshold.
[0011] Optionally, the context adapter is specifically used to: convert the format of the target base model, i.e., a context format based on preset tags to define the speaker, into a format compatible with the backup base model, i.e., a context format based on JavaScript Object Notation (JSON) objects.
[0012] This application also provides an AI model routing method for handling stateful sessions, comprising the following steps: obtaining computational resource consumption and terminal task execution status signals associated with historical model calls; identifying positive status signals from the terminal task execution status signals; and determining the number of successful execution status signals based on the positive status signals; calculating and updating the resource conversion efficiency index of each basic model by dividing the computational resource consumption by the number of successful execution status signals; and when the number of successful execution status signals is 0... When the resource conversion efficiency index of the corresponding basic model is set to a preset penalty value or delayed to the next statistical period for updating; a call request associated with a stateful session is received, and the context history of the stateful session is obtained; a target basic model is selected according to the resource conversion efficiency index; the call request is sent to the target basic model, and it is determined whether the call is successful; if the call is successful, the call request and the response received from the target basic model are updated in the context history of the stateful session; if the call fails, the cost information associated with the failed call is recorded, and a switch to a backup basic model is triggered. The switching steps include: reading the context history of the stateful session, extracting the corresponding speaker role identifier and text content by parsing the structured preset tags that define the speaker in the current data format, and mapping them to generate a JSON object array format containing role and content fields that is compatible with the backup basic model.
[0013] Optionally, the computational resource consumption includes the number of tokens consumed for model calls; the terminal task execution status signals include positive status signals and negative status signals. The positive status signals include at least one of the following: a code adoption operation trigger signal returned by the terminal application, a task completion confirmation signal, and a positive review signal; the negative status signals include at least one of the following: a code abandonment signal, a negative review signal, a continuous interaction round exceeding the limit disconnection signal, and a user-initiated session reset signal. The number of successful execution status signals only counts the number of positive status signals; the negative status signals are used to correct business result indicators, trigger demotion processing, or serve as model quality evaluation data.
[0014] Optionally, the step of triggering the switch to a backup base model specifically involves: performing the switch when the number of consecutive timeouts for the target base model reaches a preset threshold or the error rate per unit time exceeds a preset threshold.
[0015] This application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method as described in any of the preceding claims.
[0016] Compared with existing technologies, the technical solution provided in this application has the following beneficial effects: First, by establishing a closed-loop feedback between the technical cost of model calls and business result indicators, cost-benefit indicators are calculated and routing decisions are made accordingly. This allows requests to be dynamically allocated to the model with the lowest "unit business result cost," achieving cost-benefit-driven intelligent routing and maximizing value at the business level. Second, by utilizing a context adapter to transform session history during model failure switching, the continuity of stateful applications such as multi-turn dialogues can be maintained, avoiding session interruptions and resets caused by model switching, greatly improving user experience and application robustness. Furthermore, this application simplifies the access of new models to a configuration task through a unified multi-model access layer, reducing development and long-term maintenance costs. Finally, a unified monitoring and metering module provides an accurate data foundation for in-depth cost-benefit analysis, making operational decisions more scientific and precise, and improving system observability and management efficiency. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 A flowchart illustrating an AI model routing method provided in an embodiment of this application; Figure 2 This is a schematic diagram of a cost-effective closed-loop routing structure provided in an embodiment of this application.
[0019] Figure Label Explanation: S101 - Receive call request; S102 - Perform request pre-analysis; S103 - Select model based on cost-benefit indicators; S104 - Call the selected model through the adapter; S105 - Determine if the call was successful; S106 - If it fails, call the context adapter and trigger failover; S107 - If successful, record technical indicators; S108 - Receive business result indicators and update the cost-benefit database; S109 - Return results; 301 - Call request; 310 - Model selection strategy engine; 320 - Model A; 330 - Technical cost; 340 - Business result; 350 - Cost-benefit analysis module; 360 - Updated cost-benefit indicators. Detailed Implementation
[0020] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0021] This application provides an AI model routing system and method for handling stateful sessions. Specifically, the system acts as an intermediate platform, providing a unified and standardized model invocation interface to upper-layer applications, while simultaneously enabling unified access and management of various heterogeneous underlying models. The core function of this system is that it can not only route based on the real-time technical status of the models, but also calculate the cost-benefit index of each model by combining the technical cost of model invocation with the final business result, and use this as the core basis for intelligent routing decisions. Furthermore, it should be noted that for stateful application scenarios that require maintaining dialogue history, the system provided in this application can automatically convert the session context format when a model failure is detected and a switch is performed, thereby ensuring session continuity and avoiding user experience interruption.
[0022] Please see Figure 1 This is a flowchart illustrating an AI model routing method provided in an embodiment of this application. The method can be executed by the AI model routing system provided in this embodiment, which can be deployed in a server or cloud computing environment. Figure 1 As shown, the method may include: In step S101, a call request is received. The system receives an AI model call request that conforms to a unified interface specification from the upper-layer business application. This request may contain content that needs to be processed (e.g., the user's question, the code snippet to be generated), as well as metadata related to the request (e.g., session identifier, business scenario tag).
[0023] In step S102, request pre-analysis is performed. As an optional implementation, the system's internal request pre-analysis module can perform real-time natural language processing on the received call request content to extract its semantic features. These semantic features may include the request's intent type (e.g., question answering, translation, code generation), complexity (e.g., simple, medium, complex), and its professional field (e.g., law, medicine, finance), which can serve as one of the inputs for subsequent model selection.
[0024] In step S103, a model is selected based on cost-effectiveness indicators. The system's model selection strategy engine selects a target base model according to a preset routing strategy. Specifically, the model selection strategy engine first obtains the resource conversion efficiency indicators, real-time status information, and semantic matching scores corresponding to the current call request for each candidate base model within a preset statistical period. The resource conversion efficiency indicators can be calculated as "calculated resource consumption / number of successful execution status signals"; when the number of successful execution status signals is 0, the model selection strategy engine sets the resource conversion efficiency indicator of the candidate base model to a preset penalty value, or temporarily does not use the candidate base model in cost-effectiveness priority routing. Subsequently, forward normalization is applied to the semantic matching score, and backward normalization is applied to the model latency, model load, error rate, and resource conversion efficiency indicators, so that the lower the latency, the lower the load, the lower the error rate, and the lower the resource consumption per unit of successful business result, the higher the corresponding score. Finally, the comprehensive score of each candidate base model is calculated according to preset weights, and the base model with the highest comprehensive score and not in a circuit breaker state is selected as the target base model.
[0025] In step S104, the selected model is invoked via an adapter. The system uses an adapter component corresponding to the selected target base model to convert the internally unified format of the invocation request into the application programming interface format specific to that target model, and then sends the invocation to it.
[0026] In step S105, it is determined whether the call was successful. The system monitors the call process to the target basic model. If valid response data is successfully received, the call is considered successful; if a timeout occurs, an error code is returned, or the response content does not conform to the expected format, the call is considered to have failed.
[0027] In step S106, if the call fails, the context adapter is invoked and a failover is triggered. Specifically, when a call is determined to have failed, the automatic failover mechanism is triggered. Accordingly, the context adapter in this mechanism starts, reads the context history of the current session, and converts it from the format of the failed target model to a format compatible with a preset backup model. Subsequently, the system selects the backup model and re-initiates the call using the converted context and the original request (i.e., returns to step S104). Simultaneously, the system records cost information related to this failed call (e.g., call duration or a small amount of token consumption that may occur) to ensure the accuracy of the cost-benefit analysis.
[0028] In step S107, if successful, the technical metrics are recorded. Once a successful call is determined, the system's unified observability and measurement module will record detailed technical metrics logs for this call. These logs may include call time, model used, response time, and the number of tokens consumed in the request and response.
[0029] In step S108, business result metrics are received and the cost-effectiveness database is updated. The system receives the terminal task execution status signal associated with this successful call from the upper-layer business application via a specific application programming interface (API). To handle asynchronous feedback delays caused by manual code review, this system introduces an asynchronous feature matching queue mechanism: when a model call occurs, the system generates a globally unique session request ID for this interaction and sends it to the terminal along with the response; the system maintains a set time window (e.g., 24 hours) in the background. When a terminal task execution status signal carrying the session request ID is received via the API (e.g., a trigger signal generated by a developer clicking the "Accept" button), the corresponding signal is extracted from the asynchronous feature matching queue and precisely bound to the computational resource consumption log recorded in step S107, thereby triggering real-time update calculation of resource conversion efficiency metrics; if no matching trigger signal is received within the time window, the task execution status of this call is recorded as invalid. This asynchronous binding mechanism ensures that the underlying routing system can maintain data consistency in closed-loop computation even when facing uncertain delays in upper-layer business processes.
[0030] In one specific implementation, each record in the asynchronous feature matching queue includes at least: session request ID, session identifier, base model identifier, call request digest or hash value, request time, number of input tokens, number of output tokens, call status, feedback status, feedback time, expiration time, and cost information.
[0031] When the terminal application returns a terminal task execution status signal, the system queries the asynchronous feature matching queue based on the session request ID. If a record that has not expired and is not bound to feedback is found, the terminal task execution status signal is bound to the corresponding model call record, and the business result indicators of the corresponding basic model are updated according to the type of the signal. If the terminal task execution status signal is a positive status signal, the number of successful execution status signals is increased; if it is a negative status signal, the number of failure or low satisfaction status signals is increased, or the score of the model under the corresponding task type is downgraded. If no matching record is found, or the record has exceeded the preset time window, the feedback signal is marked as invalid feedback or late feedback and is not included in the calculation of resource conversion efficiency indicators in this round.
[0032] By using the above fields and matching rules, the system can bind technical costs to business results one by one, even when there is a time delay between model invocation and terminal feedback, thereby ensuring the data consistency of resource conversion efficiency index calculation.
[0033] In step S109, the result is returned. The system returns the response data received from the model and converted into a unified format to the upper-layer business application.
[0034] Through the above process, this application embodiment constructs a closed-loop system capable of dynamic learning and optimization. In this system, the selection of AI models no longer relies solely on static configurations or single technical indicators, but is tightly coupled with the ultimate business value, while ensuring high service availability and continuous experience for stateful sessions.
[0035] Example 1 In one embodiment of this application, the AI model routing system provided herein will be described in detail as achieving cost-effective intelligent routing. This embodiment corresponds to a specific application scenario: a software company wants to evaluate and select a better large language model for its AI code assistant product through comparative testing, and then dynamically route it based on cost-effectiveness.
[0036] Please see Figure 2 This is a schematic diagram of a cost-effective closed-loop routing provided in an embodiment of this application. The entire process involves several key components of the system, including a model selection strategy engine 310, a unified observability and measurement module (not shown separately, its function is reflected in the collection of technical costs 330), and a core cost-effectiveness analysis module 350.
[0037] In this embodiment, the system administrator first configures two selectable base models in the system's multi-model access layer: baseline model I and a new model H to be evaluated. Model I is considered to be the currently stable online model, and its cost and performance are known benchmarks; model H is a newer model that is claimed to be more capable but also more expensive.
[0038] To conduct a scientific comparative evaluation, the operations team configured a canary release routing strategy in the model selection strategy engine 310. This strategy stipulates that 95% of call requests (301) will be routed to the baseline model I, while the remaining 5% will be routed to the new model H. This traffic splitting allows for the collection of performance data for the new model H in real-world usage scenarios without impacting the majority of users.
[0039] When the system starts running according to this strategy, for each call request 301 (e.g., a developer requesting "implement a level-order traversal of a binary tree using Python"), the model selection strategy engine 310 will, according to the gray-scale rule, route the request to model H with a 5% probability and to model I with a 95% probability. Assuming this request is routed to model A 320, where model A refers to the selected model, which can be either model H or model I.
[0040] The system sends the request to model A and receives the generated code through the adapter corresponding to model A. After the call is completed, two types of key data are generated. The first type is the technical cost (330), which is automatically captured by the unified observability and metering module. In this embodiment, the technical cost can be precisely quantified as the fee charged by the model vendor for this call, which is typically directly related to the number of tokens consumed for input and output. This module records the detailed cost of each call; for example, the average cost of calling model I is $0.01, and the average cost of calling model H is $0.05.
[0041] The second category is business outcome 340. Understandably, quantifying and utilizing business outcomes is a crucial aspect of this application. The AI code assistant's front-end application has been modified to add a feedback mechanism. When the code generated by the model is presented to the developer, the developer can choose to "adopt" or "discard" the code. This action (adoption or rejection) serves as a business outcome metric, which is sent back to the AI model routing system of this application through a dedicated application programming interface. In this embodiment, this business outcome metric is defined as "code adoption rate."
[0042] Both the technical cost (330) and business outcome (340) data are simultaneously transmitted to the cost-benefit analysis module (350). The core function of this module is to calculate and update the cost-benefit index for each model. In this embodiment, this index is specifically defined as "unit business outcome cost," which is calculated by dividing the average technical cost of a model over a period of time by its corresponding average business outcome index.
[0043] Assume that after one week of system operation, the unified observability and measurement module and business results receiving module collect the following statistics: For baseline model I: average call cost is $0.01 / call, code adoption rate is 60%. For the new model H: average call cost is $0.05 / call, code adoption rate is 90%.
[0044] At this point, the cost-benefit analysis module 350 performs the following calculations: Cost of unit business outcome for Model I = $0.01 / 60% ≈ $0.0167 per adoption. Cost of unit business outcome for Model H = $0.05 / 90% ≈ $0.0556 per adoption.
[0045] The calculated result is the updated cost-benefit metric 360. This metric shows that although the cost per call to the new model H is five times that of the old model I, its cost of achieving a single effective adoption (i.e., the cost per unit of business outcome) is also higher. This data provides the operations team with decision-making support beyond purely technical metrics (such as latency and initial costs), enabling them to determine whether paying a higher unit adoption cost to achieve a higher code adoption rate (from 60% to 90%) aligns with their business objectives.
[0046] Furthermore, these updated cost-effectiveness metrics 360, as shown by the dotted arrows in the diagram, are fed back to the model selection strategy engine 310. In subsequent routine operations, operators can configure more complex routing strategies, such as a "cost-effectiveness first" strategy. Under this strategy, when a new request arrives, the model selection strategy engine 310 will prioritize the model with the lower "cost per unit business result." For example, for some simple auxiliary code generation tasks with low adoption rate requirements, the system might choose model I to save costs; while for core, complex, and high-value code generation tasks, the system might determine that even if the cost per unit adoption is higher, model H, with a higher adoption rate, should be prioritized to ensure final development efficiency and quality.
[0047] Through this embodiment, the technical solution of this application demonstrates how to construct a data-driven intelligent routing closed loop that is deeply integrated with business value. It not only provides a scientific A / B testing framework for model selection but also continuously optimizes the allocation of model resources during daily operation to achieve overall cost-effectiveness optimization.
[0048] Example 2 This embodiment will focus on illustrating how the AI model routing system provided in this application achieves a seamless user experience when handling stateful sessions through an automatic failover mechanism and a context adapter. This embodiment corresponds to an AI online tutoring application that answers math problems through multi-turn dialogues with students.
[0049] In this embodiment, the core components of the system include a context manager, an automatic failover mechanism, and a key sub-module embedded in the mechanism—the context adapter.
[0050] First, the system is configured with two available language models: the target base model F and the backup base model G. These two models may come from different vendors, so their requirements for the format of the session history (i.e., context) are different. For example, model F requires the context to be a plain text string, where the user's and AI's statements are delimited by specific tags; while model G requires the context to be an array of JSON objects, each of which explicitly identifies the speaking role ("role") and content ("content").
[0051] When a student begins using the AI tutoring app for consultation, a stateful session with a unique identifier is created. The system's context manager is responsible for maintaining the context history of that session.
[0052] The conversation proceeds as follows: In the first round, the student asks, "Please explain the Pythagorean theorem." The system routes this request to the target base model F. Model F successfully responds: "The Pythagorean theorem states that in a right triangle, the sum of the squares of the two legs is equal to the square of the hypotenuse, i.e., a² + b² = c²." After a successful call, the context manager records this interaction (the student's request and the model's response) in the context history associated with the session. At this point, the history can be stored in an internal standard format, or directly in the format of the target base model F. For example, it could be stored as: "[INST]Please explain the Pythagorean theorem.[ / INST] The Pythagorean theorem states that in a right triangle, the sum of the squares of the two legs is equal to the square of the hypotenuse, i.e., a² + b² = c²." In the second to fifth rounds, students continued to ask follow-up questions about theorem proofs and application examples, which the AI (Model F) answered one by one. Each successful interaction was added to the conversation's context history by the context manager, making the history increasingly longer and containing the complete context of the conversation.
[0053] In the sixth round of interaction, the student posed a new question: "If the two legs of a right triangle are 5 and 12, what is the length of the hypotenuse?" Following the established procedure, the system sent a request containing the complete history of the previous five rounds of dialogue and this new question to the target base model F. However, this call failed. The failure could be due to temporary service overload of model F causing consecutive timeouts, or its application programming interface returning an internal server error.
[0054] At this point, the automatic failover mechanism is activated. This mechanism first detects a failed call to the target base model F. For example, this mechanism can be configured to trigger failover when "the number of consecutive call timeouts reaches 3" or "the error rate per unit time (e.g., 1 minute) exceeds 20%".
[0055] Once a switchover is triggered, the mechanism determines that the request needs to be forwarded to the backup base model G. Since this is a stateful session, if a new question is sent directly to model G, model G will be unable to answer correctly due to the lack of prior dialogue information. Therefore, the context adapter intervenes before forwarding the request.
[0056] The core task of the context adapter is to perform format conversion, which involves the following steps: 1. Read the complete context history associated with the current session, maintained by the context manager. In this example, this is the text in Model F format containing the first five rounds of dialogue.
[0057] 2. Parse the history. The adapter is programmed to recognize the specific format of model F (i.e., the [INST] and [ / INST] tags) and traverse the text, recognizing the content between [INST] and [ / INST] as the user's speech, and the content after [ / INST] and before the next [INST] as the AI assistant's speech.
[0058] 3. Convert the parsed content into the format required by the backup model G. For each round of dialogue, the adapter generates a JSON object. For example, the first round of dialogue would be converted to: { "role": "user", "content": "Please explain the Pythagorean theorem."} { "role": "assistant", "content": "The Pythagorean theorem states that in a right triangle, the sum of the squares of the two legs is equal to the square of the hypotenuse, i.e., a² + b² = c²."}.
[0059] 4. The adapter performs this transformation on all five rounds of dialogue, ultimately generating an array containing 10 JSON objects.
[0060] After the format conversion is complete, the automatic fault switching mechanism encapsulates the converted, model G-compatible context history array, along with the new question raised by the student in the sixth round ("If the two legs of a right triangle are 5 and 12, what is the length of the hypotenuse?"), into a new request that conforms to the model G application interface specification, and sends it to model G.
[0061] Upon receiving the request, the backup base model G, having obtained a complete and correctly formatted dialogue history, is able to understand the background of the question and successfully perform calculations and responses: "According to the Pythagorean theorem, the length of the hypotenuse c satisfies c² = 5² + 12² = 25 + 144 = 169, therefore c = 13, and the length of the hypotenuse is 13." This response is returned to the front-end application by the system and presented to the student. From the student's perspective, although the underlying service model has changed, their question received a completely contextual and correct answer, and the entire dialogue process was smooth and natural without any sense of interruption.
[0062] Meanwhile, in the background, the unified observability and measurement module records the failed call attempts (including their cost information, such as the time consumed) and successful switch events, providing data for subsequent system operation and stability analysis. The context manager also updates the sixth round of interaction (the student's request and the model G's response) to the session history after the call succeeds, preparing for possible subsequent interactions.
[0063] Through this embodiment, the technical solution of this application demonstrates its ability to ensure high availability and continuity of experience for stateful applications. By introducing the key innovation of a context adapter in the failover process, it solves the pain point of session interruption caused by model switching in the prior art, thereby greatly improving the robustness of the application and user satisfaction.
[0064] Example 3 This embodiment demonstrates a more complex combined application scenario, aiming to illustrate how the two core mechanisms of cost-effective routing and seamless switching of stateful sessions in the AI model routing system provided in this application work together to provide comprehensive support for an AI programming assistant that provides multi-turn dialogue interaction.
[0065] This AI programming assistant is integrated into the developer's integrated development environment (IDE), allowing developers to request code generation, interpretation, optimization, and debugging via dialogue. The system backend connects to multiple code generation models, including "CodeGen-A" and "CodeGen-B," each with its own strengths in terms of cost, generation speed, code quality, and ability to follow instructions.
[0066] The workflow is as follows: 1. Initial Request and Cost-Effective Routing: A developer starts a new project and sends an initial request to the AI programming assistant: "Please implement a quicksort algorithm in Python and add comments for key steps." This request first reaches the AI model routing system of this application. The request pre-analysis module in the system analyzes the request text and extracts semantic features, such as: {"Task Type": "Code Generation", "Language": "Python", "Algorithm": "Quicksort", "Complexity": "Medium", "Additional Requirements": "Comments"}. These semantic features can be used for more refined routing decisions. Next, the model selection strategy engine 310 begins the decision-making process. At this point, the engine not only considers the semantic features of the request, but more importantly, it queries the latest cost-effectiveness metrics of each model maintained by the cost-effectiveness analysis module 350. Assuming that based on historical data (calculated by continuously collecting technology costs and code adoption rates as described in Example 1), the current "cost per unit of adoption" for model "CodeGen-A" is $0.08, while that for model "CodeGen-B" is $0.12. Based on the configured "cost-effectiveness first" strategy and the "medium complexity" characteristic of the request, the model selection strategy engine 310 determines that "CodeGen-A" is the optimal choice. Therefore, the request is routed to "CodeGen-A".
[0067] 2. Stateful Continuous Interaction: "CodeGen-A" successfully generated the code for quicksort. The developer reviewed it, expressed satisfaction, and adopted the code. This "adoption" event was fed back to the system as a business outcome (340), further solidifying "CodeGen-A's" cost-effectiveness metrics. Subsequently, the developer continued to make requests based on the generated code, forming a stateful session. Developer: "Great. Now, please modify this function to sort in-place to reduce space complexity." Developer: "The code looks good. Please write three unit test cases for this quicksort function, including a base case, a random array case, and a sorted array case." After each successful interaction, the context manager appends the request and "CodeGen-A's" response to the current session's context history. This history is maintained in a specific format required by "CodeGen-A" (e.g., a format using XML tags).
[0068] 3. Seamless switching between failures and stateful operation: When the developer makes the next request: "Now, please analyze the time complexity of this quicksort algorithm in the worst case and explain why," the system will still send a request containing the complete dialogue history and the new question to "CodeGen-A" as usual.
[0069] However, at this point, the calls to "CodeGen-A" failed five times in a row due to timeouts. The automatic failover mechanism detected this situation because one of its internal circuit breaker rules is "the number of consecutive timeouts reaches a preset threshold of 5".
[0070] The circuit breaker mechanism is triggered, temporarily marking "CodeGen-A" as unavailable and initiating a switch to the backup base model "CodeGen-B".
[0071] The context adapter starts immediately. It reads the complete conversation history containing all previous conversations about quicksort code generation, modification, and test case writing. This history is in the proprietary XML format of "CodeGen-A". The adapter parses it and converts it into a JSON object array format compatible with the alternative base model "CodeGen-B".
[0072] The transformed context, along with the developer's new questions about time complexity analysis, was sent to CodeGen-B.
[0073] 4. Closed-loop feedback and system adaptation: "CodeGen-B" understood the context of the entire programming task based on a complete and adapted context, and successfully generated a detailed analysis of the worst-case time complexity of quicksort. The developers were satisfied with this answer and gave it a positive review.
[0074] In the background, this complex processing flow also formed a closed data loop: The unified observability and measurement module recorded five failed calls to "CodeGen-A" and associated cost information, such as the waiting time before each timeout or potential network transmission costs. These failure costs are also included in the overall technical cost of "CodeGen-A," which may increase its "cost per adoption" in subsequent cost-benefit analyses, thus reducing its competitiveness.
[0075] This module also records successful calls to "CodeGen-B", including its technical costs of 330, such as token consumption.
[0076] The positive feedback from the developers, as a business outcome 340, is fed back to the cost-benefit analysis module 350. This module will use the technical cost and business outcome of this successful call to update or initialize the cost-benefit metrics 360 for "CodeGen-B" in tasks such as "code interpretation".
[0077] Through this combined embodiment, the technical solution of this application demonstrates its comprehensive capabilities in real, dynamic, and complex AI application scenarios. It can not only make the optimal model selection based on cost-effectiveness at the start of a session, but also ensure uninterrupted user experience through stateful and seamless switching in the face of sudden model failures during the session. Furthermore, the costs and outputs of all successful and unsuccessful attempts throughout the process are quantified and fed back to the routing decision system, enabling the system to continuously learn and adaptively optimize.
[0078] Example 4 This embodiment aims to detail the fundamental capabilities supporting the AI model routing system of this application. These capabilities are the cornerstone upon which advanced functions such as cost-effective routing and stateful switching in the aforementioned embodiments are based. This embodiment will describe the system's multi-model access, label-based basic routing strategy, and basic high-availability guarantee mechanism.
[0079] In a typical enterprise environment, IT administrators need to use the system in this application to uniformly manage and schedule various AI model resources.
[0080] 1. Multi-model access and configuration: The system administrator first configures three different types of basic models in the multi-model access layer through the system's management interface: Public Cloud Model A: An industry-leading, powerful general-purpose language model, but its calling cost is also relatively the highest.
[0081] Public Cloud Model B: A model known for its high cost-effectiveness, performing well in handling routine tasks at a moderate cost.
[0082] The C model for private deployment: an open-source model deployed on an internal enterprise server. Its performance may not be as good as the public cloud model, but its advantage is that it can ensure that the data does not leave the enterprise's internal network, meeting the highest data security and compliance requirements.
[0083] For each model to be integrated, the administrator needs to provide the access credentials (such as an API key), the interface address, and select or develop a corresponding adapter. The adapter's role is to convert the system's unified request format into the model-specific format. For example, model A's request body might be in XML format, model B's in JSON format, and model C's with its own custom parameter names. The adapter handles these differences, allowing upper-layer applications to ignore the underlying complexity.
[0084] 2. Configure Tag-Based Routing Policies: After model integration, administrators configure different routing policies for different business scenarios in the model selection policy engine. This policy avoids complex cost-benefit calculations and instead uses deterministic routing based on static tags carried in the request, providing a simple and efficient management method for many common business operations.
[0085] Configure a "cost-first" strategy for the "intelligent customer service" scenario: The administrator creates a routing rule stipulating that all calls originating from the intelligent customer service application and carrying the tag "scene": "customer_service" in the request metadata should follow the "cost-first" strategy. Under this strategy, the target base model is designated as the most cost-effective public cloud model B, while the backup base model is designated as the more capable but more expensive public cloud model A. This configuration aims to maximize cost savings while ensuring basic service availability.
[0086] Configure a "Privacy First" strategy for the "Product Review Analysis" scenario: For analysis tasks that require processing product review data containing user privacy, data security is the primary consideration. Therefore, the administrator creates another rule stipulating that requests carrying the tag "scene": "review_analysis" should follow the "Privacy First" strategy. Under this strategy, the target base model is uniquely designated as a privately deployed C model, and no public cloud model is set up as a backup base model to eliminate the risk of any data accidentally flowing to the public cloud.
[0087] 3. Configure circuit breaker and failover rules: To ensure high availability of the overall service, the administrator configured unified health checks and circuit breaker rules for all connected models. Specific trigger thresholds can be set in the automatic failover mechanism. For example, the administrator can configure the following rules: Timeout rule: If the response time of 5 consecutive calls to a model exceeds 10 seconds, the circuit breaker will be triggered.
[0088] Error rate rule: If a model's call error rate (e.g., the proportion of 5xx server error codes returned) exceeds 20% within a 1-minute window, a circuit breaker is triggered. These specific, configurable thresholds provide a clear basis for the system to automatically determine whether a model is "healthy".
[0089] 4. Example of system operation process: Scenario 1: Intelligent Customer Service Invocation: When a request from an intelligent customer service application arrives at the system and carries the tag "scene": "customer_service", the model selection strategy engine matches the "cost priority" strategy, and therefore routes the request to the target base model—Public Cloud Model B. Model B responds normally, and the invocation is completed.
[0090] Scenario 2: Model Failure and Automatic Switchover: One day, the provider of public cloud model B experienced a regional service failure, causing a sharp increase in the response latency of its application programming interface (API). As intelligent customer service requests continued to be sent to model B, the automatic failover mechanism's health detector detected consecutive call timeouts. When the number of consecutive timeouts reached a preset five, the circuit breaker rule was triggered. The system immediately marked model B as "unavailable" and stopped sending any requests to it during the next cooling-off period (e.g., 5 minutes). Simultaneously, because the routing strategy was configured with a backup base model A, the system automatically and seamlessly switched all subsequent new customer service requests to public cloud model A. Although the call cost temporarily increased, the intelligent customer service service itself was not interrupted, ensuring business continuity. Correspondingly, the system can also send alarm notifications to the operations team through the integrated alarm module, informing them of the model B failure.
[0091] Scenario 3: Data and Metering: Whether it's a normal call or a call during a failover, the unified observability and metering module continuously operates in the background. It tags each call with a business line label (such as "intelligent customer service" or "comment analysis") and records the specific model used (A, B, or C) and the costs incurred (token consumption). At the end of the month, the finance department can clearly see the specific expenses of each business line on each model through the reports generated by this module. This precise and categorized technical cost data is the raw data source necessary for the cost-benefit analysis in Example 1.
[0092] This embodiment provides a stable, reliable, and manageable underlying platform support for the implementation of advanced functions through a clearly defined fault-switching mechanism.
[0093] The above description is merely a preferred embodiment of this application and is not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A routing system for AI models to handle stateful sessions, characterized in that, include: The model selection strategy engine is used to obtain the computational resource consumption and terminal task execution status signals associated with historical model calls. By dividing the computational resource consumption by the number of successfully executed status signals extracted, the resource conversion efficiency index of each basic model is calculated and updated. When a call request associated with a stateful session is received, a target basic model is selected according to the resource conversion efficiency index and a call is initiated to it. A context manager is used to update the call request and the response received from the target base model to the context history associated with the stateful session after a successful call to the target base model. An automatic failover mechanism is used to record the cost information associated with the failed call when a call to the target base model fails, and to trigger a switch to a backup base model. The automatic failover mechanism includes a context adapter, which reads the context history of the stateful session when the failover occurs and converts the context history from the data format corresponding to the target base model to a data format compatible with the backup base model.
2. The system according to claim 1, characterized in that, The model selection strategy engine is also used for: When selecting the target base model based on cost-effectiveness indicators, the semantic features of the call request and the real-time status information of each base model are further combined. The semantic features include intent tags or keywords extracted through natural language processing of the call request. The real-time status information includes at least one of model latency, load, or error rate. This combination refers to the following approach during numerical normalization: for high-value priority indicators such as semantic matching degree and model capability level, forward normalization is used; for low-value priority indicators such as model latency, model load, error rate, unit business result cost, or resource conversion efficiency, reverse normalization is used, so that the lower the original value of the low-value priority indicator, the higher its normalization score. The model selection strategy engine calculates the comprehensive score as follows: Comprehensive Score = w1 × Semantic Matching Score + w2 × Resource Conversion Efficiency Reverse Normalization Score + w3 × Latency Reverse Normalization Score + w4 × Load Reverse Normalization Score + w5 × Error Rate Reverse Normalization Score, where w1 to w5... The system is pre-configured with weight coefficients, and the sum of all weight coefficients is 1; the model selection strategy engine selects the base model with the highest comprehensive score and in an available state as the target base model.
3. The system according to claim 1, characterized in that, The cost-benefit indicator is the cost per unit of business outcome; the model selection strategy engine is specifically used to calculate the cost per unit of business outcome by dividing the technology cost by the business outcome indicator.
4. The system according to claim 3, characterized in that, The computational resource consumption includes the number of tokens consumed for model calls; the terminal task execution status signals include positive status signals and negative status signals; the positive status signals include at least one of the following: code adoption operation trigger signal returned by the terminal application, task completion confirmation signal, and positive review signal; the negative status signals include at least one of the following: code abandonment signal, negative review signal, continuous interaction round exceeding limit disconnection signal, and user actively resetting session signal; the number of successful execution status signals only counts the number of positive status signals, and the negative status signals are used to correct business result indicators, trigger demotion processing, or serve as model quality evaluation data.
5. The system according to claim 1, characterized in that, The automatic fault switching mechanism is used for: When the number of consecutive timeouts for the target base model is detected to reach a preset threshold, or the error rate per unit time exceeds a preset threshold, a switch to the backup base model is triggered.
6. The system according to claim 1, characterized in that, The context adapter is specifically used for: The context data format corresponding to the target base model, which defines the speaker based on preset tags, is converted into a JSON object array format compatible with the backup base model, containing role fields and content fields. The conversion includes parsing the preset tags to extract the speaker role identifier and text content, and generating the JSON object array format according to the preset role mapping table.
7. An AI model routing method, characterized in that, Includes the following steps: The system acquires the computational resource consumption and terminal task execution status signals associated with historical model calls, identifies positive status signals from the terminal task execution status signals, and determines the number of successful execution status signals based on the positive status signals. It calculates and updates the resource conversion efficiency index of each basic model by dividing the computational resource consumption by the number of successful execution status signals. When the number of successful execution status signals is 0, the resource conversion efficiency index of the corresponding basic model is set to a preset penalty value or delayed until the next statistical period. Receive a call request associated with a stateful session and obtain the context history of the stateful session, and select a target base model based on the resource conversion efficiency index; Send the call request to the target base model and determine whether the call was successful; If the call is successful, the call request and the response received from the target base model are updated in the context history of the stateful session; If the call fails, the cost information associated with the failed call is recorded, and a switch to a backup base model is triggered. The switching steps include: reading the context history of the stateful session, extracting the corresponding speaker role identifier and text content by parsing the structured preset tags that define the speaker in the current data format, and mapping them to generate a JSON object array format containing role and content fields that is compatible with the backup base model.
8. The method according to claim 7, characterized in that, The computational resource consumption includes the number of tokens consumed for model calls; The terminal task execution status signals include positive status signals and negative status signals; the positive status signals include at least one of the following: code adoption operation trigger signal returned by the terminal application, task completion confirmation signal, and positive review signal; the negative status signals include at least one of the following: code abandonment signal, negative review signal, continuous interaction round exceeding limit disconnection signal, and user actively resetting session signal; the number of successful execution status signals only counts the number of positive status signals, and the negative status signals are used to correct business result indicators, trigger demotion processing, or serve as model quality evaluation data.
9. The method according to claim 7, characterized in that, The steps for triggering the switch to a backup base model are as follows: The switching is performed when the number of consecutive timeouts for the target base model reaches a preset threshold or the error rate per unit time exceeds a preset threshold.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 7 to 9.