A routing method and system based on model attribute capability automatic evaluation
Patent Information
- Application Number
- CN202611240340.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-17
- Publication Date
- 2026-09-25
AI Technical Summary
[0006]本发明的目的在于提供一种基于模型属性能力自动化评估的路由方法及系统,以解决现有技术中存在的以下技术问题:现有多模型路由方案中,模型能力评估粒度不足,通常仅以整体质量分或排行榜名次描述模型能力,无法区分模型在不同能力维度和不同难度任务上的差异化表现;缺少自动化、可持续更新的模型画像构建机制,当新模型入池或模型版本升级时需要大量人工介入;路由决策未充分区分硬约束与软目标,将上下文窗口不足、工具不支持、多模态不支持、敏感数据权限不满足等不可妥协的硬约束与成本、速度等可权衡的软目标混为一谈,导致路由结果不可靠;对任务需求画像刻画不足,未能将用户请求的能力需求权重、难度等级和服务约束结构化表达,使路由决策缺乏可解释性;以及缺乏将质量、成本和速度多目标统一到同一框架的路由决策机制的技术问题
[0030](1)本发明通过动态评测题库、题目能力权重向量与多维能力打分模块相结合,在多个能力维度和多个难度等级下分别计算模型的能力分数,并保留有效样本数量、置信度和能力专长标签,使每个模型的能力画像具备“分能力、分难度、带置信度”的精细刻画能力。路由决策因此可被解释为“当前任务需要哪些能力,候选模型在这些能力和难度下表现如何”,克服了现有方案仅以整体质量分或排行榜名次描述模型能力、无法区分差异化表现的技术缺陷。上述效果的关键技术手段在于相关性门控机制:当某能力维度的回答相关性低于预设阈值时,该题目的该能力维度得分被排除在外,而非记为0分参与计算。本领域技术人员通常将低相关性题目的得分记为0分处理,该做法会对模型在该能力维度的得分造成系统性低估,而本发明通过“排除而非记零”的方式从根本上消除了这一偏差,使能力画像更加客观准确,该技术效果是非显而易见的,具有预料不到性。
Smart Images

Figure CN122816907A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of large language model inference scheduling technology, and in particular to a routing method and system based on automated evaluation of model attribute capabilities. Background Technology
[0002] With the rapid increase in the number of large language models, various models exhibit significant differences in performance across dimensions such as logical reasoning, code generation, long text processing, multilingual translation, tool invocation, security compliance, invocation cost, response latency, and context window limits. A single model struggles to simultaneously achieve high output quality, service cost, and response speed across all business scenarios. Dynamic routing technology for multi-model service pools has become a key technological direction for large model inference service systems. Its core objective is to automatically match a suitable inference model from a candidate model pool after receiving a user request, combining request characteristics and system constraints. Early routing solutions primarily focused on cascading calls and cost optimization, relying on methods such as hint adaptation, model approximation, and multi-model concatenation to reduce inference overhead. They trained routing selection logic based on historical query and preference data to balance accuracy and cost. However, these solutions employ a simplistic binary decision-making approach of choosing between strong and weak models, relying on fixed historical data or end-to-end routing models, and cannot accurately characterize the differentiated performance of models across different capabilities and task difficulties.
[0003] Subsequent research into related technologies will extend routing solutions to areas such as large-scale model pool adaptation and real-world online scenario deployment. This involves constructing a graph linking tasks, queries, and models to mine interaction features, establishing multi-model routing evaluation benchmarks, forming a unified algorithm comparison framework, and refining the routing granularity from single query selection to a staged scheduling of agent-driven execution. However, most existing technologies focus only on the single-point problem of "predicting the optimal model from an input query," exhibiting several inherent shortcomings. The most prominent is the coarse granularity of model capability evaluation, using only overall scores and rankings to summarize model performance, lacking a mechanism to combine question weights and relevance gating to calculate fine-grained capability scores, resulting in a lack of precise basis for routing decisions.
[0004] Meanwhile, the existing routing architecture suffers from several engineering flaws: First, it lacks an automated model profiling and iteration mechanism. When adding new models, iterating versions, updating the question bank, or experiencing online performance drift, it cannot automatically trigger retesting and profile updates, requiring manual maintenance of routing rules. Second, it fails to distinguish between uncompromising hard constraints and negotiable soft objectives, mixing hard limitations such as context windows, tool support, multimodal capabilities, and sensitive data permissions with costs and latency in the same weighted system, easily selecting models that cannot perform tasks correctly. Third, it lacks standardized task profiling capabilities, failing to quantify the capability requirements, difficulty levels, and various service constraints of user requests, resulting in uninterpretable routing decisions. In addition, the existing solution lacks a unified multi-objective hierarchical decision-making framework, unable to first define a minimum quality threshold to screen safe candidate models, and then flexibly select the best from the qualified set based on performance, cost, and speed, easily leading to low-price, low-quality models and high-speed, low-quality models crowding out the optimal choice.
[0005] Therefore, there is an urgent need for a routing method and system that can automatically, continuously, and interpretably construct fine-grained model capability profiles, and combine structured task profiles with hierarchical routing decision mechanisms to achieve automatic model selection under multiple objective constraints of quality, cost, and speed. Summary of the Invention
[0006] The purpose of this invention is to provide a routing method and system based on automated evaluation of model attribute capabilities, in order to solve the following technical problems existing in the prior art: In existing multi-model routing schemes, the granularity of model capability evaluation is insufficient, typically describing model capabilities only using overall quality scores or leaderboard rankings, failing to distinguish the differentiated performance of models across different capability dimensions and tasks of varying difficulty; there is a lack of an automated, continuously updated model profiling mechanism, requiring significant manual intervention when new models are added to the pool or model versions are upgraded; routing decisions do not adequately distinguish between hard constraints and soft objectives, conflating uncompromising hard constraints such as insufficient context windows, tool incompatibility, multimodal incompatibility, and unmet sensitive data permissions with trade-off soft objectives such as cost and speed, leading to unreliable routing results; insufficient characterization of task requirements, failing to structurally express the weights, difficulty levels, and service constraints of user requests, resulting in a lack of interpretability in routing decisions; and the lack of a technical mechanism to unify multiple objectives such as quality, cost, and speed into a single framework for routing decisions.
[0007] To achieve the above objectives, this invention provides a routing method based on automated evaluation of model attribute capabilities, comprising the following steps:
[0008] A dynamic assessment question bank is constructed, organized by ability dimension, difficulty level, and task type. Each question carries a question ability weight vector, which represents the requirement weight of the question for each ability dimension. The model to be assessed from the model pool generates answers on the dynamic assessment question bank. The answers of the model to be assessed are scored using multi-dimensional abilities, outputting answer validity, answer relevance for each ability dimension, and ability score. Based on the question ability weight vector, the answer relevance, and the ability score, the ability score of the model to be assessed is calculated for each ability dimension and difficulty level, and a model profile including ability level tags and service attributes is generated. The model profile is stored in a model profile database. Upon receiving a user request, the system performs structured analysis on the request, generating a task profile containing the weights of each capability dimension, difficulty level, and service requirements. Based on the hard constraints in the task profile, hard filtering is performed on candidate models in the model profile database, eliminating models that do not meet the hard constraints. Based on the difficulty level and main capability dimension of the task profile, tier matching is performed on the candidate models that pass the hard filtering, determining a set of candidate models within each tier. A quality matching score is calculated for each candidate model in the candidate model set, and the final model is selected and invoked from the safe candidate set according to the operating mode. This invention, through automated model profile construction in the offline stage and structured task profile matching in the online stage, accurately aligns "what capabilities the model possesses" with "what capabilities the current task requires," and achieves accurate, sustainable, and interpretable automatic model selection through a three-layer decision-making mechanism of hard filtering, tier matching, and quality matching scoring.
[0009] Preferably, the dynamic evaluation question bank is divided into a core question bank, a dynamically expanded question bank, and a hidden question bank. The core question bank is used to perform pooling evaluation on all models to be evaluated. The dynamically expanded question bank is continuously supplemented according to business needs and model shortcomings. The hidden question bank is used to prevent pollution from the public question bank and to perform random checks. When a new model is added to the pool, the model version changes, the question bank is updated, or the online performance drifts, the evaluation of the model to be evaluated on the dynamic evaluation question bank is automatically triggered. The three-level question bank structure ensures the comprehensiveness, continuity, and anti-pollution of the evaluation, while the automatic triggering mechanism eliminates the need for manual maintenance of routing rules.
[0010] Preferably, the question ability weight vector is generated by at least one of the following: scoring model, expert rules, and manual annotation; the question ability weight vector covers all ability dimensions, and the sum of the weights of each dimension is a normalized value. Multiple weight generation methods can be flexibly selected or combined according to business scenarios, and the normalization constraint ensures the comparability of weights between different questions.
[0011] Preferably, the validity of the multi-dimensional ability scoring output includes at least one of valid, invalid, rejected, and truncated responses; when the relevance of the response in a certain ability dimension is lower than a preset relevance threshold, the ability score for that ability dimension is not included in the final ability score calculation for that ability dimension of the model under evaluation. The relevance gating mechanism avoids incorrectly including irrelevant abilities as model weaknesses or strengths, thereby improving the accuracy of ability profiling.
[0012] Preferably, the multi-dimensional ability scoring also incorporates objective evaluation metrics, including at least one of the following: unit test pass rate, exact match rate, F1 score, structured data format validation results, and code execution accuracy. Combining objective metrics with subjective scoring further improves the scoring accuracy for tasks with objective answers, such as coding and mathematics.
[0013] Preferably, the model profile includes a model identifier, capability scores for each capability dimension at each difficulty level, number of valid samples, confidence level, capability expertise labels, capability tier labels, and service attributes. The service attributes include input price, output price, cache read / write price, first token latency, output speed, high percentile latency, context window, tool support flag, multimodal support flag, sensitive data processing permissions, and current availability status. Storing the capability assessment results and service attributes uniformly as a structured model profile provides a complete and callable data interface for routing decisions.
[0014] Preferably, in the step of performing structured detection on the user request, when the number of input tokens in the user request exceeds a preset threshold, long context preprocessing is further performed on the user request to generate a session summary, document structure features, relevant fragment summaries, and a compressed global summary. The compressed global summary is then used as input to generate the task profile. This structured preprocessing mechanism before routing avoids directly feeding excessively long inputs to the routing model, reduces routing computation costs, and improves the routing accuracy for complex document tasks.
[0015] Preferably, the hard constraints include: the context window length of the candidate model is not less than the number of input tokens requested by the user; the tool support capability of the candidate model meets the tool requirements in the task profile; the multimodal support capability of the candidate model meets the multimodal requirements in the task profile; the sensitive data processing permissions of the candidate model meet the sensitive data processing requirements requested by the user; and the current service status of the candidate model is available. These explicit hard filtering conditions ensure that all candidate models entering the subsequent scoring stage possess the basic capabilities and service conditions to complete the current task.
[0016] Preferably, in automatic mode, the gear matching selects candidate gears based on the correspondence between task difficulty level and preset default gears; when no usable model is available in the candidate gears, the gear is upgraded according to a preset strategy; when the task profile contains a high-risk task marker, downgrading is not allowed. The gear matching mechanism ensures the correspondence between task difficulty and model capability level, and the constraint of not downgrading for high-risk tasks further guarantees the system's safety and compliance.
[0017] Preferably, the step of selecting the final model from the safe candidate set based on the operating mode includes: determining the minimum acceptable quality matching score based on the highest quality matching score in the candidate model set and the preset maximum quality concession amount; forming a safe candidate set by selecting candidate models whose quality matching scores are not lower than the minimum acceptable quality matching score; and within the safe candidate set, selecting the model with the highest quality matching score according to the best performance mode, or selecting the model with the lowest estimated cost according to the cost priority mode, or selecting the model with the lowest estimated latency according to the speed priority mode. This two-layer decision structure separates "whether the model is eligible to participate" from "which qualified model is optimal," achieving flexible multi-objective decision-making while ensuring a minimum quality level.
[0018] Preferably, the system further includes: recording a routing log for each routing request, wherein the routing log includes the selected model identifier, operating mode, difficulty level, quality matching score, number of input tokens, number of output tokens, first token delay, output speed, execution result, and user feedback; updating model service attributes, refreshing model profiles, adjusting routing thresholds, and expanding the evaluation question bank based on the routing log. This closed-loop calibration mechanism enables the routing system to continuously optimize as the model pool and business distribution change, achieving adaptive evolution.
[0019] This invention also provides a routing system based on automated evaluation of model attribute capabilities, comprising: an offline model profiling subsystem, which includes a dynamic evaluation question bank module, a model evaluation execution module, a multi-dimensional capability scoring module, a model capability score calculation module, and a model profiling and grading module; an online task profiling and routing decision subsystem, which includes a routing preprocessing module, a task profiling generation module, a hard filtering module, a grading matching module, a quality matching scoring module, and a running mode selection module; and a model profiling database for storing model profilings of all candidate models. The offline model profiling subsystem and the online task profiling and routing decision subsystem share data through the model profiling database, together forming a complete closed-loop routing system.
[0020] Preferably, the routing system further includes a service attribute monitoring module, used to continuously update the input price, output price, cache read / write price, first token latency, output speed, high quantile latency, context window, tool support flag, multimodal support flag, sensitive data processing permissions, and current availability status of each candidate model, and write the updated service attributes into the model profile database. Continuous updating of service attributes ensures the real-time nature of the cost, latency, and availability information upon which routing decisions are based.
[0021] Preferably, the routing system further includes a feedback calibration module, used to record routing logs for each routing request, and update model service attributes, refresh model profiles, adjust routing thresholds, and expand the evaluation questions in the dynamic evaluation question bank module based on the routing logs. The feedback calibration module feeds back online actual operation data to the offline profile construction stage, forming a complete closed-loop optimization mechanism.
[0022] Preferably, the routing preprocessing module is further configured to perform long context preprocessing on the user request when the number of input tokens in the user request exceeds a preset threshold, generate a session summary, document structure features, related fragment summary and compressed global summary, and output the compressed global summary to the task profile generation module.
[0023] Preferably, the routing system further includes a capability dimension system module, which is signal-connected to the dynamic evaluation question bank module and the task profile generation module, respectively. This module defines a capability dimension space shared by the offline model profile construction subsystem and the online task profile and routing decision subsystem. The capability dimensions include at least one of the following: intent understanding and instruction compliance, general knowledge and factual question answering, logical reasoning and mathematical ability, coding and software engineering ability, document understanding and long context ability, data and structured information processing, multilingual and translation ability, writing and style control, tool invocation and proxy execution ability, and security compliance and rejection boundaries. The capability dimension system module ensures that the offline evaluation and online task profile use the same capability coordinate system, thereby ensuring semantic consistency in the matching calculation between model capability scores and task requirement weights.
[0024] Preferably, the model profiling and grading module adopts a relative sorting model grading principle, grading for each capability dimension separately. Specifically, within any capability dimension, models are first sorted according to their capability scores at the corresponding expert difficulty level, and models ranking in the top first preset proportion are marked as expert models for that capability dimension. After removing models already marked as expert models, the remaining models are sorted according to their capability scores at the corresponding complexity level, and models ranking in the top second preset proportion are marked as strong models for that capability dimension. After removing models already marked as expert and strong models, the remaining models are sorted according to their capability scores at the corresponding standard difficulty level, and models ranking in the top third preset proportion are marked as standard models for that capability dimension; the remaining models are marked as weak models for that capability dimension.
[0025] Specifically, the model profiling and grading module adopts a relative ranking principle for model grading and generates model grading labels for each capability dimension. Specifically, for each capability dimension, the models in the model pool are first sorted from highest to lowest according to their capability scores at the L4 difficulty level, and the top 25% of models are marked as expert models for that capability dimension. Then, models marked as expert models are removed from the model pool, and the remaining models are sorted from highest to lowest according to their capability scores at the L3 difficulty level, with the top 33% marked as strong models for that capability dimension. Next, models marked as expert and strong models are removed, and the remaining models are sorted from highest to lowest according to their capability scores at the L2 difficulty level, with the top 50% marked as standard models for that capability dimension; the remaining models are marked as weak models for that capability dimension.
[0026] This tiering method doesn't simply divide the same ranking results into top, middle, and bottom tiers, nor does it create expert models by superimposing multiple strong model labels. Instead, it filters models step-by-step based on L4, L3, and L2 difficulty scores. This approach allows the system to prioritize identifying expert models that excel on high-difficulty tasks, and then identify strong models suitable for complex tasks and standard models suitable for standard tasks from the remaining models. This avoids the fixed score threshold becoming ineffective under different model pool sizes, changes in overall model capabilities, or model version upgrades. Expert models, strong models, standard models, and weak models are all labels for a model on a specific capability dimension, not labels for the model as a whole. Therefore, a model might be an expert model on capability x and a weak model on capability y. Subsequent task profiling and tier matching also correspond to the capability dimension and difficulty of the current user's question. For example, if capability x has the highest weight in a question's capability dimension vector, and the question difficulty is L4, then only models that are expert models on capability x (L4 difficulty corresponds to expert models) are selected as candidate models.
[0027] The present invention also provides an electronic device, including a processor and a memory, wherein the memory stores computer program instructions, and the processor executes the computer program instructions to implement all the steps of the above-described routing method.
[0028] The present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements all the steps of the above-described routing method.
[0029] Compared with the prior art, the present invention has the following beneficial effects:
[0030] (1) This invention combines a dynamic evaluation question bank, a question ability weight vector, and a multi-dimensional ability scoring module to calculate the model's ability score under multiple ability dimensions and multiple difficulty levels, while retaining the number of valid samples, confidence level, and ability expertise labels. This enables each model's ability profile to have a finely detailed characterization ability that is "differentiated by ability, difficulty, and confidence level." The routing decision can therefore be interpreted as "what abilities are required for the current task, and how do candidate models perform under these abilities and difficulties," overcoming the technical defects of existing solutions that only describe model abilities with overall quality scores or rankings and cannot distinguish differentiated performances. The key technical means to achieve the above effect lies in the relevance gating mechanism: when the relevance of the answer to a certain ability dimension is lower than a preset threshold, the score of that ability dimension for that question is excluded, rather than being recorded as 0 points for calculation. Those skilled in the art usually record the score of low-relevance questions as 0 points, which will cause a systematic underestimation of the model's score in that ability dimension. However, this invention fundamentally eliminates this bias by "excluding rather than recording zero," making the ability profile more objective and accurate. This technical effect is not obvious and has unexpectedness.
[0031] (2) When a new model is added to the pool, the model version changes, the question bank is updated, or the online performance drifts, the present invention automatically triggers the model evaluation execution module to re-evaluate and refresh the model profile and ranking, without the need for manual intervention to maintain routing rules. This automatic triggering mechanism covers four scenarios: new model addition to the pool, model version upgrade, question bank content update, and online performance drift, realizing a fundamental shift from "manual maintenance of routing rules" to "event-driven automatic updates". Existing solutions (such as the progressive decision funnel method described in CN121882260A) have a full-link data closed loop, but their closed loop only updates the routing strategy weights and does not link the re-evaluation of offline model profiles. When the model's capabilities themselves change, the profile data on which the routing decision depends is still outdated. The present invention, on the other hand, directly triggers online feedback to the offline evaluation stage, realizing a deeper level of closed-loop optimization, fundamentally solving the technical problem that the model profile cannot be automatically updated with the dynamic changes of the model pool, and significantly reducing the operation and maintenance costs of large-scale multi-model service platforms.
[0032] (3) This invention first performs hard filtering based on context window length, tool support capability, multimodal support capability, sensitive data processing permissions, and current availability. Then, it calculates the quality matching score for candidate models that pass the hard filtering, thus avoiding situations where a model with a high quality score cannot complete the current task due to insufficient context window or lack of tool support. Existing solutions incorporate context window constraints into a weighted calculation system. Under this system, a model with a context window of only 32,000 tokens but a high overall quality score (e.g., a quality matching score of 8.60) may still be selected as the final model because its quality advantage outweighs its disadvantage of insufficient window. However, this model cannot complete the task when faced with a request with 180,000 input tokens, directly causing routing failure in actual deployment. This invention treats uncompromising conditions such as context window and tool support as pre-filtering conditions rather than weighting factors, fundamentally eliminating such unreliable routing results and making the routing results engineering reliable. This technical effect is unexpected because the mainstream design idea of existing solutions is to incorporate all constraints into a weighted system rather than processing them hierarchically.
[0033] (4) This invention adopts a two-layer operation mode selection structure: the first layer filters out a safe candidate set through quality retention constraints, ensuring that the quality matching score of all candidate models is not lower than the minimum acceptable quality baseline; the second layer selects the model with the highest quality matching score in the best performance mode, the model with the lowest estimated cost in the cost priority mode, or the model with the lowest estimated latency in the speed priority mode within the safe candidate set. The key innovation of the two-layer structure is to separate "qualification screening" from "optimal selection": the first layer ensures that all candidate models entering the second layer decision meet the quality baseline, preventing low-quality models from being wrongly selected due to cost or speed advantages. The single-layer weighted method of the existing solution has a systemic risk - when a model has extremely low cost or extremely fast speed, even if its quality is significantly lower than the optimal model, it may still win in the comprehensive weight calculation, resulting in a decline in user experience; the two-layer structure of this invention eliminates this risk from the architectural level, and under the premise of ensuring the quality baseline, it can flexibly balance the effect, cost and response experience in different business scenarios, solving the technical problem that the existing solution lacks the ability to unify multiple objectives of quality, cost and speed into the same framework.
[0034] (5) This invention records the routing log of each routing request through the feedback calibration module, including fields such as selected model identifier, running mode, quality matching score, first token delay, output speed, execution result, failure reason, and user feedback. The above information is used to update model service attributes, refresh model profile, adjust routing thresholds, and expand the evaluation question bank. Unlike the coarse-grained closed loop of existing solutions that only adjust the routing strategy weights, the closed-loop calibration of this invention achieves a precise closed loop of "perceived drift → targeted retesting → profile refresh": when the feedback calibration module detects that the online failure rate of a certain model in a specific capability (such as long context processing) continues to rise, the system not only adjusts the routing threshold, but also precisely triggers the retesting of the model on the question bank related to that capability, and updates the new evaluation results to the model profile, so that the profile data on which subsequent routing decisions are based is consistent with the actual capabilities of the model. This precise closed-loop mechanism enables the routing system to continuously self-optimize as the model pool changes and business distribution drifts, achieving long-term adaptability of routing decisions and solving the technical problem of the lack of online feedback and offline profile update linkage mechanism in existing solutions. Attached Figure Description
[0035] Figure 1 This is a schematic diagram of the overall architecture of a routing system that automatically evaluates model attribute capabilities.
[0036] Figure 2 A flowchart illustrating the method for building profiles for offline models;
[0037] Figure 3 This is a flowchart illustrating the online model routing method. Detailed Implementation
[0038] The present invention will now be described in further detail with reference to the accompanying drawings.
[0039] This invention provides a routing method and system based on automated assessment of model attribute capabilities. Its core improvement lies in accurately aligning "what capabilities the model possesses" with "what capabilities the current task requires" by matching offline automated model profiling with online structured task profiling. Furthermore, through a three-layer decision-making mechanism of hard filtering, tier matching, and quality matching scoring, it solves the technical problems of coarse capability assessment granularity, unsustainable profile updates, lack of separation between hard constraints and soft objectives, and lack of a unified framework for multi-objective decision-making in existing solutions.
[0040] Example 1
[0041] like Figure 1 As shown in the figure, this embodiment provides a routing system based on automated evaluation of model attribute capabilities. The system consists of an offline model profile construction subsystem, an online task profile and routing decision subsystem, a model profile database, a service attribute monitoring module, and a feedback calibration module.
[0042] The offline model profiling construction subsystem includes: a dynamic evaluation question bank module, a capability dimension system module, a question capability weight labeling module, a model evaluation execution module, a multi-dimensional capability scoring module, a model capability score calculation module, and a model profiling and grading module.
[0043] The online task profiling and routing decision subsystem includes: a routing preprocessing module, a task profiling generation module, a hard filtering module, a gear matching module, a quality matching and scoring module, and an operating mode selection module.
[0044] The dynamic evaluation question bank module is used to store model evaluation questions and their structured metadata. Each question includes at least the question number, task type, difficulty level, consecutive difficulty scores, primary ability, secondary ability, reference answer, scoring criteria, objective indicators, input token range, expected output length, tool requirements, multimodal requirements, security sensitivity markers, domain tags, language tags, and metadata version.
[0045] The capability dimension system module is used to define the capability space shared by model evaluation and task profiling. This embodiment adopts ten core capability dimensions, numbered C1 to C10, namely: C1 Intent Understanding and Instruction Compliance (understanding user goals, constraints, formats, and implicit intent), C2 General Knowledge and Fact-Based Question Answering (fact accuracy, common sense, and domain-based knowledge), C3 Logical Reasoning and Mathematical Ability (multi-step reasoning, mathematical calculation, and complex problem decomposition), C4 Code and Software Engineering Ability (code generation, debugging, testing, and interface usage), C5 Document Understanding and Long Context Ability (long text, multiple documents, cross-segment integration, and evidence location), C6 Data and Structured Information Processing (table analysis, structured extraction, and statistical interpretation), C7 Multilingual and Translation Ability (cross-language understanding, translation, and localization), C8 Writing and Style Control (summary rewriting, style control, and official document writing), C9 Tool Invocation and Agent Execution Ability (function calls, retrieval, code execution, and multi-step tasks), and C10 Security Compliance and Rejection Boundaries (high-risk identification, privacy protection, and compliant rejection). The competency dimension system module, dynamic assessment question bank module, and task profile generation module are connected by signals to ensure that offline assessment and online task profile use the same competency coordinate system.
[0046] The question ability weighting module is used to generate the required weights for each ability dimension for each question based on the question, the reference answer, and the scoring criteria. ,in Indicates the question number. This represents the ability dimension number. Weights can be generated by a strong scoring model, expert rules, or manual annotation and are stored in the question bank metadata. This is the ability weight vector for each question. Covering all capability dimensions, constraints: (The sum of the weights for each dimension is a normalized value of 1). Where: This represents the total number of all capability dimensions. Indicates the first The ability weight vector of the question.
[0047] The model evaluation execution module is connected to the dynamic evaluation question bank module. When a new model is added to the pool, the model version changes, the question bank is updated, or the online performance drifts, the module calls the model to be evaluated to generate answers on the dynamic evaluation question bank and sends the questions, reference answers, scoring criteria, and model answers to the multi-dimensional ability scoring module.
[0048] The multi-dimensional capability scoring module is signal-connected to the model evaluation execution module and is used to determine the validity of the model's answers, the relevance of each capability dimension, and the corresponding scores. Its output includes a validity flag (validity), a capability list (capability_list), and a brief rationale (brief_rationale). Validity can take values such as valid, invalid, refusal, or truncated. Each capability dimension in the capability_list contains a relevance value and a score.
[0049] The model ability score calculation module and the multi-dimensional ability scoring module are connected by a signal to calculate the ability score based on the question's ability weight. Relevance of the answer and ability score Calculate the ability score of model m under ability k and difficulty l. The specific calculation method is as follows: The summation range is within the range of difficulty. Lower, ability All valid questions with a relevance level not lower than the preset relevance threshold Simultaneously, record the effective sample count (effective_sample_count) and the confidence level.
[0050] When the relevance of an answer to a certain ability dimension is lower than a preset relevance threshold, the ability score for that dimension is not included in the model's final ability score calculation for that dimension, instead of being recorded as 0. This relevance gating mechanism avoids incorrectly including irrelevant abilities as model weaknesses or strengths, thereby improving the accuracy of ability profiling.
[0051] The model profiling and grading module is signal-connected to the model capability score calculation module. It is used to convert the model's capability and difficulty scores into a model profile, and further generate capability specialty tags, service attribute tags, and capability grading tags. The capability grading includes four levels: weak model, standard model, strong model, and expert model, which correspond to task difficulty levels L1, L2, L3, and L4, respectively.
[0052] The model profile includes the following fields: model_id (model identifier), capability_scores (capability scores for each capability dimension at each difficulty level, with each dimension including L1, L2, L3, L4 scores, effective_sample_count, and confidence), specialty_tags (capability and specialty tags), service_attributes (service attributes), availability_status (current availability status), profile_version (profile version), and updated_at (update time).
[0053] The service_attributes include: input_price, output_price, cache_read_price, cache_write_price, avg_first_token_latency_ms (average first token latency in milliseconds), avg_output_tokens_per_second (average output speed in tokens per second), p95_latency_ms (high quantile latency in milliseconds), context_window (context window length), tool_supported (tool support flag), multimodal_supported (multimodal support flag), and sensitive_data_allowed (sensitive data processing permissions).
[0054] The service attribute monitoring module communicates with the model profile database to continuously update the aforementioned service attributes of each candidate model and writes the updated service attributes into the model profile database. Continuous updates of service attributes ensure the real-time nature of cost, latency, and availability information used for routing decisions.
[0055] The routing preprocessing module performs structured inspections on user requests, including detecting the number of input tokens, the number of historical context tokens, whether images, audio, video, file types, whether tools are required, and whether sensitive data rules are met. For excessively long context requests with an input token count exceeding a preset threshold, the routing preprocessing module can also generate session summaries, document structure features, top-k relevant fragment summaries, and compressed global summaries to replace the full-text input routing model, thereby reducing routing computation costs.
[0056] The task profile generation module is signal-connected to the routing preprocessing module and is used to output a task profile based on user requests, routing preprocessing results, and system constraints. The task profile must include at least `capability_need` (the weights of requirements for each capability dimension). The fields include difficulty_level (task difficulty level), tool_required (whether tool calls are required), multimodal_required (whether multimodal support is required), contains_sensitive_data (whether sensitive data is contained), input_tokens (number of input tokens), file_count (number of files), file_types (file types), user_mode (user running mode), and router_input_summary (route input summary).
[0057] The hard filtering module is signal-connected to the task profiling module and is used to eliminate candidate models that do not meet the hard constraints before calculating the quality matching score. The hard constraints include: the context window length of the candidate model is not less than the number of input tokens currently requested (…). The tool support capabilities of the candidate models meet the requirements of the task tools. satisfy ); Indicates the candidate model number, referring to the first... One model, This indicates the task requested by the current user.
[0058] The candidate model's multimodal support capability meets the task's multimodal requirements; the candidate model's sensitive data processing permissions meet the sensitive data processing requirements of the current request; the candidate model's current service status is available. ).
[0059] The task difficulty matching module is signal-connected to the hard filtering module and is used to select candidate task difficulty levels based on the task's difficulty level and the main task capability dimension. The correspondence between task difficulty levels and default model difficulty levels is as follows: L1 (single-step, factual, simple format conversion, or low-risk task) corresponds to the weak model difficulty level; L2 (clear task boundaries, requiring certain comprehensive capabilities) corresponds to the standard model difficulty level; L3 (multi-step, multi-constraint, multi-document, or highly specialized) corresponds to the strong model difficulty level; L4 (highly specialized, high-risk, multi-tool, multi-round planning, or complex context) corresponds to the expert model difficulty level.
[0060] In automatic mode, the gear matching module selects candidate gears according to the default correspondence mentioned above; users can also specify fixed gears. If there is no available model in the corresponding gear, the gear is upgraded according to the preset strategy; when the task profile contains a high-risk task marker, downgrading is not allowed.
[0061] The quality matching scoring module is signal-connected to the gear matching module and is used to evaluate each candidate model within the candidate gear. Calculate quality matching score .in Indicates the current task. Indicates the task's impact on capabilities Demand weight, . Representation Model In ability Difficulty The ability score is used to determine the degree of match between the model's capability profile and the task requirement profile.
[0062] The operation mode selection module is signal-connected to the quality matching scoring module and is used to determine the final model based on the optimal performance mode, cost-priority mode, or speed-priority mode. This module employs a two-layer structure to complete the decision: the first layer determines the model based on the highest quality matching score among the candidate models. With the system's maximum allowable mass concession Determine the minimum acceptable quality baseline. And filter out those with a quality match score of not less than Safety candidate set .
[0063] The second layer is in the safe candidate set. Within the system, the model with the highest quality matching score is selected based on the best performance mode, the model with the lowest estimated cost is selected based on the cost priority mode, or the model with the lowest estimated latency is selected based on the speed priority mode. This two-layer structure separates "whether a model is eligible to participate" from "which of the eligible models is the best," preventing low-quality models from being mistakenly selected due to cost or speed advantages.
[0064] The feedback calibration module communicates with the operation mode selection module and the dynamic evaluation question bank module respectively. It is used to record the routing log for each routing request, and refresh the model profile, adjust the routing threshold, and update the evaluation question bank based on the routing log. The routing log contains the following fields: request_id, selected_model (selected model identifier), routing_mode (operation mode), difficulty_level (difficulty level), candidate_count_after_filter (number of candidates after filtering), quality_score (quality matching score), input_tokens (number of input tokens), output_tokens (number of output tokens), first_token_latency_ms (first token latency in milliseconds), output_tokens_per_second (output speed in tokens per second), success (execution result), failure_reason (failure reason), and user_feedback (user feedback).
[0065] Through the above system structure, the offline model profiling construction subsystem and the online task profiling and routing decision subsystem achieve data sharing through the model profiling database. The feedback calibration module feeds back the actual online operation data to the offline profiling construction process, forming a complete closed-loop routing system.
[0066] Example 2
[0067] like Figure 2 As shown, this embodiment provides an offline implementation process for a routing method based on automated evaluation of model attribute capabilities, using an enterprise multi-model large language model service platform as an application scenario. This platform integrates multiple candidate models, including a low-cost general-purpose model (Model_A), a model with strong code capabilities (Model_B), a model with strong long-context capabilities (Model_C), and a multimodal model (Model_D). The system pre-configures ten capability dimensions and four difficulty levels, and the dynamic evaluation question bank consists of a core question bank, a dynamically expanded question bank, and a hidden question bank.
[0068] Step S101: Construct a dynamic assessment question bank.
[0069] The dynamic evaluation question bank module organizes questions according to "ability dimension × difficulty level × task type," and is divided into three parts: a core question bank, a dynamically expanded question bank, and a hidden question bank. The core question bank is used for evaluation of all models entering the pool, covering task types such as instruction following, knowledge question answering, mathematical reasoning, coding, long document comprehension, table analysis, multilingual translation, writing, tool calling, and security compliance; the dynamically expanded question bank is continuously supplemented with new questions based on business needs and model shortcomings; the hidden question bank is used to prevent contamination from the public question bank and to conduct random checks to ensure the authenticity of the evaluation results.
[0070] Step S102: Label the question metadata and ability weights.
[0071] The question ability weight annotation module generates an ability weight vector for each question. Taking the code debugging question as an example, its ability weights are: coding ability 0.50, logical reasoning 0.25, instruction following 0.15, and long context 0.10, indicating the distribution of ability weights required to solve the question. Taking the financial statement analysis question as an example, its ability weights are: long context 0.30, data analysis 0.25, logical reasoning 0.20, writing 0.15, and safety compliance 0.10.
[0072] Step S103: Perform an evaluation on each model in the model pool.
[0073] When Model_A, Model_B, Model_C, and Model_D are added to the model pool, the model evaluation execution module is automatically triggered, calling each model to generate answers on the evaluation question bank. For coding questions, the system combines unit test results and scoring models for evaluation; for open-ended writing questions, long document question-and-answer questions, and tool calling questions, the system combines reference answers, scoring criteria, and strong scoring models for multi-dimensional scoring. The system automatically triggers re-evaluation when a new model is added, an old model is upgraded, the model vendor adjusts its version, the question bank is updated, or online performance drifts.
[0074] Step S104: Perform multi-dimensional capability scoring.
[0075] The multi-dimensional ability scoring module outputs the validity of the answer and the correlation between each ability dimension based on the question, the reference answer, the scoring criteria, and the model answer. Ability score And a brief reason for the scoring. If a model does not provide any code or debugging logic in the coding question, the relevance of the code ability may be lower than the preset relevance threshold (e.g., 0.5), and the code ability score for that question will not be included in the calculation of the model's code ability score; if the answer is truncated, it can be marked as truncated and the confidence level can be reduced.
[0076] Step S105: Calculate the model's ability score under each ability dimension and difficulty level. .
[0077] The model ability score calculation module generates scores based on ability and difficulty. For example, Model_B scores 8.6 on code capability L3 with a high confidence level, and the number of valid samples is above a preset low-confidence sample count threshold (e.g., 10); Model_C scores 8.4 on long context capability L3 with a high confidence level. The low-confidence sample count threshold can be set to 10, meaning that when the number of valid samples in a certain capability dimension is lower than this threshold, the corresponding confidence level is marked as low.
[0078] Step S106: Generate confidence scores.
[0079] The model capability score calculation module simultaneously records the effective sample count and confidence level for each capability dimension at each difficulty level, which are used for reliability assessment in subsequent routing decisions.
[0080] Step S107: Create a model image.
[0081] The model profiling and grading module integrates the aforementioned capability scores, confidence levels, number of valid samples, and service attributes into a structured model profile and generates capability specialty labels. For example, the system labels Model_B as coding_strong (strong coding ability) and Model_C as long_context_strong (strong long context ability), and records service attributes such as input price, output price, first token delay, output speed, context window, and tool support for each model.
[0082] Step S108: The model is classified and written into the model profile database.
[0083] The model profiling and grading module generates ability gradation labels based on ability scores and writes the complete model profile into the model profile database. Ability gradations are divided into four levels: weak model, standard model, strong model, and expert model, corresponding to the default model selection for difficulty levels L1 to L4, respectively.
[0084] Specifically, the model ranking uses a relative sorting method, which is executed separately for each capability dimension. Taking the code capability dimension as an example, assuming the model pool contains 20 models, each of which has obtained capability scores at difficulty levels L1, L2, L3, and L4, the system first sorts the models from highest to lowest according to their capability scores at the L4 difficulty level. The top 25% of models, i.e., the top 5 models, are marked as expert models in the code capability dimension.
[0085] Subsequently, the system removed the five models that had been marked as expert models from the model pool, sorted the remaining 15 models from high to low according to their ability scores at the L3 difficulty level of coding ability, and marked the top 33% of the remaining models, i.e., the top five models, as strong models in the coding ability dimension.
[0086] Next, the system continues to remove models already marked as expert models and strong models. The remaining 10 models are then sorted from highest to lowest according to their ability scores at the L2 difficulty level of coding ability. The top 50% of these remaining models (the top 5 models) are marked as standard models in the coding ability dimension. Finally, the remaining models that are not classified as expert, strong, or standard models are marked as weak models in the coding ability dimension.
[0087] For example, if Model_A and Model_B rank in the top 25% at the L4 coding ability difficulty level, they are labeled as expert models, not just strong models. If Model_C is not in the top 25% at L4, but ranks in the top 33% at the L3 coding ability level among the remaining models, then Model_C is labeled as a strong model. If Model_D is not in the expert or strong model set, but ranks in the top 50% at the L2 coding ability level among the remaining models, then Model_D is labeled as a standard model. The remaining models are labeled as weak models.
[0088] When a new Model_E is added to the model pool or an existing model version is upgraded, the system recalculates its capability scores across all capability dimensions and difficulty levels, and re-executes the aforementioned relative ranking and grading process. Because the grading is based on relative ranking within the model pool, rather than a fixed absolute score threshold, it can adapt to changes in the size of the model pool, improvements in overall model capabilities, and differences in capability distribution between different business model pools. Through this offline phase, the system generates a structured model profile for each model in the model pool, categorized by capability, difficulty, and service attributes, providing a complete data foundation for online routing decisions.
[0089] Example 3
[0090] like Figure 3As shown, this embodiment provides an online implementation process for a routing method based on automated evaluation of model attribute capabilities, using the enterprise multi-model platform for processing user financial statement comparison tasks as a specific scenario.
[0091] The user input request reads: "Please read the three annual financial reports I uploaded, compare the changes in revenue, gross profit margin, and cash flow of the three companies over the past three years, and output a table and an investment risk analysis." The user also uploaded three PDF files, with a total input length of approximately 180,000 tokens.
[0092] Step S201: Receive user request and perform structured detection.
[0093] The pre-processing module performs structured inspection on user requests. The inspection results are: input_tokens = 180,000, file_count = 3, file_types = pdf, multimodal_required = false, require_file_parsing = true, and contains_sensitive_data = false. Since the number of input tokens exceeds a preset threshold, the system enables long context preprocessing to extract the document structure, key financial tables, top-k relevant fragment summaries, and compressed global summaries for each financial report, replacing the full-text input routing model.
[0094] Step S202: Determine if the context is too long and perform pre-routing compression and digest.
[0095] The routing preprocessing module generates a session summary, document structure features, relevant fragment summaries, and a compressed global summary. The compressed global summary is then output to the task profile generation module as input for task profile generation. This mechanism reduces routing computation costs and improves routing accuracy for complex document tasks.
[0096] Step S203: Generate a task profile.
[0097] The task profile generation module outputs a structured task profile, with the required weights for each capability dimension. The task's difficulty levels are: instruction compliance 0.10, logical reasoning 0.15, long context 0.30, data analysis 0.25, writing 0.10, and security compliance 0.10; difficulty_level is L3; tool_required is false; and input_tokens is 180000. The primary capability of this task is long context, and its difficulty level is L3.
[0098] Step S204: Perform hard filtering.
[0099] The hard filtering module eliminates candidate models that do not meet the hard constraints based on model profiles and service attributes. Specifically, it eliminates models with fewer than 180,000 tokens in their context window, currently unavailable models, and models that cannot reliably handle PDF parsing results or long document inputs. Assume that Model_A and Model_C remain after filtering.
[0100] Step S205: Determine whether the user has specified a model gear and perform gear matching.
[0101] The tier matching module selects candidate models from the strong model tiers within the long context capability dimension based on the task master's long context capability and difficulty level L3. If the user does not specify a tier, selection is performed automatically; if the user explicitly specifies a strong model or expert model tier, sorting is only performed within the user-specified tiers.
[0102] Step S206: Calculate the candidate model quality matching score .
[0103] The quality matching scoring module calculates the quality matching scores for Model_A and Model_C respectively: Q(Model_A,x)=8.42, Q(Model_C,x)=8.05.
[0104] Step S207: Determine the operating mode and perform the final model selection.
[0105] The operation mode selection module makes the final decision based on the user's selected operation mode. If the user selects the best performance mode, then Model_A, which has the highest quality matching score, is selected. If the user selects the cost-first mode, then the maximum quality concession corresponding to L3 difficulty is selected. If we set it to 0.35, then =8.42-0.35=8.07. Only models with a quality-match score of 8.07 or higher are included in the safe candidate set. Model_C's score of 8.05 does not meet the requirement, so Model_A is still selected. If Model_C's score is 8.12 and its cost is significantly lower than Model_A, then the cost-priority mode will select Model_C from the safe candidate set. If the user selects the speed-priority mode, then the model with the lowest estimated latency will be selected from the safe candidate set.
[0106] Step S208: Call the final model to generate an answer and return it to the user.
[0107] Step S209: Record online feedback.
[0108] The system records information such as selected_model, routing_mode, difficulty_level, quality_score, first_token_latency_ms, output_tokens_per_second, success, and user_feedback. When a model's failure rate increases repeatedly in long document tasks, the feedback calibration module can trigger a retest of that model using a long context-related question bank and update its long context capability score or availability status.
[0109] Example 4
[0110] This embodiment provides an extended implementation of ultra-long context preprocessing and closed-loop calibration.
[0111] Building upon Example 3, when the routing preprocessing module detects that the number of input tokens requested by the user exceeds a preset threshold, it executes a long context preprocessing procedure, specifically generating the following: a conversation summary (compressing and summarizing historical dialogues), document structure features (extracting the chapter structure, table positions, and key financial indicators from the PDF file), a relevant fragment summary (retrieving the top-k fragments most relevant to the current question and generating a summary), and a compressed global summary (integrating the above information to generate a global summary). The compressed results are output to the task profile generation module, replacing full-text input and significantly reducing routing computation overhead.
[0112] The workflow of the feedback calibration module is as follows: The system records a complete routing log for each routing request, including fields such as request_id, selected_model, routing_mode, difficulty_level, candidate_count_after_filter, quality_score, input_tokens, output_tokens, first_token_latency_ms, output_tokens_per_second, success, failure_reason, and user_feedback.
[0113] Based on the routing logs, the feedback calibration module performs the following calibration operations: updates model service attributes (refreshes service_attributes based on measured latency and cost data); refreshes model profiles (triggers a re-evaluation by the model evaluation execution module when online performance deviates significantly from offline evaluation results); and adjusts routing thresholds (dynamically adjusts the maximum quality concession for each difficulty level based on business feedback). Expand the assessment question bank (add new questions to the dynamically expanded question bank based on online failure cases and question bank gaps).
[0114] For high-risk tasks such as those involving healthcare, law, finance, privacy, security, and minors, the system not only calculates the quality matching score but also checks the security compliance capability score (corresponding to C10 score), security compliance tags, and sensitive data processing permissions (sensitive_data_allowed) in the model profile. If a model does not meet the minimum security capability requirements, it cannot be selected as the final model, even if it is cheaper or faster. When downgrading for high-risk tasks, only upward upgrades are allowed; downgrading to lower-level models is not permitted to ensure the system's security compliance.
[0115] Example 5
[0116] This embodiment provides a comparison of parameter configurations and effects.
[0117] In a specific parameter configuration embodiment, the initial values of the key system parameters are set as follows: the relevance threshold is set to 0.5, meaning that when the relevance of an answer in a certain ability dimension is lower than 0.5, the score for that ability dimension will not be included in the model's final ability score calculation for that ability dimension; the low confidence sample number threshold is set to 10, meaning that when the number of valid samples in a certain ability dimension is lower than 10, the corresponding confidence level is marked as low confidence; and the quality difference threshold... The difficulty levels are set as follows: L1 corresponds to 0.70, L2 to 0.50, L3 to 0.35, and L4 to 0.20. These parameters are initial values only; the system can dynamically adjust them based on manual sampling and online feedback.
[0118] Table 1 compares the core capabilities of the routing system of the present invention with those of existing solutions.
[0119] Table 1. Comparison of the core capabilities of this invention and existing solutions
[0120] Granularity of capability assessment Overall quality score or ranking Divided by ability dimension and difficulty level Image update method Manual maintenance of routing rules Automatically trigger evaluation and refresh profile Hard constraint handling Unified weight calculation, without distinguishing hard constraints First perform hard filtering, then calculate the quality matching score. Multi-objective decision making Lack of a unified framework Two-layer structure: quality retention constraints + operation mode selection Extremely long context Direct input of routing model Pre-routing compression and digest preprocessing Route interpretability Black box decision Explainable ability profile and task profile matching
[0121] Comparative Example 1
[0122] Comparative Example 1 refers to the routing system of Examples 2 and 3, except that the relevance gating mechanism is removed. That is, when the relevance of the answer to a certain ability dimension is lower than the preset relevance threshold, the score of that ability dimension of the question (recorded as 0 points) is still included in the final ability score calculation of that ability dimension of the model. The rest of the structure is the same as that of Examples 2 and 3.
[0123] In Comparative Example 1, if Model_B provides no code logic in a certain coding question, its coding ability relevance is 0.2 (below the threshold of 0.5), but the coding ability score (0 points) for that question is included in the calculation, causing Model_B's coding ability score A_{Model_B, coding, L3} to be artificially lowered. Based on this, the system may incorrectly label Model_B as having average coding ability rather than strong coding ability, thus incorrectly selecting other models in the coding task routing.
[0124] Comparing Example 2 with Comparative Example 1, it can be seen that the relevance gating mechanism effectively avoids incorrectly including irrelevant capabilities in the model's weaknesses or strengths, making the model's capability profile more accurate and the routing decisions more reliable. This effect is unexpected because, before the introduction of the relevance gating mechanism, those skilled in the art would typically record low-relevance questions as 0 points instead of excluding them from the calculation.
[0125] Comparative Example 2
[0126] Comparative Example 2 follows the routing method of Example 3, except that the hard filtering step S204 is removed, and all candidate models directly enter the level matching and quality matching scoring stage. The remaining steps are the same as in Example 3.
[0127] In Comparative Example 2, if a candidate model's context window contains only 32,000 tokens, far fewer than the 180,000 tokens required for the current request, but the model has a high quality match score (let's say 8.60), then after removing hard filtering, this model might be selected as the final model. However, due to the insufficient context window, the model is unable to complete the current task, leading to routing failure.
[0128] Comparing Example 3 with Comparative Example 2, it can be seen that the introduction of the hard filtering step effectively prevents models with high quality scores but unable to complete the task from being selected, significantly improving the engineering reliability of the routing results. This effect is unexpected because existing solutions typically incorporate context window constraints and quality scores into a unified weighted calculation, rather than treating them as pre-processing hard filtering conditions.
[0129] Example 6
[0130] This embodiment provides some alternatives to the embodiments described above.
[0131] The capability dimension system is not limited to ten core capability dimensions. It can be expanded to more dimensions depending on the business scenario, such as visual understanding, audio understanding, video understanding, scientific research capabilities, legal capabilities, medical capabilities, financial analysis capabilities, and intelligent agent long-term planning capabilities. Alternatively, some capabilities can be merged into fewer dimensions.
[0132] Difficulty levels are not limited to four levels (L1 to L4); three, five, or continuous difficulty scores (ranging from 0 to 1) can also be used. Continuous difficulty scores can be directly used for interpolation calculations in quality matching scoring, for example, by weighted interpolation between adjacent difficulty levels based on continuous difficulty scores. .
[0133] Question Ability Weighting The generation method is not limited to strong scoring models; it can also be generated by expert manual annotation, rule template generation, voting by multiple scoring models, historical statistical learning, or a combination thereof.
[0134] Multidimensional ability scoring is not limited to strong scoring models alone. It can also be combined with objective evaluation indicators, such as unit test pass rate, exact match rate, F1 score, JSON format validation results, SQL execution accuracy, human review results, or integrated results of multiple scorers, to further improve the scoring accuracy of tasks with objective answers, such as code and mathematics.
[0135] The model grading method is not limited to a fixed score threshold. It can also adopt relative sorting, clustering, quantiles, business rules, or adaptive grading based on historical online success rates.
[0136] Task profile generation methods are not limited to large language model prompts; they can also be generated by small routing models, traditional text classifiers, embedding models with nearest neighbor retrieval, graph neural networks, rule systems, or multi-model integration.
[0137] The utility function for selecting the operating mode is not limited to single-index ranking; the Nash / Cobb-Douglas combined utility function can also be used. ,in , , Representing the model respectively Normalized revenue values in terms of quality, cost, and speed. , , These represent the weighting indices for quality, cost, and speed, respectively, with a sum of 1. This function emphasizes multi-dimensional balance, preventing extreme disadvantages in a single dimension from being offset by high weights. It is suitable for scenarios that are sensitive to quality but still require consideration of efficiency. The above alternative is compatible with the default security candidate set framework and can be flexibly switched according to task type, business risk level, and user preferences.
[0138] Quality matching scoring functions are not limited to linear weighted matching scores. Alternatively, nonlinear functions, learned ranking models, Bayesian uncertainty estimation, reinforcement learning strategies, Pareto front ranking, or scoring functions with added confidence penalties can be used.
[0139] Cost estimation methods are not limited to static unit prices; dynamic calculations are also possible: Input token count × input unit price + estimated output token count × output unit price, plus cache read / write costs, tool call costs, and other additional costs. For scenarios billed by model instance occupancy time, duration × instance unit price is used. For batch processing scenarios, batch-allocated costs are employed. Latency estimation methods are not limited to first token latency; average total latency, high-quantile latency (such as p95 latency), output token throughput rate, or risk-adjusted latency can also be used. Considering both average latency and tail risk, it is suitable for high service level protocol scenarios that are sensitive to long-tail latency; among which This is the risk aversion coefficient, with a value range of [0,1]. The larger the value, the more sensitive the risk is to tail delay. For the model The average expected delay; Confidence level Conditional risk value, i.e., delay exceeding Expected delay at quantiles.
[0140] Long context processing methods are not limited to summarization and compression; they can also be implemented using context compression, top-k fragment retrieval, structured metadata extraction, sampling windows, sliding window statistics, or hierarchical routing.
[0141] The data sources for feedback calibration are not limited to user feedback, but can also come from manual sampling, task success rate, tool execution results, automatic evaluation models, retry logs, exception logs, or cost delay monitoring systems.
[0142] The routing system of this invention can be deployed on cloud-based model service platforms, enterprise-internal model gateways, local private inference gateways, or agent platforms; it can be used for ordinary chat requests as well as for model selection before each tool call by the agent. Capability dimensions, difficulty levels, service attributes, and operating modes can all be expanded according to business needs, without relying on a fixed model or a fixed routing algorithm.
[0143] The preferred embodiments of the present invention have been described in detail above. It should be understood that those skilled in the art can make numerous modifications and variations based on the concept of the present invention without creative effort. Therefore, all technical solutions that can be obtained by those skilled in the art based on the concept of the present invention through logical analysis, reasoning, or limited experimentation on the basis of existing technology should be within the scope of protection defined by the claims.
Claims
1. A routing method based on automated evaluation of model attribute capabilities, characterized in that, Includes the following steps: A dynamic assessment question bank is constructed, which is organized according to ability dimensions, difficulty levels and task types. Each question carries a question ability weight vector, which represents the requirement weight of the question for each ability dimension. The model to be evaluated from the model pool is used to generate answers on the dynamic evaluation question bank; The model to be evaluated is given a multi-dimensional ability score, and the output is the validity of the answer, the relevance of the answer in each ability dimension and the ability score. Based on the question ability weight vector, the answer relevance, and the ability score, calculate the ability score of the model to be evaluated under each ability dimension and difficulty level. When the answer relevance of a certain ability dimension is lower than a preset relevance threshold, the ability score of that ability dimension is not included in the final ability score calculation of the model to be evaluated in that ability dimension. Generate a model profile containing ability level tags and service attributes based on the ability score, and store the model profile in the model profile database. Receive user requests, perform structured detection on the user requests, and generate a task profile that includes the requirement weights of each capability dimension, difficulty level, and service requirements; Based on the hard constraints in the task profile, perform hard filtering on the candidate models in the model profile database to remove models that do not meet the hard constraints. Candidate levels are determined by combining the difficulty level and the main ability dimension of the task profile. Level matching is performed on the candidate models that pass the hard filtering, and the candidate models corresponding to the candidate levels are combined to form a candidate model set. Calculate the quality matching score for each candidate model in the candidate model set; The minimum acceptable quality matching score is determined based on the highest quality matching score in the candidate model set and the preset maximum quality concession amount. Candidate models with quality matching scores not lower than the minimum acceptable quality matching score are formed into a safe candidate set. The final model is selected and called from the safe candidate set according to the operating mode.
2. The method according to claim 1, characterized in that, The dynamic evaluation question bank is divided into a core question bank, a dynamically expanded question bank, and a hidden question bank. The core question bank is used to perform pooling evaluation on all models to be evaluated. The dynamically expanded question bank is continuously supplemented according to business needs and model shortcomings. The hidden question bank is used to prevent pollution of the public question bank and to conduct random checks. When a new model is added to the pool, the model version changes, the question bank is updated, or the online performance drifts, the evaluation of the model to be evaluated on the dynamic evaluation question bank is automatically triggered.
3. The method according to claim 1, characterized in that, The question ability weight vector is generated by at least one of the scoring model, expert rules, and manual annotation; the question ability weight vector covers all ability dimensions, and the sum of the weights of each dimension is a normalized value.
4. The method according to claim 1, characterized in that, The validity of the multidimensional ability scoring output includes at least one of valid, invalid, rejected, and truncated responses; when the relevance of the response in a certain ability dimension is lower than a preset relevance threshold, the ability score of that ability dimension will not be included in the final ability score calculation of the model to be evaluated in that ability dimension.
5. The method according to claim 4, characterized in that, The multidimensional capability scoring also incorporates objective evaluation metrics, which include at least one of the following: unit test pass rate, exact match rate, F1 score, structured data format verification results, and code execution accuracy.
6. The method according to claim 1, characterized in that, The model profile includes model identifier, ability score of each ability dimension at each difficulty level, number of valid samples, confidence level, ability expertise label, ability tier label and service attributes; the service attributes include input price, output price, cache read / write price, first token latency, output speed, high percentile latency, context window, tool support flag, multimodal support flag, sensitive data processing permissions and current availability status.
7. The method according to claim 1, characterized in that, In the step of performing structured detection on the user request, when the number of input tokens of the user request exceeds a preset threshold, long context preprocessing is also performed on the user request to generate a session summary, document structure features, related fragment summary and compressed global summary, and the compressed global summary is used as input to generate the task profile.
8. The method according to claim 1, characterized in that, The hard constraints include: the context window length of the candidate model is not less than the number of input tokens requested by the user; the tool support capability of the candidate model meets the tool requirements in the task profile; the multimodal support capability of the candidate model meets the multimodal requirements in the task profile; the sensitive data processing permissions of the candidate model meet the sensitive data processing requirements requested by the user; and the current service status of the candidate model is available.
9. The method according to claim 1, characterized in that, In automatic mode, the gear matching selects candidate gears based on the correspondence between the task difficulty level and the preset default gears; when there are no available models among the candidate gears, the gear is upgraded according to a preset strategy. Downgrading is not allowed when the task profile contains high-risk task markers.
10. The method according to claim 1, characterized in that, The step of selecting the final model from the safe candidate set according to the operating mode includes: determining the minimum acceptable quality matching score based on the highest quality matching score in the candidate model set and the preset maximum quality concession amount; forming a safe candidate set by selecting candidate models whose quality matching scores are not lower than the minimum acceptable quality matching score; and selecting the model with the highest quality matching score in the safe candidate set according to the best performance mode, or the model with the lowest estimated cost in the cost priority mode, or the model with the lowest estimated latency in the speed priority mode.
11. The method according to claim 1, characterized in that, Also includes: Record routing logs for each routing request. The routing logs include the selected model identifier, operating mode, difficulty level, quality matching score, number of input tokens, number of output tokens, first token delay, output speed, execution result, and user feedback. Update model service attributes, refresh model profiles, adjust routing thresholds, and expand the evaluation question bank based on the routing logs.
12. A routing system based on automated evaluation of model attribute capabilities, characterized in that, include: An offline model profiling construction subsystem, comprising: The dynamic assessment question bank module is used to store assessment questions organized by ability dimension, difficulty level and task type, as well as the question ability weight vector of each question; The model evaluation execution module is signal-connected to the dynamic evaluation question bank module and is used to call the model to be evaluated to generate answers to the evaluation questions in the dynamic evaluation question bank module; The multi-dimensional ability scoring module is signal-connected to the model evaluation execution module and is used to output the answer validity, answer relevance of each ability dimension and ability score of the model to be evaluated. The model ability score calculation module is signal-connected to the multi-dimensional ability scoring module and is used to calculate the ability score of the model to be evaluated under each ability dimension and difficulty level based on the question ability weight vector, the answer relevance and the ability score. The model profile and grading module is signal-connected to the model capability score calculation module, and is used to generate a model profile containing capability grade tags and service attributes based on the capability score, and store the model profile in the model profile database. The online task profiling and routing decision subsystem includes: The routing preprocessing module is used to perform structured detection on user requests and output the structured detection results. The task profile generation module is signal-connected to the routing preprocessing module and is used to generate a task profile containing the requirement weights, difficulty levels and service requirements of each capability dimension based on the user request and the structured detection results. The hard filtering module is signal-connected to the task profile generation module and is used to perform hard filtering on candidate models in the model profile database according to the hard constraints in the task profile, and remove models that do not meet the hard constraints. The gear matching module is signal-connected to the hard filtering module and is used to perform gear matching on the candidate models that have passed the hard filtering according to the difficulty level and the main ability dimension of the task profile, and to determine the set of candidate models within the candidate gear. The quality matching scoring module is signal-connected to the gear matching module and is used to calculate the quality matching score for each candidate model in the candidate model set. The operation mode selection module is signal-connected to the quality matching and scoring module, and is used to select the final model from the safety candidate set according to the operation mode; The model profile database is communicatively connected to the model profile and grading module and the hard filtering module, respectively, and is used to store the model profiles of all candidate models.
13. The system according to claim 12, characterized in that, It also includes a service attribute monitoring module, which is communicatively connected to the model profile database. The service attribute monitoring module is used to continuously update the input price, output price, cache read / write price, first token delay, output speed, high percentile delay, context window, tool support flag, multimodal support flag, sensitive data processing permissions, and current availability status of each candidate model, and write the updated service attributes into the model profile database.
14. The system according to claim 12, characterized in that, It also includes a feedback calibration module, which is communicatively connected to the operation mode selection module and the dynamic evaluation question bank module, respectively. The feedback calibration module is used to record the routing log of each routing request, and update the model service attributes, refresh the model profile, adjust the routing threshold and expand the evaluation questions in the dynamic evaluation question bank module according to the routing log.
15. The system according to claim 12, characterized in that, The routing preprocessing module is also used to perform long context preprocessing on the user request when it detects that the number of input tokens in the user request exceeds a preset threshold, generate a session summary, document structure features, related fragment summary and compressed global summary, and output the compressed global summary to the task profile generation module.
16. The system according to claim 12, characterized in that, It also includes a capability dimension system module, which is signal-connected to the dynamic evaluation question bank module and the task profile generation module, respectively, and is used to define the capability dimension space shared by the offline model profile construction subsystem and the online task profile and routing decision subsystem; the capability dimensions include at least one of the following: intent understanding and instruction compliance, general knowledge and factual question answering, logical reasoning and mathematical ability, code and software engineering ability, document understanding and long context ability, data and structured information processing, multilingual and translation ability, writing and style control, tool invocation and proxy execution ability, security compliance and rejection boundary.
17. The system according to claim 12, characterized in that, The model profiling and grading module adopts a relative sorting model grading principle, grading each capability dimension separately. In any ability dimension, the models are first sorted according to their ability scores for the corresponding expert difficulty level in that ability dimension, and the models that rank in the first preset proportion are marked as expert models in that ability dimension. After removing models that have been marked as expert models, the remaining models are sorted according to their ability scores for the corresponding complexity level in this ability dimension, and the models that rank in the second-to-last preset proportion are marked as strong models in this ability dimension. After removing models that have been marked as expert models and strong models, the remaining models are sorted according to their ability scores at the corresponding standard difficulty level in this ability dimension. The models that rank in the top three of the preset proportion are marked as standard models in this ability dimension. The remaining models are labeled as weak models in this capability dimension.
18. An electronic device, characterized in that, The method includes a processor and a memory, the memory being electrically connected to the processor, the memory storing computer program instructions, and the processor executing the computer program instructions to implement the method of any one of claims 1 to 11.
19. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is loaded and executed by the processor, it implements the method according to any one of claims 1 to 11.
Citation Information
Patent Citations
Large language model intelligent routing method and device, electronic equipment and storage medium
CN121882260A