A heterogeneous large model interface aggregation management and control and intelligent routing system
Patent Information
- Application Number
- CN202610970108.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-01
- Publication Date
- 2026-09-25
AI Technical Summary
[0003]现有异构大模型接口聚合技术多采用简单的字段映射方式实现接口适配,对不同能力等级模型的可观测性差异支持不足,导致后续路由机制难以统一适配各类模型;同时,现有智能路由方案多在请求发起阶段完成模型选择,或仅在接口发生明确失败时触发备用调用,无法在模型生成过程中实时感知认知不确定性并动态调整调用路径;此外,现有技术缺乏对生成内容中不确定片段的精准定位能力,也未建立基于不确定性等级的分级资源调度机制,容易出现资源冗余浪费或输出质量不稳定的情况
1、针对现有技术接口适配性差且缺乏生成过程风险感知的问题,本发明采用统一接入与异构适配模块,通过字段映射表将不同业务端请求转换为统一请求对象,并将各模型返回结果转换为统一模型事件,同时引入二进制位掩码形式的可观测能力标记,使同一套路由机制能够适配不同能力等级的大模型接口;结合认知不确定性探测与片段定位模块,按滑动窗口提取多维度不确定性子指标,计算窗口级认知不确定性指数并定位不确定片段,能够精准识别生成过程中的风险区域,为后续动态路由决策提供依据;
Smart Images

Figure CN122816936A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence large model application technology, and more specifically, it relates to a heterogeneous large model interface aggregation management and intelligent routing system. Background Technology
[0002] With the rapid development of large-scale artificial intelligence model technology, the mixed invocation of various types of heterogeneous large models has become a common requirement for various business systems, and the aggregation management and intelligent routing technology of large model interfaces have also become a research focus.
[0003] Existing heterogeneous large model interface aggregation technologies mostly use simple field mapping to achieve interface adaptation, which is insufficient in supporting the observability differences of models with different capability levels, making it difficult for subsequent routing mechanisms to uniformly adapt to various models. At the same time, existing intelligent routing solutions mostly complete model selection at the request initiation stage, or only trigger backup calls when an interface fails explicitly, failing to perceive and recognize uncertainty in real time during model generation and dynamically adjust the call path. In addition, existing technologies lack the ability to accurately locate uncertain segments in the generated content, and have not established a hierarchical resource scheduling mechanism based on uncertainty level, which can easily lead to resource redundancy and waste or unstable output quality.
[0004] To address the aforementioned issues, this invention proposes a heterogeneous large model interface aggregation management and intelligent routing system. Summary of the Invention
[0005] To address the shortcomings of existing technologies, the present invention aims to provide a heterogeneous large model interface aggregation management and intelligent routing system.
[0006] To achieve the above objectives, the present invention provides the following technical solution: A heterogeneous large model interface aggregation management and intelligent routing system includes a unified access and heterogeneous adaptation module, a request profiling and model profiling module, a cognitive uncertainty detection and fragment localization module, an intelligent routing control and shadow switching module, and an upgrade and downgrade resource scheduling and result management module. The unified access and heterogeneous adaptation module is used to verify and preprocess model call requests, convert different business requests into unified request objects, and convert model return results into unified model events through interface adapters corresponding to each accessed model. The request profiling and model profiling modules are used to generate request profiling based on a unified request object, and to generate model profiling based on model registration information, interface health detection results, and system call trajectories. The cognitive uncertainty detection and fragment localization module is used to process unified model events according to a sliding window, extract uncertainty signals and locate uncertain fragments; The intelligent routing control and shadow switching module is used to select the primary model, backup model and verification model according to the request profile, model profile and cognitive uncertainty index, and to perform shadow routing, switching or arbitration during the generation process; The upgrade and downgrade resource scheduling and result management module is used to schedule model resources, verification resources and output buffers according to the uncertainty level, and to splice, verify and record the trajectory of the final output results.
[0007] Furthermore, the unified access and heterogeneous adaptation module receives a model invocation request that includes user input content, multi-turn context, system prompt content, output format constraints, tool invocation constraints, cost strategy, latency strategy, risk strategy, and streaming output flag. It then performs identity verification, tenant permission verification, input length verification, parameter integrity verification, and sensitive field preprocessing on the model invocation request to form a unified request object that includes at least the request identifier, tenant identifier, user identifier, session context, current input content, input modality, output format requirements, scope of callable tools, streaming output flag, budget strategy, risk strategy, and request timestamp.
[0008] Furthermore, the interface adapter includes a field mapping table, which stores the correspondence between unified request fields and target model-specific fields. When the target model does not support a field in the unified request object, the interface adapter ignores the field. When the target model does not support candidate token probability, visible inference summary, or structured output, the interface adapter writes an observable capability flag in the unified model event. The observable capability flag uses a binary bitmask to represent the model capability support status.
[0009] Furthermore, the request profiling and model profiling module parses the current input content, context, output requirements, and business strategies to generate a request profiling that includes task type, task complexity, risk level, output strictness, context length, real-time preference, cost preference, and quality preference. The task type is determined by a rule parser, a lightweight classification model, or a combination of both. When the recognition results of the rule parser and the lightweight classification model are inconsistent, the task label with higher risk is selected as the primary label according to the risk strategy, and the other task label is recorded as the auxiliary label.
[0010] Furthermore, the model profile includes a static capability profile and a dynamic operational profile. The static capability profile includes the model vendor, model type, context window capability, streaming output support status, candidate token probability support status, visible inference summary support status, structured output support status, tool call support status, supported input modalities, applicable task tags, security level, regional restrictions, concurrency restrictions, and billing method. The dynamic operational profile includes the recent first token response status, average generation status, interface error status, rate limiting status, historical performance under the task type, user feedback results, current load, and health status. After each call, the dynamic profile score under the corresponding task type is updated using a sliding average method.
[0011] Furthermore, the cognitive uncertainty detection and fragment localization module divides the sliding window according to the number of tokens, sentence boundaries, paragraph boundaries, code block boundaries, or structured field boundaries, and extracts uncertainty sub-indicators from self-doubt semantics, probability entropy, semantic conflict, self-consistency divergence, generation anomaly, and interface stability for each window; wherein, when the observability flag in the unified model event indicates that the candidate token probability is unobservable, the weight of the sub-indicator corresponding to the probability entropy is reset to invalid, and the weights of the remaining valid sub-indicators are renormalized.
[0012] Furthermore, the cognitive uncertainty detection and fragment localization module also calculates the candidate fragment risk score based on the window-level cognitive uncertainty index, the persistence of similar uncertain causes, and the task criticality of the content covered by the window. When the candidate fragment risk score reaches the fragment threshold, the corresponding interval is marked as an uncertain fragment, and adjacent uncertain fragments are merged into the same fragment when the interval between them meets the merging condition. The merged fragment is extended to the nearest natural boundary and a fragment description object is generated.
[0013] Furthermore, the intelligent routing control and shadow switching module performs hard filtering on candidate models based on request profiles and model profiles, and determines the initial routing score based on task capability matching degree, historical output stability, current availability, format support, security policy matching degree, expected cost, expected response burden, recent error risk, and rate limiting risk. The module selects the model that meets the hard filtering conditions as the main model, and selects backup and verification models based on the capability gaps of the main model. The hard filtering conditions include insufficient context capabilities, tenant permission mismatch, current model unavailable, unsupported input modality, unsupported necessary output format, unmet regional policy, unmet security level, and unmet budget policy.
[0014] Furthermore, the intelligent routing control and shadow switching module employs a state machine for control. This state machine includes a main model output state, a shadow routing state, a switching preparation state, a backup model takeover state, an arbitration state, and a completion state. When the cognitive uncertainty index reaches the shadow routing threshold and the continuous window triggering condition is met, the system enters the shadow routing state, where the backup model generates a shadow stream based on the continuation context. When the cognitive uncertainty index reaches the switching threshold and the shadow stream meets the continuation condition, the system enters the switching preparation state, determining the switching boundary and performing prefix alignment on the backup model output. When the cognitive uncertainty index reaches the arbitration threshold, or when the main model and backup model conflict on key conclusions, key values, code logic, contract terms, or tool parameters, the system enters the arbitration state.
[0015] Furthermore, the upgrade / degrade resource scheduling and result management module divides the current execution state into low uncertainty, medium uncertainty, high uncertainty, and extremely high uncertainty. In the low uncertainty state, the current main model continues to be used, unnecessary shadow streams are canceled, the verification frequency is reduced, and the stable buffer content is released. In the medium uncertainty state, shadow routing is started, the generation length of the backup model is limited, and the positioning frequency of uncertain fragments is increased. In the high uncertainty state, the backup model is switched to, the verification model is started, the output buffer is expanded, and the release of unstable fragments is paused. In the extremely high uncertainty state, arbitration, conservative output, or clarification processes are entered. The upgrade / degrade resource scheduling and result management module also records the running trajectory and updates the initial routing weight and backup model priority in the model profile based on the running trajectory.
[0016] Compared with the prior art, the present invention has the following beneficial effects: 1. To address the issues of poor interface adaptability and lack of risk awareness in the generation process in existing technologies, this invention employs a unified access and heterogeneous adaptation module. This module converts requests from different business ends into a unified request object through a field mapping table and converts the return results of each model into a unified model event. Simultaneously, it introduces observable capability markers in the form of binary bitmasks, enabling the same routing mechanism to adapt to large model interfaces with different capability levels. Combined with a cognitive uncertainty detection and fragment localization module, it extracts multi-dimensional uncertainty sub-indicators using a sliding window, calculates the window-level cognitive uncertainty index, and locates uncertain fragments. This allows for accurate identification of risk areas in the generation process, providing a basis for subsequent dynamic routing decisions. 2. To address the issues of static routing and unreasonable resource scheduling in existing technologies, this invention employs an intelligent routing control and shadow switching module. Based on request profiles and model profiles, it completes initial route screening and scoring. During the generation process, a state machine sequentially triggers shadow routing, switching preparation, standby takeover, or arbitration states according to the uncertainty index, achieving dynamic path adjustment during the generation process. Simultaneously, through an upgrade / downgrade resource scheduling and result management module, the execution state is divided into four uncertainty levels and corresponding to different resource scheduling strategies. This helps to optimize the efficiency of system resource utilization and reduce unnecessary redundant calls while ensuring output quality. Attached Figure Description
[0017] Figure 1 This is a block diagram of a heterogeneous large model interface aggregation management and intelligent routing system. Figure 2 This is a flowchart illustrating the implementation of the cognitive uncertainty detection and fragment localization module of the present invention. Figure 3 This is a flowchart illustrating the implementation of the intelligent routing control and shadow switching module of the present invention. Detailed Implementation
[0018] Example, refer to Figure 1 This embodiment of a heterogeneous large model interface aggregation management and intelligent routing system includes a unified access and heterogeneous adaptation module, a request profile and model profile module, a cognitive uncertainty detection and fragment localization module, an intelligent routing control and shadow switching module, and an upgrade and downgrade resource scheduling and result management module. Mod1, Unified Access and Heterogeneous Adaptation Module: This module is located at the system entry point and is responsible for request standardization and model interface difference shielding. Specifically, it converts requests from different business ends into a unified request object and converts data returned by different model interfaces into a unified model event for subsequent modules to process directly.
[0019] The module receives model invocation requests submitted by clients, business systems, or third-party applications. The requests include user input content, multi-turn context, system prompts, output format constraints, tool invocation constraints, cost strategies, latency strategies, risk strategies, and streaming output flags, etc. The system performs identity verification, tenant permission verification, input length verification, parameter integrity verification, and sensitive field preprocessing on requests. Identity verification is achieved by parsing the access token carried in the request and matching it with the set of valid tokens stored in the system. Tenant permission verification queries the pre-configured permission list based on the tenant identifier to verify whether the tenant has the permission to call the target model. Input length verification limits the number of characters or tokens in the input content based on the context window capability of different models. Parameter integrity verification checks whether the required fields in the request exist and are in a valid format. Sensitive field preprocessing filters or de-identifies personal privacy information and sensitive content contained in the request through regular expression matching and semantic recognition. Through the above processing, the original request is converted into a unified request object, which includes at least the request identifier, tenant identifier, user identifier, session context, current input content, input modality, output format requirements, scope of callable tools, streaming output flag, budget strategy, risk strategy, and request timestamp. The module also configures an interface adapter for each connected model. The interface adapter converts the unified request object into the target model interface through a field mapping table. The field mapping table stores the correspondence between the unified request fields and the target model-specific fields, including model identifier, context message, sampling parameters, maximum output length, whether to enable streaming output, whether to return candidate token probability, whether to enable tool calls, and whether to return visible inference summary. For fields that the target model does not support, the interface adapter automatically ignores the field and does not pass it to the target model. For the model's returned results, the interface adapter converts them into unified model events. A unified model event includes at least the request identifier, model identifier, event type, token sequence number, token text, generation timestamp, candidate token probability information, visible inference summary information, interface status, completion reason, exception flag, and cumulative usage. Event types include three categories: token generation events, completion events, and exception events. Interface status includes four categories: normal, timeout, rate limiting, and error. Completion reasons include four categories: normal termination, length limit, user termination, and abnormal termination. When the target model does not support candidate token probabilities, visible inference summaries, or structured outputs, the interface adapter does not forge them. Instead, it records observable capability flags in the unified model event. The observable capability flags are stored in the form of binary bitmasks, with each bit corresponding to a model capability. A bit value of 1 indicates that the capability is supported, and a bit value of 0 indicates that the capability is not supported. During subsequent cognitive uncertainty calculations, the system adjusts the signals involved in the calculation based on the observable capability flags, so that the same routing mechanism can adapt to large model interfaces with different capability levels.
[0020] Mod2, Request Portrait and Model Portrait Module.
[0021] The request profile describes the task, risks, and constraints of the current request, while the model profile describes the capabilities, status, and adaptability of each model. Together, they determine the initial route and subsequent backup paths.
[0022] The request profile is generated by a unified request object. The module parses the current input content, context, output requirements and business strategies to obtain task type, task complexity, risk level, output strictness, context length, real-time preference, cost preference and quality preference. Task types are determined by a rule parser, a lightweight classification model, or a combination of both. The rule parser identifies task types by matching explicit features in the input content, including code block markers, formula symbols, contract clause formats, function call syntax, structured field identifiers, and image / text input markers. The lightweight classification model takes the semantic vector of the input content as input and outputs the probability distribution of each task type. Task types include question answering, summarizing, code generation, mathematical reasoning, contract analysis, data interpretation, tool calling, or multimodal understanding. If the results of the rule parser and the classification model are inconsistent, the system selects the higher-risk task label as the primary label based on the risk strategy and records the other label as the auxiliary label. The lightweight classification model can be fine-tuned using a pre-trained small language model. In this embodiment, the Distil BERT lightweight pre-trained language model is used as the base model. This model is a general-purpose simplified semantic understanding model that can fully inherit the text semantic extraction capabilities of the base model. It also has the characteristics of being lightweight, having low computational consumption, and adapting to real-time inference, which is fully suitable for the semantic classification scenario requirements of this system. Furthermore, the fine-tuning data comes from the labeled samples in the system's historical call records. Task complexity is determined based on input length, contextual dependency, number of constraints, whether multi-step reasoning is required, whether structured output is required, and whether external tool calls are involved. The scores of each dimension are linearly weighted and fused to obtain the final task complexity score. Risk level is determined based on task domain, output consequences, whether professional judgment is involved, whether compliance requirements are included, and whether the user requires a definitive conclusion. The scores of each dimension are linearly weighted and fused to obtain the final risk level score. Output strictness is determined based on whether the output must meet a fixed format, fixed fields, code syntax, function parameter format, or numerical derivation process. The scores of each dimension are linearly weighted and fused to obtain the final output strictness score. All of the above profile fields can be mapped to dimensionless scores, limited to the range of 0 to 1. The model profile includes a static capability profile and a dynamic operational profile. The static capability profile includes the model vendor, model type, context window capabilities, whether streaming output is supported, whether candidate token probability is supported, whether visible inference summaries are supported, whether structured output is supported, whether tool calls are supported, input modalities are supported, applicable task tags, security level, regional restrictions, concurrency limits, and billing method. The static capability profile is entered by the administrator during model registration, and the system periodically verifies the accuracy of the static capabilities. The dynamic operational profile includes the recent first token response status, average generation status, interface error status, rate limiting status, historical performance under task type, user feedback results, current load, and health status. Model profiling does not rely on external knowledge bases; its data comes from model registration information, interface health check results, and the system's own call history. Interface health checks are implemented by periodically sending heartbeat requests to the model. These heartbeat requests use simple text input to detect the model's response status and response time. For short-term anomalies such as interface rate limiting, network fluctuations, or service unavailability, the system updates the temporary health status in real time. For cases where a model performs stably or unstably in a specific task type over a long period, the system updates the dynamic profiling score for the corresponding task type after each call. The update method uses a moving average, retaining the weight of recent call results higher than that of older call results.
[0023] Mod3, Cognitive Uncertainty Detection and Fragment Localization Module.
[0024] like Figure 2 As shown, uncertainty signals are extracted from streaming output, probability distribution, visible inference summary, format constraints and semantic consistency, and the overall uncertainty is further located to specific text fragments.
[0025] The system handles unified model events using sliding windows. Windows are divided based on the number of tokens, sentence boundaries, paragraph boundaries, or structured field boundaries. Regular text can use longer windows, while code, contracts, formulas, function calls, and structured output can use finer-grained windows. Window size and sliding step can be dynamically adjusted according to task type. In some implementations, the size of a regular text window can be set to a number of tokens, and the sliding step can be set to a certain percentage of the window size. The size of a code and structured output window can be set to a single statement or a single field, and the sliding step can be set to 1. Each window is bound to the token number, character offset, sentence boundary, paragraph boundary, code block boundary, or structured field boundary it covers. The module first calculates a self-doubt semantic score. The system identifies semantic components within the window, such as negation, rollback, re-examination, conditional assumptions, uncertainty, path restart, conclusion retention, and pre- and post-correction. The identification method uses a combination of rule sets, lightweight semantic classifiers, and context position judgment. The system does not only count the number of times a word appears, but also judges whether the expression affects key content. Uncertain expressions appearing in conclusions, numerical values, code logic, contract terms, function parameters, or structured fields have higher weights; expressions appearing in polite instructions, scope hints, or non-key transitional phrases have lower weights. The self-doubt semantic score is calculated using the following formula: ; in, Indicates the first The semantic score of self-doubt in each window; Indicates the first A collection of sentence elements or tokens within a window; Indicates the first The sentence element in the first The intensity of uncertain expression in each window is determined by the rule set and the output of the semantic classifier. The rule set outputs the base intensity value, and the semantic classifier outputs the correction coefficient. The two are multiplied to obtain the final intensity value, which is limited to the range of 0 to 1. This indicates that the sentence element is relative to the request. The position weight is determined based on whether the sentence element is located in the conclusion, numerical value, code logic, tool parameter, structured field, or factual assertion position. The weight of key positions can be set to a higher value, and the weight of non-key positions can be set to a lower value. The denominator represents the sum of all position weights within the window, and the system ensures that it is not zero before calculation. When the model returns the candidate token probabilities, the system further calculates the probability entropy score: ; in, Indicates the first The probability entropy score of each window; Indicates the number of tokens used in the probability entropy calculation; Indicates the first The number of candidate tokens available at each token position; Indicates the first The token position is the first The probability of each candidate token is determined either directly by the model interface or by normalizing the candidate scores. Normalization is achieved by dividing the scores of all candidate tokens by the sum of their total scores. If the number of candidates for a given token position is insufficient to calculate the entropy value, that token position is excluded from the calculation. Calculation; through After normalization, a higher entropy value indicates a more significant divergence in the model's choice at that position; For models that do not return candidate probabilities, the system resets the probability entropy sub-index weights to invalid and supplements the uncertainty judgment through semantic conflict, format deviation, self-correction frequency and generation of abnormal signals. The semantic conflict score is determined by the consistency between the current output and the user request, system constraints, the output stable prefix, and the output format. It is obtained by linearly weighting the inconsistency probability output by the natural language inference classifier, the number of errors output by the structured format parser, the number of errors output by the code syntax parser, and the number of violations output by the rule checker. The score is limited to the range of 0 to 1. The self-consistency divergence score is determined by the difference in conclusions between the current segment of the main model and the limited-length sampled segment or the candidate segment of the backup model. It is obtained by calculating the semantic similarity between the segments. The lower the similarity, the higher the self-consistency divergence score. The score is limited to the range of 0 to 1. The generation anomaly score is determined by events such as first token response anomaly, sudden change in generation speed, interface error, output interruption, security flag triggering, and repeated correction of tool parameters. Domain experts assign different scores according to the severity of the event, with the score limited to the range of 0 to 1. The interface stability score is used to characterize the real-time operational reliability of the current model interface within the streaming generation window period. It is obtained by normalizing multiple operational characteristics based on the continuous response status of the interface within the window, instantaneous rate limiting fluctuations, frequency of abnormal errors, and network link transmission jitter. The higher the frequency of interface fluctuations, errors, and interruptions, the higher the corresponding interface stability score, with the score limited to the range of 0 to 1. The sub-indicators are combined into a cognitive uncertainty index: ; in, Indicates the first The cognitive uncertainty index of a window; Indicates the first The normalized uncertainty sub-indices can be derived from self-doubting semantics, probability entropy, semantic conflict, self-consistency divergence, generation anomaly, and interface stability. Indicates a request Next The weights of each sub-indicator are determined by the task type, risk level, output strictness, and model observability. The weights of semantic conflict and self-consistency divergence can be appropriately increased in high-risk tasks, and the weights of format deviation can be appropriately increased in high-output strictness tasks. Indicates the number of sub-indicators involved in the integration; This means that the fusion result is limited to the range of 0 to 1. If a certain type of signal is unobservable, its weight is not included in the calculation, and the remaining effective weights are renormalized. In obtaining Then, the fragment localization logic maps window-level uncertainty to a specific output region, and the system calculates a fragment risk score for each candidate fragment: ; in, Indicates from position Arrive at the location Candidate fragment risk score; This represents the set of windows that cover the candidate segment; This indicates the number of windows in the window set. Indicates the first The cognitive uncertainty index of a window; The persistence coefficient is obtained by normalizing the occurrence of the same uncertain cause consecutively. The more consecutive occurrences, the higher the persistence coefficient, which is limited to the range of 0 to 1. The criticality coefficient of a task is determined by whether the window covers the conclusion, numerical value, clause, function signature, conditional judgment, structured field, or tool parameter. The criticality coefficient of a task in a critical area can be set to a higher value, while that in a non-critical area can be set to a lower value, limited to the range of 0 to 1. when When the segment threshold is reached, the system marks the interval as an uncertain segment. The segment threshold is dynamically adjusted according to the request risk level. The segment threshold for high-risk tasks can be set to a lower value. If the interval between adjacent uncertain segments is less than the system's preset interval, they are merged into the same segment. The preset interval can be set according to the text type. The preset interval for ordinary text can be set to several tokens, while the preset interval for code and structured output can be set to a single statement or a single field. The merged segment is extended to the nearest natural boundary, which includes the end position of a sentence, the end position of a paragraph, the end position of a code block, the end position of a structured field, the boundary of a function parameter, or the boundary of an inference step. Each uncertain segment corresponds to a segment description object, which includes at least the start and end positions, overlay text, triggering reason, segment risk score, whether it has been sent to the client, whether replacement is allowed, whether a backup model correction is needed, whether a verification model review is needed, and whether it affects the final conclusion. Triggering reasons include self-doubt semantics, excessively high probability entropy, semantic conflict, self-consistency divergence, generation anomalies, etc. For segments that have not yet been sent to the client, the system can directly replace or delete them in the output buffer. For segments that have already been sent to the client, the system corrects the bias through subsequent continuation text and avoids resending stable prefixes.
[0026] Mod4, Intelligent Routing Control and Shadow Switching Module.
[0027] like Figure 3 As shown, not only is the model selected at the beginning of the request, but the call path is also changed in real time during the generation process based on uncertainty, fragment position, model status and resource strategy.
[0028] The intelligent routing control first performs initial routing, and then performs hard filtering on candidate models based on request profile and model profile. Hard filtering conditions include insufficient context capabilities, mismatched tenant permissions, model currently unavailable, unsupported input modality, unsupported necessary output format, unmet regional policy, unmet security level, and unmet budget policy. The filtered candidate models then enter the scoring stage. The initial route score is calculated using the following formula: ; in, Representation Model In response to the request Initial route score; Indicates the first Positive metrics include task capability matching degree, historical output stability, current availability, format support degree, and security policy matching degree. Each positive metric is normalized to the range of 0 to 1. Indicates the first The inverse indicators include projected costs, projected response burden, near-term error risk, and flow restriction risk. Each inverse indicator is normalized to the range of 0 to 1. Indicates the first The weight of positive indicators is determined by the quality preference of the request. The higher the quality preference, the higher the weight of task capability matching degree and historical output stability. Indicates the first The weights of the reverse-index are determined by the cost preference and real-time preference of the request. The higher the cost preference, the higher the weight of the expected cost; the higher the real-time preference, the higher the weight of the expected response burden. and These represent the number of positive and negative indicators, respectively. Before being entered into the formula, all the above indicators are normalized to dimensionless scores. All weights are non-negative and normalized, so that different indicators can be compared under the same dimension. The model with the highest score and meeting the hard requirements is selected as the main model. At the same time, backup and validation models are selected based on the capability gaps of the main model. The backup model is not required to be the same type as the main model, but rather to complement the uncertainty risk. If the main model is not stable enough in the inference task, the backup model with stronger inference ability is preferred. If the health of the main model interface declines, the backup model with a more stable health status is preferred. If the main model may deviate from the format, the backup model with stronger structured output ability is preferred. The validation model is selected based on the model with the most stable historical performance in this task type. The system also determines an uncertainty threshold based on request risk and output strictness: ; in, Indicates a request The The uncertainty thresholds are, in order, the shadow routing threshold, the handover threshold, and the arbitration threshold; Indicates the first The basic threshold is determined by the system policy configuration, and the basic threshold can be dynamically adjusted according to the overall system operation. and These represent the lower and upper limits allowed for the threshold at this level, respectively, to prevent routing failures caused by thresholds that are too high or too low. This indicates a request for a normalized risk level value; This indicates that the output is a strict normalized value; Indicates the risk level for the first The shrinkage coefficient of the risk level threshold; the higher the risk level, the more significant the threshold shrinkage. Indicates the strictness of the output for the first... The threshold contraction coefficient is calculated based on the level of strictness; the higher the output stringency, the more pronounced the threshold contraction. The system performs sequential correction after threshold calculation to ensure... Maintain the triggering relationship from low to high; Intelligent routing control is implemented using a state machine, which includes at least the main model output state, shadow routing state, switchover preparation state, standby model takeover state, arbitration state, and completion state; the system according to The system performs state transitions based on uncertain fragment descriptions, model health status, output buffer status, and budget status. To avoid frequent switching, the system sets continuous window trigger conditions, minimum dwell time, switching cooldown conditions, and maximum number of switching. The continuous window trigger condition is set to a consecutive window period. The threshold is exceeded; the minimum dwell time can be set to the minimum time that the model must remain running after a model switch; the switch cooldown time can be set to the minimum interval between two switches; the maximum number of switches can be set to the upper limit of model switches allowed in a single request. when Below When the system maintains the main model output, it allows for upgrade and downgrade resource scheduling to reduce unnecessary verification frequency; when achieve When the continuous window condition is met, the system enters the shadow routing state; in this state, the main model continues to be generated, and the backup model starts in the background, but its output is not directly sent to the user. When the shadow route starts, the system constructs a connection context, which includes the original user request, necessary session context, stable output prefix, description of uncertain segments, output format constraints, control instructions to prevent duplicate stable prefixes, the position that needs to be corrected or continued, the current risk level, and the expected connection method. The backup model generates a shadow stream based on this connection context. The shadow stream is only used by the system to determine takeover capability and is not directly displayed to the user. The system continuously evaluates the shadow stream's first token response status, generation continuity, prefix consistency, format compliance, ability to correct uncertain segments, and its own cognitive uncertainty. The shadow stream evaluation calculates scores for each of these dimensions and linearly weights them to obtain a comprehensive evaluation score. If the comprehensive evaluation score reaches a preset standard, the shadow stream is considered to meet the continuation conditions; if the main model subsequently... Upon returning to a safe range, the system cancels the shadow stream and releases spare resources, if the main model... achieve Furthermore, the shadow stream meets the continuation conditions, and the system enters the switching preparation state; In the switch preparation state, the system first determines the switch boundary, and the last stable natural boundary is selected as the switch boundary first to avoid cutting from the middle of sentences, code statements, structured fields or function parameters; the content before the switch boundary is retained as a stable prefix, and the buffer content after the switch boundary that has not yet been sent is marked as content to be replaced; Subsequently, the system performs prefix alignment on the output of the backup model. If the backup model repeatedly generates a stable prefix, the system deletes duplicate content based on the longest common semantic fragment, sentence similarity, and structural boundary matching, retaining only contiguous fragments. The longest common semantic fragment is determined by calculating the semantic vector similarity between two text fragments; sentence similarity is calculated using cosine similarity; structural boundary matching is achieved by recognizing natural boundary markers in the text. If the backup model corrects an uncertain fragment, and the corrected content satisfies format constraints and semantic consistency constraints, then the corrected fragment is adopted. If the output of the backup model conflicts with the stable prefix in key conclusions, the system does not perform a direct switch but enters an arbitration state. when achieve If the main model and the backup model conflict on key conclusions, key values, code logic, contract terms, or tool parameters, the system enters arbitration mode. In arbitration mode, the verification model or validator performs consistency, format, risk, and security checks on the candidate outputs. Consistency checks are performed by comparing the similarity of key content among multiple candidate outputs. Format checks are performed by a structured parser to verify whether the output conforms to the specified format. Risk checks are performed by a content security model to check whether the output contains harmful content. Security checks are performed by permission checks to verify whether the output conforms to the security policy. If the detection results still cannot support stable output, the system enters clarification processing, requesting necessary conditions from the user instead of forcibly generating a definitive conclusion. After the switch is completed, the client still receives streaming responses through the same request identifier. The system can generate short transition phrases to ensure semantic coherence, but does not expose the model name, supplier information or underlying switch details; the transition phrases only serve a connecting function and are not used as factual evidence. Through the above processing, intelligent routing is no longer just a backup call after interface failure, but proactively plans backup paths based on cognitive uncertainty in the early stages of model generation, and selects a more suitable takeover model according to the type of uncertain fragments; format deviation is prioritized for routing to models with stronger structured output capabilities, code logic risks are prioritized for routing to code models or code verification processes, inference instability is prioritized for routing to inference models, and interface anomalies are prioritized for routing to peer models with better health status.
[0029] Mod5, the module for upgrading and downgrading resource scheduling and result management.
[0030] By dynamically controlling model resources, verification resources, and output strategies based on the uncertainty level, the system can improve reliability in complex tasks and reduce redundant calls in stable tasks.
[0031] The system categorizes the current execution state into low uncertainty, medium uncertainty, high uncertainty, and extremely high uncertainty; low uncertainty corresponds to... States below the shadow routing threshold; medium uncertainty corresponds to The state where the shadow routing threshold has been reached but the handover threshold has not been reached; high uncertainty corresponds to... The state where the switching threshold has been reached but the arbitration threshold has not been reached; extremely high uncertainty corresponds to... The arbitration threshold has been reached or a conflict of key conclusions has occurred. Under low uncertainty, the system continues to use the current main model, cancels unnecessary shadow streams, reduces the verification frequency, and releases the contents of the stable buffer. Under medium uncertainty, the system starts shadow routing, limits the generation length of the backup model, and increases the frequency of locating uncertain fragments. Under high uncertainty, the system can switch to a more reliable model, start the verification model, expand the output buffer, and pause the release of unstable fragments. Under extremely high uncertainty, the system enters the arbitration, conservative output, or clarification process. The scheduling action is selected according to the following formula: ; in, Indicates the first The scheduling action can be selected from a single window; Indicates the current level of uncertainty. The set of actions that can be selected below; Indicates a request The preference weight vector for the improvement effects of cost, latency, stability and reliability is a dimensionless quantity and is normalized. The order of the components is cost, latency, stability and reliability. Indicates action In the The cost vector under each window comprises normalized incremental cost, normalized expected delay, normalized output disturbance risk, and reliability improvement inverse vector, each component being normalized to the range of 0 to 1; (symbol) This represents the vector inner product; cost-priority scenarios increase the weight of the cost component, low-latency scenarios increase the weight of the latency component, and high-risk tasks increase the weight of the stability and reliability improvement component. In the results management stage, if only the main model output is available and the uncertainty remains at a low level, the system performs necessary formatting and finalization on the output before returning it. If there are backup model continuation segments, the system will concatenate the stable prefix, the corrected segment, and the subsequent output, and check for duplicate paragraphs, semantic jumps, format breaks, incorrect numbering, unclosed code blocks, and missing structured fields. If there are multiple candidate answers, the system will prioritize the candidate segment that simultaneously satisfies the request constraints, stable prefix consistency, format constraints, and verification results. If key conclusions cannot be agreed upon, the system will output the definite parts and prompt for the conditions that need to be supplemented. For structured output, the system performs formal repair, which includes filling in parentheses, fixing field separators, validating field types, maintaining continuous numbering, closing code blocks, and conforming to user-specified formats. Filling in parentheses is achieved by traversing the output text, counting the number of left and right parentheses, and filling in missing parentheses. Fixing field separators is achieved by replacing invalid separators with specified separators. Validating field types is achieved by matching field values with expected types. Maintaining continuous numbering is achieved by regenerating the numbering sequence. Closing code blocks is achieved by identifying code block markers and filling in missing closing markers. This repair does not introduce new factual content; it only addresses output format errors. The system also records the execution trajectory, which includes request profile, model profile version, initial route score, primary model, backup model, uncertainty curve, threshold trigger position, uncertainty segment position, shadow route startup time, switching boundary, scheduling action, arbitration result, final output status, interface exception, and user feedback. The execution trajectory is stored locally in a structured format, and the storage period can be set according to the system storage capacity and business needs. The model profile is updated periodically based on the execution trajectory, and the update period can be set to several hours or several days. If a model frequently triggers high uncertainty in a certain type of task or is taken over by a backup model, its initial route weight for that type of task is reduced. If a backup model can reduce uncertainty and pass the consistency check after taking over, its priority as a backup model for that type of task is increased. Weight adjustment adopts an incremental adjustment method, and the adjustment range can be set according to the system strategy. When the main model interface times out, experiences rate limiting, stream interruption, or returns an exception code, the system first checks if there is an available shadow stream. If an available shadow stream exists, the output is continued according to the seamless switching process. If not, a backup model is selected again, and the already stable output content is used as the continuation context. If a security policy is triggered, the system does not bypass the restriction through normal model switching, but instead enters the security processing flow, returning a rejection, downgrade explanation, or clarification request. The security processing flow performs different operations according to the level of the security policy. Low-level security policies can return a downgrade explanation, while high-level security policies can directly return a rejection message.
[0032] Through the detailed description of the above embodiments, the heterogeneous large model interface aggregation management and intelligent routing system of the present invention constructs a complete unified management and dynamic intelligent routing system for heterogeneous large models through the coordinated cooperation of five core modules. The unified access and heterogeneous adaptation module shields the interface differences of different large models at the entry point, realizing standardized processing of requests and results; the request profiling and model profiling modules respectively characterize task requirements and model capabilities, providing basic support for routing decisions; the cognitive uncertainty detection and fragment localization module realizes real-time perception and accurate location of risks during the generation process, breaking through the limitations of traditional routing that only relies on interface states; the intelligent routing control and shadow switching module dynamically adjusts the call path during the generation process through a state machine mechanism, realizing a shift from passive fault detection to proactive risk avoidance; the upgrade / downgrade resource scheduling and result management module dynamically allocates resources according to the uncertainty level and completes result splicing, verification, and trajectory recording, forming a closed-loop operation optimization mechanism. The present invention can adapt to heterogeneous large models with different capability levels, effectively improving the system's operational stability and output reliability in mixed large model call scenarios.
[0033] The preset parameters in the above formulas shall be set by those skilled in the art according to the actual situation.
[0034] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more sets of available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium. A semiconductor medium can be a solid-state drive.
[0035] It should be understood that in the various embodiments of this application, the order of the above-mentioned processes does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0036] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0037] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0038] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0039] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0040] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A heterogeneous large-scale model interface aggregation management and intelligent routing system, characterized in that, It includes a unified access and heterogeneous adaptation module, a request profiling and model profiling module, a cognitive uncertainty detection and fragment localization module, an intelligent routing control and shadow switching module, and an upgrade and downgrade resource scheduling and result management module; The unified access and heterogeneous adaptation module is used to verify and preprocess model call requests, convert different business requests into unified request objects, and convert model return results into unified model events through interface adapters corresponding to each accessed model. The request profiling and model profiling modules are used to generate request profiling based on a unified request object, and to generate model profiling based on model registration information, interface health detection results, and system call trajectories. The cognitive uncertainty detection and fragment localization module is used to process unified model events according to a sliding window, extract uncertainty signals and locate uncertain fragments; The intelligent routing control and shadow switching module is used to select the primary model, backup model and verification model according to the request profile, model profile and cognitive uncertainty index, and to perform shadow routing, switching or arbitration during the generation process; The upgrade and downgrade resource scheduling and result management module is used to schedule model resources, verification resources and output buffers according to the uncertainty level, and to splice, verify and record the trajectory of the final output results.
2. The heterogeneous large model interface aggregation management and intelligent routing system according to claim 1, characterized in that, The unified access and heterogeneous adaptation module receives a model invocation request that includes user input, multi-turn context, system prompts, output format constraints, tool invocation constraints, cost policy, latency policy, risk policy, and streaming output flag. It then performs identity verification, tenant permission verification, input length verification, parameter integrity verification, and sensitive field preprocessing on the model invocation request to form a unified request object that includes at least the request identifier, tenant identifier, user identifier, session context, current input content, input modality, output format requirements, scope of callable tools, streaming output flag, budget policy, risk policy, and request timestamp.
3. The heterogeneous large model interface aggregation management and intelligent routing system according to claim 2, characterized in that, The interface adapter includes a field mapping table, which stores the correspondence between unified request fields and target model-specific fields; when the target model does not support a field in the unified request object, the interface adapter ignores that field. When the target model does not support candidate token probabilities, visible inference summaries, or structured output, the interface adapter writes an observable capability flag in the unified model event. The observable capability flag represents the model capability support status in the form of a binary bitmask.
4. The heterogeneous large model interface aggregation management and intelligent routing system according to claim 1, characterized in that, The request profiling and model profiling modules parse the current input content, context, output requirements, and business strategies to generate a request profiling that includes task type, task complexity, risk level, output strictness, context length, real-time preference, cost preference, and quality preference. The task type is determined by a rule parser, a lightweight classification model, or a combination of both. When the recognition results of the rule parser and the lightweight classification model are inconsistent, the task label with higher risk is selected as the primary label according to the risk strategy, and the other task label is recorded as the auxiliary label.
5. The heterogeneous large model interface aggregation management and intelligent routing system according to claim 4, characterized in that, The model profile includes a static capability profile and a dynamic operational profile. The static capability profile includes the model vendor, model type, context window capability, streaming output support status, candidate token probability support status, visible inference summary support status, structured output support status, tool call support status, supported input modalities, applicable task tags, security level, regional restrictions, concurrency restrictions, and billing method. The dynamic operational profile includes the recent first token response status, average generation status, interface error status, rate limiting status, historical performance under the task type, user feedback results, current load, and health status. After each call, the dynamic profile score under the corresponding task type is updated using a moving average method.
6. The heterogeneous large model interface aggregation management and intelligent routing system according to claim 1, characterized in that, The cognitive uncertainty detection and fragment localization module divides sliding windows according to the number of tokens, sentence boundaries, paragraph boundaries, code block boundaries, or structured field boundaries, and extracts uncertainty sub-indicators from self-doubt semantics, probability entropy, semantic conflict, self-consistency divergence, generation anomaly, and interface stability for each window; wherein, when the observability flag in the unified model event indicates that the candidate token probability is unobservable, the weight of the sub-indicator corresponding to the probability entropy is reset to invalid, and the weights of the remaining valid sub-indicators are renormalized.
7. The heterogeneous large model interface aggregation management and intelligent routing system according to claim 6, characterized in that, The cognitive uncertainty detection and fragment localization module also calculates the risk score of candidate fragments based on the window-level cognitive uncertainty index, the persistence of similar uncertainty causes, and the task criticality of the content covered by the window. When the risk score of a candidate segment reaches the segment threshold, the corresponding interval is marked as an uncertain segment. When the interval between adjacent uncertain segments meets the merging condition, they are merged into the same segment. The merged segment is extended to the nearest natural boundary and a segment description object is generated.
8. The heterogeneous large model interface aggregation management and intelligent routing system according to claim 1, characterized in that, The intelligent routing control and shadow switching module performs hard filtering on candidate models based on request profiles and model profiles. It determines the initial routing score based on task capability matching degree, historical output stability, current availability, format support, security policy matching degree, expected cost, expected response burden, recent error risk, and rate limiting risk. It selects models that meet the hard filtering conditions as the main model and selects backup and verification models based on the capability gaps of the main model. The hard filtering conditions include insufficient context capabilities, tenant permission mismatch, current model unavailable, unsupported input modality, unsupported necessary output format, unmet regional policy, unmet security level, and unmet budget policy.
9. A heterogeneous large model interface aggregation management and intelligent routing system according to claim 8, characterized in that, The intelligent routing control and shadow switching module is controlled by a state machine, which includes the main model output state, shadow routing state, switching preparation state, standby model takeover state, arbitration state, and completion state. When the cognitive uncertainty index reaches the shadow routing threshold and the continuous window triggering condition is met, the system enters the shadow routing state, and the backup model generates the shadow stream based on the continuation context. When the cognitive uncertainty index reaches the switching threshold and the shadow stream meets the continuation conditions, the switching preparation state is entered, the switching boundary is determined, and the output of the backup model is prefix aligned. Arbitration is initiated when the cognitive uncertainty index reaches the arbitration threshold, or when the primary model and the backup model conflict on key conclusions, key values, code logic, contract terms, or tool parameters.
10. A heterogeneous large model interface aggregation management and intelligent routing system according to claim 1, characterized in that, The upgrade / downgrade resource scheduling and result management module divides the current execution state into low uncertainty, medium uncertainty, high uncertainty, and extremely high uncertainty. In the low uncertainty state, the current main model continues to be used, unnecessary shadow streams are canceled, the verification frequency is reduced, and the stable buffer content is released. In the medium uncertainty state, shadow routing is started, the generation length of the backup model is limited, and the positioning frequency of uncertain fragments is increased. In the high uncertainty state, the backup model is switched to, the verification model is started, the output buffer is expanded, and the release of unstable fragments is paused. In the extremely high uncertainty state, arbitration, conservative output, or clarification processes are entered. The upgrade / downgrade resource scheduling and result management module also records the running trajectory and updates the initial routing weight and backup model priority in the model profile based on the running trajectory.