Large language model dynamic selection method for multi-dimensional quantitative scoring
By using a multi-dimensional quantitative scoring method to dynamically select large language model services, the problems of slow task response and resource waste in enterprises have been solved, achieving fast and high-quality task processing and improved user experience.
Patent Information
- Application Number
- CN202511453563.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-13
- Publication Date
- 2026-01-09
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
When enterprises use large language models, the workload is heavy and the response is not immediate, resulting in serious waste of resources. Differences in user needs lead to unreasonable model selection, resulting in slow response and poor performance.
By employing a multi-dimensional quantitative scoring method, the system comprehensively evaluates model capabilities, network latency, geographical location, data security, context length, service costs, current load, and task suitability, and dynamically selects the best model service.
It improved task response speed, reduced waiting time, ensured data security, saved resources, and enhanced user experience and result accuracy.
Smart Images

Figure CN121303199A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of large model technology, and more specifically to a dynamic selection method for large language models with multi-dimensional quantitative scoring. Background Technology
[0002] Many large language model services exist in enterprises, but some problems still exist in actual user use and intelligent agent operation, which hinder the application of large language models in actual business scenarios.
[0003] When the large language model used by the agent is performing a task with a heavy workload, it cannot immediately execute the current task, nor can it reasonably and timely call other large language model services, resulting in a slow response time for the current task. When users use the agent, the difficulty of the questions raised or the tasks to be completed by different users varies. The current common approach is to select the large language model based on the most difficult user scenario, but this will cause unnecessary waste of resources for enterprises. Furthermore, the tasks that a given large language model is good at are relatively fixed, so it will be less effective in solving various problems of different users. Summary of the Invention
[0004] To overcome the aforementioned deficiencies in the prior art, this invention provides a dynamic selection method for large language models with multi-dimensional quantitative scoring. This method scores the available large model services for each enterprise across multiple dimensions, selects the best large model service through quantitative evaluation, and achieves fast, high-quality, and secure completion of dialogue requests. This saves resources and costs for enterprises while improving user experience and satisfaction, thereby solving the problems existing in the aforementioned background technology.
[0005] This invention provides the following technical solution: a dynamic selection method for a large language model with multi-dimensional quantitative scoring, comprising the following steps: Step 1: Label and sort all candidate models. ; Step 2: Evaluate the model's capability and accuracy; Step 3: Rate network latency and geographic location; Step 4: Conduct data security level scoring; Step 5: Perform a supporting context length score; Step Six: Rate the service fees and token prices; Step 7: Score the current service load; Step 8: Score the task domain suitability; Step 9: Obtain the final total score and select the final model.
[0006] Preferably, the scoring of model capability and accuracy specifically involves: Obtain the raw score of model capability, i.e., the model accuracy score; the model accuracy is the model precision rate; assuming the accuracy rate of all candidate models on the benchmark test is denoted as... The formula is expressed as: ;in, Indicates the first The accuracy scores of each candidate model Indicates the first The accuracy of each candidate model. This represents the minimum accuracy among all candidate models. This represents the maximum accuracy among all candidate models; Obtain the parameter size score; record the number of parameters for all candidate models as... The formula is expressed as: ;in, Indicates the first The parameter size score of each candidate model. This represents the maximum number of parameters among all candidate models. This represents the minimum number of parameters among all candidate models; The final comprehensive ability score is obtained as the final score of the model's ability and accuracy, and is expressed as a weighted average: ;in, Indicates the first The comprehensive capability score of each candidate model. Weights representing accuracy.
[0007] Preferably, the scoring of network latency and geographical location specifically involves: Determine the user's geographical location; Collect the geographical locations of the model service nodes; Calculate the geographical distance from the user to each model node; Scoring is obtained by distance normalization.
[0008] Preferably, the distance normalization score is expressed as follows: ;in, Indicates the first Distance scores between candidate models and users This represents the value that is furthest from the user among all candidate models. Indicates the first The distance between each candidate model and the user This represents the closest distance to the user among all candidate models.
[0009] Preferably, the data security level scoring is set to a security level of 1-5. When the model security level is greater than or equal to the user task security level, 1 point is awarded; when the model security level is less than the user task security level, 0 points are awarded. Expressed as a formula: ;in, Indicates the first Safety level scores for each candidate model Indicates the first The model safety level of each candidate model. This indicates the security level of this user task.
[0010] Preferably, the context length score for providing support is expressed by the formula: ;in, Indicates the first The context length of each candidate model supports scoring; Indicates the first The maximum context length supported by each candidate model; This indicates the actual context length required for this user task.
[0011] Preferably, the scoring of the current service load is expressed by the formula: ;in, Indicates the first The service load scores of the candidate models; where the minimum queue size is... The maximum number of people in the queue is ; Indicates the first The number of tasks currently queued in the current model.
[0012] Preferably, obtaining the final total score and selecting the final model specifically involves: ; in, Indicates the first The final total score of the candidate models; , , , , , , These are the corresponding weight coefficients; Indicates the first Token price scores for each candidate model.
[0013] The technical effects and advantages of this invention are as follows: This invention, through steps two through eight, facilitates the measurement of model performance and accuracy based on model parameter scale; dynamically selects available model service nodes based on user geographic location, prioritizing services with closer network distance and faster response speeds to reduce waiting time and improve user experience; automatically matches model services with compliant security levels based on task data sensitivity, prioritizing private and high-security models for highly sensitive data to ensure data compliance and security; considers the maximum number of input and output tokens that the scoring model can handle to ensure that complex tasks and long text inputs are not affected by model context limitations in task completion; monitors the current task queuing and load status of each model service in real time, prioritizing model nodes with sufficient available resources and faster response times to avoid long waiting times; and prioritizes the model that best meets the current task requirements to improve the professionalism and accuracy of the results. Attached Figure Description
[0014] Figure 1 The flowchart shows the dynamic selection method for the large language model of the multi-dimensional quantitative scoring of this invention. Detailed Implementation
[0015] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. In addition, the forms of the various structures described in the following embodiments are merely illustrative. The dynamic selection method of a large language model for multi-dimensional quantitative scoring involved in the present invention is not limited to the structures described in the following embodiments. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0016] like Figure 1 As shown, this invention provides a dynamic selection method for large language models with multi-dimensional quantitative scoring, including the following steps: Step 1: Label and sort all candidate models. The model in question is the large language model. Step 2: Score the model's capabilities and accuracy; this measures the model's performance and accuracy, including the model's parameter size, which directly affects the model's task completion and output quality. Step 3: Score network latency and geographical location; dynamically select available model service nodes based on the user's geographical location, prioritizing services with closer network distance and faster response speed to reduce waiting time and improve user experience; Step 4: Conduct data security level scoring; based on the data sensitivity of the task, automatically match model services with compliant security levels, prioritizing private models with high security levels for highly sensitive data to ensure data compliance and security; Step 5: Perform a supporting context length score; measure the maximum number of input and output tokens the model can handle to ensure that the model's context limitations do not affect task completion when dealing with complex tasks and long text inputs. Step Six: Score the service fees and token prices; dynamically compare the billing methods and token unit prices of each model service, and prioritize the model service with higher cost performance while meeting the needs, to help enterprises save costs; Step 7: Score the current service load; this is used to monitor the current task queuing and load of each model service in real time, and prioritize model nodes with sufficient available resources and faster response to avoid long waiting times. Step 8: Score the task domain suitability; by analyzing the similarity between the user's task content and the domains that each model excels in, prioritize the model that best meets the current task requirements to improve the professionalism and accuracy of the results. Step 9: Obtain the final total score and select the final model.
[0017] In this embodiment, it should be specifically noted that the scoring of model capability and accuracy is linearly normalized according to model ranking or benchmark test scores. For example, the best model is scored as 1 point, the weakest model as 0 points, and models in the middle are linearly distributed proportionally. Authoritative model test scores can be collected, such as evaluation scores of mainstream large models on various authoritative benchmarks from Open LLM Leaderboard (e.g., HuggingFace, lmsys.org). The comprehensive score of model capability is calculated from two parts: model accuracy and parameter size. Appropriate weights can be set for model accuracy and parameter size to represent their proportions. Specifically: Obtain the raw score of model capability, i.e., the model accuracy score; the model accuracy is the model precision rate; assuming that the accuracy of all candidate models on a certain benchmark such as MMLU is known, denoted as . The formula is expressed as: ;in, Indicates the first The accuracy scores of each candidate model Indicates the first The accuracy of each candidate model. This represents the minimum accuracy among all candidate models. This represents the maximum accuracy among all candidate models; Obtain the parameter size score; record the number of parameters for all candidate models as... The formula is expressed as: ;in, Indicates the first The parameter size score of each candidate model. This represents the maximum number of parameters among all candidate models. This represents the minimum number of parameters among all candidate models; The final comprehensive ability score is obtained as the final score of the model's ability and accuracy, and is expressed as a weighted average: ;in, Indicates the first The comprehensive capability score of each candidate model. Weights representing accuracy.
[0018] In this embodiment, it should be specifically noted that the scoring of network latency and geographical location is normalized according to the actual latency time, with the lowest latency recorded as 1 point and the highest latency recorded as 0 points; specifically: Determine the user's geographical location; resolve the user's IP address to obtain the user's approximate city, province, country, and latitude and longitude. Collect the geographical locations of the model service nodes; each model service node needs to record its actual deployment geographical coordinates, such as latitude and longitude; Calculate the geographic distance from the user to each model node; calculate the great circle distance between two points using the Haversine formula or the spherical cosine formula; Scoring is obtained by distance normalization; the nearest service node gets 1 point, the farthest service node gets 0 points, and the rest are obtained by linear interpolation.
[0019] In this embodiment, it should be specifically noted that the Haversine formula is expressed as: ;in, Indicates distance, , and This represents the dimensions of users and model service nodes. , and Indicates longitude. This represents the Earth's radius.
[0020] In this embodiment, it should be specifically noted that the score obtained by distance normalization is represented as follows: ;in, Indicates the first Distance scores between candidate models and users This represents the value that is furthest from the user among all candidate models. Indicates the first The distance between each candidate model and the user This represents the closest distance to the user among all candidate models.
[0021] In this embodiment, it should be specifically noted that the data security level scoring can be set to a security level of 1-5. A score of 1 is awarded if the required security level for the task is greater than the model's security level; otherwise, a score of 0 is awarded. The principle for scoring the data security level is to determine the minimum security level required to process the task based on its content, data type, and degree of sensitivity. The reference standards for scoring the data security level are the company's own data protection standards, domestic and international regulations, and industry practices. The data security level is scored automatically, i.e., by using an NLP model to detect whether the text contains keywords such as sensitive personal information, financial information, health information, and government affairs information. The standards for the user task security level are shown in the table below: grade Typical Tasks Data content Examples of safety requirements 1 (Public) General Questions and Answers No sensitive information No special requirements 2 (Internal) Enterprise Knowledge Base Non-public information of enterprises Internal circulation only 3 (Sensitive) User feedback, customer service Contact information may be involved. Encrypted transmission, log auditing 4 (Restricted) Medical and financial analysis Personal identity / health data Strict encryption and desensitization 5 (Top Secret) Strategy and Politics State / Company Secrets Compliance and Qualifications / Local Deployment The model safety level standards are shown in the table below: grade Deployment Form Compliance and Qualification Applicable Scenarios 1 Public cloud / SaaS none Public data, non-sensitive 2 Public cloud / SaaS Basic encryption, partial compliance General business of enterprises 3 Private Cloud Domestic mainstream compliance (information security level protection) sensitive corporate data 4 Local private deployment Full compliance abroad Industry-level sectors such as finance and healthcare 5 Local isolation deployment High-level compliance + physical isolation Government, military industry, core secrets If the model's security level is greater than or equal to the user's task security level, then 1 point is awarded. If the model security level is less than the user task security level, then 0 points are awarded. Expressed as a formula: ;in, Indicates the first Safety level scores for each candidate model Indicates the first The model safety level of each candidate model. This indicates the security level of this user task.
[0022] In this embodiment, it should be specifically explained that the context length scoring for support is as follows: Obtain parameters from the official materials of the model manufacturer and confirm and save them when the user registers the model; Define the maximum number of input tokens that the model can process in a single run; When registering a model, users need to manually enter or the system can automatically read and register the model parameters. The determination of the required context length for a user task is specifically as follows: Collect task input data, including historical dialogues, documents, and contextual fragments; Use a tokenizer to count tokens in the input content; The maximum input length is calculated, and the number of tokens in the actual input content is the required context length.
[0023] In this embodiment, it should be specifically noted that the context length score for providing support is expressed by the formula: ;in, Indicates the first The context length of each candidate model supports scoring; Indicates the first The maximum context length supported by each candidate model; This indicates the actual context length required for this user task.
[0024] In this embodiment, it should be specifically explained that the scoring of service fees and token prices is as follows: Inverse normalization is applied to prices, with the lowest price being 1 point and the highest being 0 points; When registering a model, users manually enter the price per 1000 tokens, which is saved in the model's metadata. The calculation formula is as follows: ;in, Indicates the first Token price scores for each candidate model; This represents the highest unit price among all models. This represents the lowest unit price among all models. Indicates the first The price of each candidate model.
[0025] In this embodiment, it should be specifically noted that the scoring of the current service load is normalized according to the number of queued tasks, with the lightest load receiving 1 point and the heaviest receiving 0 points. In the model load collection method, the collection entity is the model gateway, which monitors the load information of all registered model service nodes in real time. The collection content is the current number of queued tasks, and the collection method is to actively pull the status of model service nodes every N seconds. The load is quantified; the larger the number of queued tasks, the higher the load and the slower the response. The lowest queued task receives 1 point, and the highest queued task receives 0 points, with the rest being linear interpolation. The model gateway achieves unified management of large language model interfaces, permissions, logs, and keys through a standardized gateway. This enables enterprises to efficiently and uniformly manage large model resources, control model usage, and form a closed loop of access → management → usage → monitoring.
[0026] like Indicates the first The current number of queued tasks for each model is given, where the minimum queue size is... The maximum number of people in the queue is ; expressed as a formula: ;in, Indicates the first The service load scores of each candidate model.
[0027] In this embodiment, it should be specifically explained that the task domain fit scoring uses text similarity, with a perfect match scoring 1 point and an irrelevant match scoring 0 points. During model registration, the model's preferred domain tags or text are registered and saved. When a task arrives, the task domain is automatically determined, and text similarity is calculated using Embedding cosine similarity. The text is then normalized to generate a fit score; a higher score indicates better fit. In model domain registration and saving, the model's preferred domain tags or keywords must be manually entered during registration and stored as a structured tag array. Alternatively, commas can be used to separate the text. Automatic determination is performed during user task domain identification. The task text input by the user is automatically identified using an NLP classification model, which can be employed. The specific steps for scoring task domain fit are as follows: Concatenate the model domain labels into a short text using delimiters; Generate vectors for both the user task text and the model domain text; Calculate the cosine similarity as the fit score; Expressed as a formula: ;in, This is the similarity score between the current task and the model domain, with a value of [value missing]. ; Cosine similarity represents the similarity calculation result between the task content and the model domain, and its value is [value missing]. It can be normalized to .
[0028] In this embodiment, it should be specifically explained that obtaining the final total score and selecting the final model specifically refers to: ; in, Indicates the first The final total score of the candidate models; , , , , , , These are the corresponding weighting coefficients, all ranging from 0 to 1. This embodiment does not impose specific limits on the specific values of the weighting coefficients, which can be set by those skilled in the art within the range of values. Select the one with the highest final total score The candidate models were selected as the final service models.
[0029] In conclusion, the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
[0030] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A dynamic selection method for a large language model with multi-dimensional quantitative scoring, characterized in that: Includes the following steps: Step 1: Label and sort all candidate models. ; Step 2: Evaluate the model's capability and accuracy; Step 3: Rate network latency and geographic location; Step 4: Conduct data security level scoring; Step 5: Perform a supporting context length score; Step Six: Rate the service fees and token prices; Step 7: Score the current service load; Step 8: Score the task domain suitability; Step 9: Obtain the final total score and select the final model.
2. The method for dynamic selection of a large language model with multi-dimensional quantitative scoring according to claim 1, characterized in that: The specific steps for scoring the model's capability and accuracy are as follows: Obtain the raw score of model capability, i.e., the model accuracy score; the model accuracy is the model precision rate; assuming the accuracy rate of all candidate models on the benchmark test is denoted as... The formula is expressed as: ;in, Indicates the first The accuracy scores of each candidate model Indicates the first The accuracy of each candidate model. This represents the minimum accuracy among all candidate models. This represents the maximum accuracy among all candidate models; Obtain the parameter size score; record the number of parameters for all candidate models as... The formula is expressed as: ;in, Indicates the first The parameter size score of each candidate model. This represents the maximum number of parameters among all candidate models. This represents the minimum number of parameters among all candidate models; The final comprehensive ability score is obtained as the final score of the model's ability and accuracy, and is expressed as a weighted average: ;in, Indicates the first The comprehensive capability score of each candidate model. Weights representing accuracy.
3. The method for dynamic selection of a large language model with multi-dimensional quantitative scoring according to claim 2, characterized in that: The specific steps for scoring network latency and geographic location are as follows: Determine the user's geographical location; Collect the geographical locations of the model service nodes; Calculate the geographical distance from the user to each model node; Scoring is obtained by distance normalization.
4. The method for dynamic selection of a large language model with multi-dimensional quantitative scoring according to claim 3, characterized in that: The score obtained by distance normalization is represented as follows: ;in, Indicates the first Distance scores between candidate models and users This represents the value that is furthest from the user among all candidate models. Indicates the first The distance between each candidate model and the user This represents the closest distance to the user among all candidate models.
5. The method for dynamic selection of a large language model with multi-dimensional quantitative scoring according to claim 4, characterized in that: The data security level scoring is set to a security level of 1-5. When the model security level is greater than or equal to the user task security level, 1 point is awarded; when the model security level is less than the user task security level, 0 points are awarded. Expressed as a formula: ;in, Indicates the first Safety level scores for each candidate model Indicates the first The model safety level of each candidate model. This indicates the security level of this user task.
6. The method for dynamic selection of a large language model with multi-dimensional quantitative scoring according to claim 5, characterized in that: The context length score used for support is expressed by the formula: ;in, Indicates the first The context length of each candidate model supports scoring; Indicates the first The maximum context length supported by each candidate model; This indicates the actual context length required for this user task.
7. The method for dynamic selection of a large language model with multi-dimensional quantitative scoring according to claim 6, characterized in that: The scoring of the current service load is expressed by the formula: ;in, Indicates the first The service load scores of the candidate models; where the minimum queue size is... The maximum number of people in the queue is ; Indicates the first The number of tasks currently queued in the current model.
8. The method for dynamic selection of a large language model with multi-dimensional quantitative scoring according to claim 7, characterized in that: The specific steps for obtaining the final total score and selecting the final model are as follows: ; in, Indicates the first The final total score of the candidate models; , , , , , , These are the corresponding weight coefficients; Indicates the first Token price scores for each candidate model.