Multi-level large model scheduling and knowledge base adaptive optimization method and system
By calculating the complexity of tasks and maturity of processing results, dynamically dispatching models of different scales, and optimizing the model and knowledge base based on the processing results, the problems of waste of computing resources and failure to optimize the knowledge base in the existing technology are solved, and more efficient computing and resource utilization are achieved.
Patent Information
- Application Number
- CN202510680266.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-26
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-05-26
AI Technical Summary
The existing technology is difficult to dynamically schedule models of different sizes according to task complexity, resulting in waste of computing resources and high computing costs. The static storage and retrieval methods of the knowledge base cannot be optimized in coordination with the model scheduling mechanism, increasing the frequency of calling the super-large models.
By obtaining the complexity of the pending tasks and the maturity of the test processing results, dynamically schedule models of different sizes, select appropriate models for processing, and optimize relevant small models or knowledge bases based on the processing results, implement adaptive optimization of the model and knowledge base.
It reduces calls to super-large models, reduces computing costs and resource consumption, improves system response efficiency and processing capabilities of small models and knowledge bases, and reduces duplicate calculations.
Smart Images

Figure CN120216149A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and specifically relates to a multi-level large model scheduling and knowledge base adaptive optimization method and system. Background Art
[0002] Currently, ultra-large-scale artificial intelligence models (such as large models with 671B parameters) demonstrate excellent capabilities in handling complex tasks, but their extremely high demand for computing resources leads to a significant increase in computing costs and energy consumption. In practical applications, not all tasks require the invocation of such large models. Small-scale models (such as models with 7B, 13B, and 30B parameter scales) can already provide sufficiently accurate answers in many conventional scenarios. In the prior art, some studies have proposed methods such as model compression, knowledge distillation, and mixture of experts (MoE) to optimize the model computing efficiency and reduce resource consumption.
[0003] Relying on super-large models for computing not only causes waste of computing resources but also goes against the requirements of economy and sustainable development. Existing solutions have not effectively addressed how to dynamically schedule different-scale models according to task complexity to achieve optimal computing efficiency. In addition, current knowledge bases usually adopt static storage and retrieval methods and fail to be synergistically optimized with the model scheduling mechanism, resulting in the difficulty of effectively reducing the invocation frequency of super-large models, thereby further exacerbating the computing burden. Summary of the Invention
[0004] In view of this, the present invention provides a multi-level large model scheduling and knowledge base adaptive optimization method and system to solve the problem of being unable to dynamically schedule different-scale models according to task complexity.
[0005] In a first aspect, the present invention provides a multi-level large model scheduling and knowledge base adaptive optimization method, and the method includes: Obtain a task to be processed, preprocess the task to be processed, and calculate the complexity of the task to be processed; Use a small model to process the task to be processed, obtain a test processing result, and calculate the maturity of the test processing result; Determine a scheduling strategy for the large model according to the complexity and maturity, and use the scheduling strategy to process the task to be processed to obtain a processing result; Optimize the small model or knowledge base related to the scheduling strategy based on the processing result.
[0006] The multi - level large - model scheduling and knowledge - base adaptive optimization method provided by the present invention calculates the complexity of the task to be processed and tests the maturity of the processing results, selects models of different scales accordingly, reduces the invocation of super - large models, lowers the computing cost, saves computing resources, improves the system response efficiency, optimizes small models or knowledge - bases using historical processing results, enhances the processing capabilities of small models and knowledge - bases, continuously accumulates knowledge, and reduces repeated calculations.
[0007] In an alternative embodiment, after pre - processing the task to be processed, the complexity of the task to be processed is calculated, including: Perform semantic analysis on the task to be processed, and conduct multi - category classification according to the semantic analysis results to determine the type of the task to be processed; Obtain the first dynamic factor for calculating complexity, and the first dynamic factor includes: text length, proportion of proper nouns, knowledge - base matching degree, number of reasoning steps, context span, data complexity, uncertainty; Analyze the task to be processed based on the type of the task to be processed and the first dynamic factor to determine the values of the first dynamic factors of the task to be processed; Calculate the complexity of the task to be processed based on the values of the first dynamic factors of the task to be processed and the corresponding preset weights.
[0008] The multi - level large - model scheduling and knowledge - base adaptive optimization method provided by the present invention selects the first dynamic factors related to large - model scheduling according to experience and historical data, and assigns reasonable weights to each first dynamic factor based on historical data, which more scientifically reflects the actual contributions of various factors in complexity calculation, so that the complexity calculation result is more in line with the real situation and provides more reliable data support for subsequent decision - making.
[0009] In an alternative embodiment, use a small model to process the task to be processed to obtain test processing results, and calculate the maturity of the test processing results, including: Obtain the second dynamic factor for calculating maturity, and the second dynamic factor includes: quality, relevance to the problem, stability, user feedback score, context coherence, information integrity, degree of uncertainty; Analyze the test processing results based on the second dynamic factor to obtain the values of the second dynamic factors of the test processing results; Determine the maturity of the test processing results based on the values of the second dynamic factors of the test processing results and the corresponding preset weights.
[0010] The multi-level large model scheduling and knowledge base adaptive optimization method provided by the present invention uses a small model for trial processing, which can quickly complete calculations and reduce operating costs. By selecting a second dynamic factor related to maturity and calculating the maturity of the test processing results based on the second dynamic factor, on the premise of ensuring a certain processing quality, the cost-effective advantage of the small model is fully utilized, and computing resources are saved.
[0011] In an alternative embodiment, determining the scheduling strategy of the large model according to complexity and maturity includes: If the complexity is less than the preset complexity threshold and the maturity is not less than the preset maturity threshold, then select a small model to process the task to be processed; If the complexity is not less than the preset complexity threshold and the maturity is not less than the preset maturity threshold, then select a small model + knowledge base to process the task to be processed; If the complexity is not less than the preset complexity threshold and the maturity is less than the preset maturity threshold, then select a large model to process the task to be processed.
[0012] In an alternative embodiment, determining the scheduling strategy of the large model according to complexity and maturity further includes: If the complexity is less than the preset complexity threshold and the maturity is less than the preset maturity threshold, then obtain the type of the task to be processed and determine whether there is a corresponding knowledge base according to the type of the task to be processed; If there is a corresponding knowledge base, then select a small model + knowledge base to process the task to be processed; If there is no corresponding knowledge base, then select a large model to process the task to be processed.
[0013] The multi-level large model scheduling and knowledge base adaptive optimization method provided by the present invention accurately selects the most suitable large model according to the complexity of the task to be processed and the maturity of the test processing results, saves computing resources and improves response efficiency on the premise of ensuring processing quality.
[0014] In an alternative embodiment, optimizing the small model or knowledge base related to the scheduling strategy based on the processing results includes: From the processing results, screen out the large model processing results obtained by using the large model for processing; Obtain the user feedback and problem-solving rate of the large model processing results, and determine high-quality processing results according to the user feedback and problem-solving rate; Cache the high-quality processing results or store them in the corresponding knowledge base to realize the optimization of the knowledge base.
[0015] The multi-level large model scheduling and knowledge base adaptive optimization method provided by the present invention optimizes small models or knowledge bases by analyzing processing results, improves the scheduling strategy, enables the scheduling strategy to better adapt to the requirements of different types of tasks, allocates resources more precisely, avoids waste or shortage of resources, enables the system to process more tasks with limited resources, and improves the overall performance and stability of the system.
[0016] In an alternative embodiment, the adaptive optimization of small models or knowledge bases related to the scheduling strategy based on the processing results further includes: If the data in the knowledge base is greater than the preset capacity threshold, a fine-tuning data set is constructed using the data in the knowledge base; The corresponding small model is optimized using the fine-tuning data set.
[0017] The multi-level large model scheduling and knowledge base adaptive optimization method provided by the present invention optimizes small models by combining short-term memory, long-term memory, and knowledge solidification, improves the reasoning ability of small models, enables them to approach the performance of large models on specific tasks, and saves computing resources on the premise of ensuring processing quality.
[0018] In a second aspect, the present invention provides a multi-level large model scheduling and knowledge base adaptive optimization system, which includes: A complexity calculation module, configured to obtain a task to be processed, preprocess the task to be processed, and calculate the complexity of the task to be processed; A maturity calculation module, configured to process the task to be processed using a small model, obtain a test processing result, and calculate the maturity of the test processing result; A large model scheduling module, configured to determine a scheduling strategy for the large model according to the complexity and maturity, and process the task to be processed using the scheduling strategy to obtain a processing result; A knowledge base optimization module, configured to optimize a small model or a knowledge base related to the scheduling strategy based on the processing result.
[0019] In a third aspect, the present invention provides a computer device, including: a memory and a processor, which are communicatively connected to each other, the memory stores computer instructions, and the processor executes the computer instructions to execute the method according to the first aspect or any corresponding embodiment thereof.
[0020] In a fourth aspect, the present invention provides a computer-readable storage medium, on which computer instructions are stored, and the computer instructions are used to cause a computer to execute the method according to the first aspect or any corresponding embodiment thereof. Description of the Drawings
[0021] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the accompanying drawings required for the description of the specific embodiments or the prior art. Obviously, the accompanying drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.
[0022] Figure 1 is a schematic flowchart of a multi-level large model scheduling and knowledge base adaptive optimization method according to an embodiment of the present invention; Figure 2 is a schematic diagram of the system architecture in the multi-level large model scheduling and knowledge base adaptive optimization method according to an embodiment of the present invention; Figure 3 is a schematic flowchart of another multi-level large model scheduling and knowledge base adaptive optimization method according to an embodiment of the present invention; Figure 4 is a schematic flowchart of the scheduling process of the multi-level large model scheduling strategy in the multi-level large model scheduling and knowledge base adaptive optimization method according to an embodiment of the present invention; Figure 5 is a schematic flowchart of the process of storing high-quality answers in the multi-level large model scheduling and knowledge base adaptive optimization method according to an embodiment of the present invention; Figure 6 is a structural block diagram of a multi-level large model scheduling and knowledge base adaptive optimization system according to an embodiment of the present invention; Figure 7 is a schematic diagram of the hardware structure of a computer device according to an embodiment of the present invention. Specific Embodiments
[0023] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present invention.
[0024] The embodiments of the present invention provide a multi-level large model scheduling and knowledge base adaptive optimization method. By calculating the complexity of the task to be processed and testing the maturity of the processing results, different-scale models are selected accordingly. At the same time, the knowledge base or small model is optimized based on the processing results to achieve the effects of improving the processing capabilities of the small model and the knowledge base and enhancing the system response efficiency.
[0025] According to an embodiment of the present invention, an embodiment of a multi-level large model scheduling and knowledge base adaptive optimization method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0026] In this embodiment, a multi-level large model scheduling and knowledge base adaptive optimization method is provided, which can be used in the above computer system. Figure 1 It is a flowchart of the multi-level large model scheduling and knowledge base adaptive optimization method according to an embodiment of the present invention, as Figure 1 shown, this process includes the following steps: Step S101, obtain the task to be processed, and calculate the complexity of the task to be processed after preprocessing the task to be processed.
[0027] Specifically, the task to be processed can be the problem text input by the user. The problem text input by the user is cleaned, including removing stop words, word segmentation, grammar parsing, etc., to improve the understanding accuracy.
[0028] Semantic analysis is performed on the cleaned text, and the problems are classified into multiple categories according to the business. Calculate the complexity of the task to be processed based on the classified problem text. For example, the longer the text length, the higher the complexity. Therefore, the text length can be used as an influencing factor for calculating the complexity. In addition, multiple influencing factors can be determined according to historical experience or experimental data, and the values of each influencing factor are determined based on the problem to be processed. Finally, calculate the complexity of the problem to be processed based on the values of each influencing factor. This is only an example, but not limited thereto.
[0029] Step S102, use a small model to process the task to be processed, obtain the test processing result, and calculate the maturity of the test processing result.
[0030] Specifically, according to the type of the task to be processed, select a suitable small model to process the task to be processed and try to answer the user's question. There are a wide variety of small models, covering dedicated models in multiple fields such as natural language processing, computer vision, and data analysis. Before processing the task to be processed, it is necessary to select a suitable small model according to the task type, data characteristics, and performance requirements. For example, for text classification tasks, a lightweight model based on the Transformer architecture, such as DistilBERT, can be selected; for image recognition tasks, MobileNet series models can be selected, which have high computing power and a small model size.
[0031] After receiving the preprocessed task data, the adapted small model performs calculations according to its internal algorithm logic and trained parameters. In this process, the small model extracts, analyzes, and maps data features, and outputs processing results. For example, in the sentiment analysis task of natural language processing, the small model will perform semantic understanding and sentiment tendency judgment on the input text, and output positive, negative, or neutral sentiment classification results, which is only an example and not limited thereto.
[0032] Determine multiple influencing factors that affect the maturity of the processing result based on historical experience or experimental data, determine the values of each influencing factor based on the test processing result, and finally calculate the maturity of the test processing result according to the values of each influencing factor.
[0033] Step S103, determine the scheduling strategy of the large model according to the complexity and maturity, and use the scheduling strategy to process the task to be processed to obtain the processing result.
[0034] Specifically, the complexity of the problem to be processed determines the difficulty and required resources for the large model to process the task. If a simple scheduling strategy is adopted for a high-complexity task, it may lead to a decrease in the maturity of the processing result. For example, in the multi-modal sentiment analysis task, due to the high complexity of the fusion of text, image, and audio data, if the large model only allocates computing resources conventionally, it may not be able to fully extract features due to insufficient resources, resulting in a decline in maturity indicators such as the accuracy rate and recall rate of sentiment classification. The maturity evaluation result can inversely guide the accuracy of complexity analysis. If the maturity of the large model's processing result does not meet the expectation, and it is found through analysis that the complexity estimation is insufficient, for example, the originally thought simple text summarization task is actually more complex due to the involvement of professional domain knowledge, then it is necessary to re-examine the complexity calculation and adjust the subsequent scheduling strategy.
[0035] Based on the above analysis, the complexity of the problem to be processed can be divided into multiple levels, and a maturity threshold for the test processing result can be set. When performing multi-level large model scheduling, according to the relationship between the complexity level of the problem to be processed, the maturity of the test processing result and the maturity threshold, dynamically schedule large models of different scales to achieve the optimal task processing efficiency. For low-complexity tasks, such as basic text keyword extraction, a relatively small-scale large model or multiple small models can be allocated to process in parallel, making full use of edge computing resources to reduce the pressure on the core computing nodes. For low-complexity tasks, if the maturity of the test processing result is not lower than the maturity threshold, it means that when using a small model to process this task, the processing effect can meet the expectation, and the low-complexity task can be directly processed using a small model, which is only an example and not limited thereto.
[0036] The scheduling strategies for large models can include: small models, small models + knowledge bases, large models, and extra-large models. The main differences between small models, large models, and extra-large models lie in the scale of the number of parameters and the computing resources required. This is only an example and not limited thereto.
[0037] Step S104, optimize the small model or knowledge base related to the scheduling strategy based on the processing result.
[0038] Specifically, the processing results include high-quality processing results that meet user needs and general processing results that do not meet user needs. For high-quality processing results, it can be considered that the corresponding large model has good performance. The high-quality processing results and the corresponding tasks to be processed can be used as new knowledge to supplement the knowledge base or as optimization samples to optimize the parameters of the small model. For example, through hyperparameter tuning techniques, such as grid search, random search, or Bayesian optimization, etc., to find the optimal combination of hyperparameters, improve the training effect and prediction accuracy of the model; as the knowledge base continues to expand, there may be situations of knowledge redundancy or duplication, and it is necessary to integrate and optimize the knowledge base, merge duplicate knowledge, streamline the expression, and improve the readability and maintainability of the knowledge. At the same time, according to the feedback of the processing results, adjust the priority of the knowledge, place commonly used and important knowledge in a more prioritized position, facilitate the small model to quickly obtain it, and improve the efficiency of processing tasks.
[0039] As Figure 2 shown, it is a schematic diagram of the system architecture of this embodiment. After the user inputs a question, it is first classified, and then the large model is scheduled to call the knowledge base or model. If the large model is called to answer the question, knowledge memory is generated to optimize the knowledge base or small model, and the optimized knowledge base or small model is used to feedback to the large model scheduling. When a new question input by the user is received again, the answering results of the small model and / or knowledge base will be improved, thus realizing the dynamic scheduling strategy of the multi-level large model.
[0040] The multi-level large model scheduling and knowledge base adaptive optimization method provided in this embodiment calculates the complexity of the task to be processed, tests the maturity of the processing results, selects models of different scales accordingly, reduces the call to the extra-large model, reduces the computing cost, saves computing resources, improves the system response efficiency, optimizes the small model or knowledge base using historical processing results, improves the processing capabilities of the small model and knowledge base, continuously accumulates knowledge, and reduces repeated calculations.
[0041] In this embodiment, a multi-level large model scheduling and knowledge base adaptive optimization method is provided, which can be used in the above computer system. Figure 3 It is a flowchart of the multi-level large model scheduling and knowledge base adaptive optimization method according to the embodiment of the present invention. As Figure 3 shown, this process includes the following steps: Step S201: Obtain the task to be processed, preprocess the task to be processed, and calculate the complexity of the task to be processed.
[0042] Specifically, the above step S201 includes: Step S2011: Perform semantic analysis on the task to be processed, and perform multi-category classification according to the semantic analysis results to determine the type of the task to be processed.
[0043] Specifically, a multi-modal semantic analysis module (including BERT sentence vector encoding, domain keyword matching, and syntactic dependency parsing) can be used to perform semantic analysis and in-depth understanding of the user's question, and based on the three-level business classification system: ① domain attribution (such as medical, financial, legal), ② task type (such as diagnosis, contract review, knowledge Q&A), ③ computational complexity (quantified based on semantic entropy and inference step length), generate structured classification labels, and then drive the dynamic scheduling decision of downstream heterogeneous models. For example, classify 'Analysis of the efficacy of liver cancer targeted drugs' as [Medical - Treatment Plan Evaluation - High Complexity], so as to intelligently match it to the 30B tumor combined treatment special model, rather than the 100-billion-level general model. This is only an example, but not limited to this.
[0044] Step S2012: Obtain the first dynamic factor of the computational complexity. The first dynamic factor includes: text length, proportion of proper nouns, knowledge base matching degree, number of inference steps, context span, data complexity, uncertainty.
[0045] Specifically, the calculation formula for the computational complexity is: (1) Among them, C represents the problem complexity score, L represents the text length of the problem, T represents the proportion of proper nouns (calculated by the proportion of words in the existing knowledge base), K represents the knowledge base matching degree (the more existing knowledge, the lower the complexity), R represents the multi-step inference requirement (if the problem requires multiple logical steps to solve, the complexity increases), S represents the context span involved in the problem (such as the intersection of multiple knowledge domains, the complexity increases), D represents the data dimension complexity (such as involving multi-dimensional tables, graph data, etc., the complexity increases), P represents the uncertainty of the problem (such as the problem has a high semantic ambiguity, the complexity increases).
[0046] w1 - w7 represent the first weighting factors of each first dynamic factor (preset weights, adjustable). The determination process of the first weighting factor can be based on historical data and obtained through least squares fitting. This process is a mature existing technology and will not be elaborated here. The magnitude of the first weighting factor represents the influence degree of the first dynamic factor on the complexity. The larger the first weighting factor, the greater the influence of the corresponding first dynamic factor on the complexity, and the smaller the first weighting factor, the smaller the influence of the corresponding first dynamic factor on the complexity.
[0047] Step S2013: Analyze the task to be processed based on the type of the task to be processed and the first dynamic factors, and determine the values of the first dynamic factors of the task to be processed.
[0048] Specifically, determine the values of the first dynamic factors according to the type and other characteristics of the task to be processed, which specifically include: obtaining the text length of the task to be processed, and performing normalization processing on the text length to obtain the value of the text length, that is, L = the number of text characters of the task to be processed / the maximum allowable number of characters; calculating the proportion of the number of proper nouns in the task to be processed in the total number of words in the corresponding knowledge base as the proper noun ratio, that is, T = the number of proper nouns / the total number of times in the corresponding knowledge base; using the relevance score and vector similarity matching result of the information retrieved and recalled from the knowledge base to determine the knowledge base matching degree K (specifically including: a knowledge retrieval: recall relevant entries of relevant questions from the knowledge base, and obtain a candidate knowledge fragment set through vector retrieval (FAISS / Annoy) or keyword inverted index; b semantic matching: calculate the similarity between the question and the candidate knowledge through the vector cosine similarity of BERT / SimCSE to obtain a list of relevance scores; c score normalization: convert the matching score into a standardized K value through Softmax / Min-Max normalization to obtain the top-k best results); if the question contains "causal relationship", "multi-level logic", etc., determine the number of reasoning steps as the reasoning requirement R; determine the number of knowledge domains according to the type of the task to be processed, and determine the context span S based on the number of knowledge domains; determine the data complexity D by judging whether it involves large-scale data sets, multiple data structures, etc.; calculate the semantic uncertainty P using Neuro-Linguistic Programming (such as analyzing the ambiguity degree through a syntactic tree).
[0049] Step S2014: Calculate the complexity of the task to be processed based on the values of the first dynamic factors of the task to be processed and the corresponding preset weights.
[0050] Specifically, substitute the values of the first dynamic factors of the task to be processed and the corresponding preset weights into formula (1) for weighted calculation to obtain the complexity of the task to be processed.
[0051] The multi-level large model scheduling and knowledge base adaptive optimization method provided in this embodiment selects the first dynamic factors related to large model scheduling according to experience and historical data, and assigns reasonable weights to each first dynamic factor based on historical data, which more scientifically reflects the actual contributions of various factors in complexity calculation, so that the complexity calculation result is more in line with the real situation and provides more reliable data support for subsequent decisions.
[0052] Step S202: Use a small model to process the task to be processed, obtain the test processing result, and calculate the maturity of the test processing result.
[0053] Specifically, the above-mentioned step S202 includes: Step S2021: Obtain the second dynamic factor for calculating maturity. The second dynamic factor includes: quality, relevance to the problem, stability, user feedback score, context coherence, information completeness, degree of uncertainty.
[0054] Specifically, use a suitable small model to attempt to process the task to be processed, obtain the test processing result. The calculation formula for the maturity of the test processing result is: (2) where M represents the answer maturity score, Q represents the answer quality, E represents the relevance of the answer to the question, V represents the stability of the answer, U represents the user feedback, B represents the context coherence (whether the answer is consistent with the existing knowledge base content), I represents the information completeness (whether the answer covers all key points of the question), and A represents the ambiguous expression (the degree of ambiguity or uncertainty in the question or answer).
[0055] v1 - v7 represent the second weighting factors corresponding to the second dynamic factors (preset weights, adjustable). The determination process of the second weighting factors can be based on historical data and obtained through least - squares fitting. This process is a mature existing technology and will not be elaborated here. The magnitude of the second weighting factor represents the degree of influence of the second dynamic factor on complexity. The larger the second weighting factor, the greater the influence of the corresponding second dynamic factor on complexity, and the smaller the second weighting factor, the smaller the influence of the corresponding second dynamic factor on complexity.
[0056] Step S2022: Analyze the test processing result based on the second dynamic factor to obtain the values of each second dynamic factor of the test processing result.
[0057] Specifically, natural language processing (NLP) technology can be used to automatically score the text content. Through NLP scoring, such as calculating the answer quality based on BLEU, ROUGE, and BERTScore , where, represents the overlap rate evaluation similarity between the generated text and the reference text calculated based on word - level matching, represents the degree of phrase overlap between the generated text and the reference text measured based on recall rate, represents the similarity between the generated text and the reference text evaluated from the semantic level based on the pre - trained language model, 、 、 respectively represent , , The weights of; Vector similarity calculation can be used to determine the relevance between the answer and the question , where represents the semantic vector of the question, represents the semantic vector of the answer; By the multiple sampling method, the same question is input into the model N times (such as 5 times), and the average value of the semantic similarity of multiple answers is calculated as the stability V of the answer; Based on the user's feedback score, such as whether the user adopts the answer as the user feedback U; Compare the consistency score of the knowledge base answer , as the context coherence B, where represents the answer generated by the system, the standard answer in the knowledge base; Determine the information integrity I through keyword coverage or expert annotation; Determine the ambiguous expressions A by counting the number of ambiguous, unclear, and ambiguous words in the answer
[0058] Step S2023, based on the value of the second dynamic factor of the test processing result and the corresponding preset weight, determine the maturity of the test processing result
[0059] Specifically, substitute the value of the second dynamic factor of the test processing result and the corresponding preset weight into formula (2) for weighted calculation to obtain the maturity of the test processing result
[0060] The multi-level large model scheduling and knowledge base adaptive optimization method provided in this embodiment uses a small model for trial processing, which can quickly complete calculations and reduce operating costs. By selecting the second dynamic factor related to maturity and calculating the maturity of the test processing result based on the second dynamic factor, under the premise of ensuring a certain processing quality, the cost-effective advantage of the small model is fully utilized, and computing resources are saved
[0061] Step S203, determine the scheduling strategy of the large model according to the complexity and maturity, and use the scheduling strategy to process the task to be processed to obtain the processing result
[0062] Specifically, determining the scheduling strategy of the large model according to the complexity and maturity in the above step S203 includes Step S2031, if the complexity is less than the preset complexity threshold and the maturity is not less than the preset maturity threshold, select a small model to process the task to be processed
[0063] Step S2032, if the complexity is not less than the preset complexity threshold and the maturity is not less than the preset maturity threshold, select a small model + knowledge base to process the task to be processed
[0064] Step S2033, if the complexity is not less than the preset complexity threshold and the maturity is less than the preset maturity threshold, then select a large model to process the task to be processed.
[0065] Specifically, select the scheduling strategy of the large model according to the problem complexity C and the answer maturity M. As Figure 4 shown, it is a schematic diagram of the scheduling process of the multi-level large model scheduling strategy, which specifically includes: Set the preset complexity threshold and the preset maturity threshold according to experience. If the complexity is less than the preset complexity threshold and the maturity is not less than the preset maturity threshold, it means that the complexity is low and the maturity is high, then directly use the small model to answer the question; if the complexity is not less than the preset complexity threshold and the maturity is not less than the preset maturity threshold, it means that the complexity is high and the maturity is high, then select the small model + knowledge base to answer the question; if the complexity is not less than the preset complexity threshold and the maturity is less than the preset maturity threshold, it means that the complexity is high and the maturity is low, then select a large model (such as 671B) to answer the question.
[0066] The scheduling strategy of the large model is expressed by the formula: D = f(C, M), where D represents the model scheduling strategy, and f(C,M) can be a decision tree or a scheduling model based on deep learning, which is a mature existing technology and will not be elaborated here.
[0067] In some alternative embodiments, determining the scheduling strategy of the large model according to the complexity and maturity further includes: Step S2034, if the complexity is less than the preset complexity threshold and the maturity is less than the preset maturity threshold, then obtain the type of the task to be processed, and determine whether there is a corresponding knowledge base according to the type of the task to be processed.
[0068] Step S2035, if there is a corresponding knowledge base, then select the small model + knowledge base to process the task to be processed.
[0069] Step S2036, if there is no corresponding knowledge base, then select a large model to process the task to be processed.
[0070] Specifically, as Figure 4 shown, if the complexity is less than the preset complexity threshold and the maturity is less than the preset maturity threshold, it means that the complexity is low and the maturity is low, then preferentially try to use the small model + knowledge base to answer the question. If there is no corresponding knowledge base, then call a larger model to answer the question.
[0071] The multi-level large model scheduling and knowledge base adaptive optimization method provided by this embodiment of the present invention accurately selects the most suitable large model according to the complexity of the task to be processed and the maturity of the test processing results, saves computing resources and improves response efficiency on the premise of ensuring the processing quality.
[0072] Step S204: Optimize the small model or knowledge base related to the scheduling strategy based on the processing result.
[0073] Specifically, the above-mentioned step S204 includes: Step S2041: Screen out the large model processing results obtained by processing with the large model from the processing results.
[0074] Step S2042: Obtain the user feedback and problem-solving rate of the large model processing results, and determine the high-quality processing results according to the user feedback and problem-solving rate.
[0075] Specifically, the processing results obtained by processing tasks with the large model are very accurate. Therefore, screen out the large model processing results obtained by processing with the large model from the processing results, and determine the high-quality processing results through user feedback and problem-solving rate. Generally speaking, judge whether the user adopts, the user satisfaction, and whether the problem is solved according to the feedback U given by the user's answer to the large model processing results. The large model processing results adopted by the user and with the problem solved can be used as high-quality processing results.
[0076] Step S2043: Cache the high-quality processing results or store them in the corresponding knowledge base to optimize the knowledge base.
[0077] Specifically, store the high-quality processing results in the knowledge base corresponding to the type of the task to be processed. As Figure 5 shown, it is a schematic flowchart of storing high-quality answers, and the storage methods include short-term memory, long-term memory, and knowledge solidification.
[0078] The short-term memory is used to cache the answers to recently frequent questions to improve the call efficiency. The answers to recently frequent questions (such as the frequency is TOP100) can be stored using the Remote Dictionary Server (Redis). The key is the question fingerprint (MD5 hash), and the value is the answer + metadata (such as the call times, the last update time).
[0079] The long-term memory uses the clustering algorithm to summarize and integrate the knowledge based on the knowledge accumulated from long-term use, form a knowledge topic tree, and store it in the knowledge base in the form of data vectors to solve problems such as repeated and scattered knowledge, and achieve more organized knowledge classification.
[0080] The multi-level large model scheduling and knowledge base adaptive optimization method provided in this embodiment optimizes the small model or knowledge base by analyzing the processing results, improves the scheduling strategy, enables the scheduling strategy to better adapt to the needs of different types of tasks, allocate resources more accurately, avoid waste or shortage of resources, enables the system to process more tasks with limited resources, and improves the overall performance and stability of the system.
[0081] In some alternative embodiments, for the adaptive optimization of small models or knowledge bases related to scheduling policies based on processing results, it further includes: If the data in the knowledge base is greater than the preset capacity threshold, use the data in the knowledge base to construct a fine-tuning data set.
[0082] Use the fine-tuning data set to optimize the corresponding small model.
[0083] Specifically, the capacity of the knowledge base has an upper limit. When the capacity occupied by the data stored in the knowledge base through long-term memory reaches the preset capacity threshold, the knowledge base can no longer store new knowledge in the form of long-term memory. Then, it is necessary to extract high-frequency and high-confidence question-and-answer pairs (questions and corresponding answers) from the knowledge base, construct an instruction fine-tuning data set, and use the fine-tuning data set to fine-tune the corresponding type of small-scale model (such as LLaMA-7B) to make it have stronger specific-domain reasoning capabilities and form a small model expert in the vertical domain. During the fine-tuning process, reasonable, compliant, and correct judgments are made on the knowledge content to ensure the credibility of the knowledge.
[0084] In a specific embodiment, when upgrading the knowledge base for liver cancer targeted therapy, in the liver cancer drug use consultation, the large model (671B model) is frequently asked "how to handle drug resistance". The internal processing process for the answer to this question includes: Short-term memory cache: Search in the short-term high-frequency question-and-answer to see if there is this question. If there is, return the cached answer. If the corresponding cache record cannot be found in the short-term cache, it is necessary to generate an answer through the model.
[0085] Screen high-quality answers: Comprehensively calculate the maturity M of the answer based on factors such as the user feedback result U, and determine whether the maturity is greater than 0.8 through a preset threshold, such as 0.8, to decide whether to classify it as a high-quality answer and store it in the relevant knowledge base.
[0086] Long-term memory clustering: Cluster similar questions, such as "what to do when liver cancer targeted drugs fail", "treatment plan after sorafenib resistance", "applicable conditions of regorafenib", etc., and integrate and summarize them into a theme "subsequent treatment of liver cancer drug resistance" through clustering for knowledge integration of questions and answers.
[0087] Knowledge solidification and fine-tuning: Organize a certain type of question-and-answer pair into a type of fine-tuning data set, for example, {"instruction": "What are the treatment options for liver cancer patients after sorafenib resistance?", "output": "According to the RESORCE trial, it is recommended to switch to regorafenib..."}, and then fine-tune the same type of small model to adapt to the pattern of this type of question, so as to obtain a dedicated small model in a specific domain.
[0088] The multi-level large model scheduling and knowledge base adaptive optimization method provided in this embodiment optimizes small models by combining short-term memory, long-term memory, and knowledge solidification, improves the reasoning ability of small models, makes their performance on specific tasks close to that of large models, and saves computing resources on the premise of ensuring processing quality.
[0089] In this embodiment, a multi-level large model scheduling and knowledge base adaptive optimization system is also provided. This system is used to implement the above-mentioned embodiments and preferred implementation manners, and those that have been described will not be repeated. As used hereinafter, the term "module" can be a combination of software and / or hardware that can achieve a predetermined function. Although the systems described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.
[0090] This embodiment provides a multi-level large model scheduling and knowledge base adaptive optimization system, as Figure 6 shown, including: A complexity calculation module 601, configured to obtain a task to be processed, preprocess the task to be processed, and calculate the complexity of the task to be processed.
[0091] A maturity calculation module 602, configured to use a small model to process the task to be processed, obtain a test processing result, and calculate the maturity of the test processing result.
[0092] A large model scheduling module 603, configured to determine a scheduling strategy for the large model according to the complexity and maturity, and use the scheduling strategy to process the task to be processed to obtain a processing result.
[0093] A knowledge base optimization module 604, configured to optimize the small model or knowledge base related to the scheduling strategy based on the processing result.
[0094] In some alternative implementation manners, the complexity calculation module 601 includes: A task classification unit, configured to perform semantic analysis on the task to be processed, perform multi-category classification according to the semantic analysis result, and determine the type of the task to be processed.
[0095] A first dynamic factor determination unit, configured to obtain a first dynamic factor for calculating complexity. The first dynamic factor includes: text length, proportion of proper nouns, knowledge base matching degree, number of reasoning steps, context span, data complexity, uncertainty.
[0096] A first dynamic factor value determination unit, configured to analyze the task to be processed based on the type of the task to be processed and the first dynamic factor, and determine the values of the first dynamic factors of the task to be processed.
[0097] A complexity calculation unit for calculating the complexity of a task to be processed based on the values of the first dynamic factors of the task to be processed and the corresponding preset weights.
[0098] In some alternative embodiments, the maturity calculation module 602 includes: A second dynamic factor determination unit for obtaining the second dynamic factors for calculating maturity, where the second dynamic factors include: quality, relevance to the problem, stability, user feedback score, context coherence, information integrity, degree of uncertainty.
[0099] A second dynamic factor value calculation unit for analyzing the test processing results based on the second dynamic factors to obtain the values of the second dynamic factors of the test processing results.
[0100] A maturity calculation unit for determining the maturity of the test processing results based on the values of the second dynamic factors of the test processing results and the corresponding preset weights.
[0101] In some alternative embodiments, the large model scheduling module 603 includes: A first scheduling unit for selecting a small model to process the task to be processed if the complexity is less than a preset complexity threshold and the maturity is not less than a preset maturity threshold.
[0102] A second scheduling unit for selecting a small model + knowledge base to process the task to be processed if the complexity is not less than a preset complexity threshold and the maturity is not less than a preset maturity threshold.
[0103] A third scheduling unit for selecting a large model to process the task to be processed if the complexity is not less than a preset complexity threshold and the maturity is less than a preset maturity threshold.
[0104] A fourth scheduling unit for obtaining the type of the task to be processed and determining whether there is a corresponding knowledge base according to the type of the task to be processed if the complexity is less than a preset complexity threshold and the maturity is less than a preset maturity threshold.
[0105] A first scheduling subunit for selecting a small model + knowledge base to process the task to be processed if there is a corresponding knowledge base.
[0106] A second scheduling subunit for selecting a large model to process the task to be processed if there is no corresponding knowledge base.
[0107] In some alternative embodiments, the knowledge base optimization module 604 includes: A processing result screening unit for screening out the large model processing results obtained by processing using a large model from the processing results.
[0108] A high-quality processing result determination unit is configured to obtain user feedback and problem-solving rates of large model processing results, and determine high-quality processing results based on the user feedback and problem-solving rates.
[0109] A knowledge base optimization unit is configured to cache the high-quality processing results or store them in the corresponding knowledge base to optimize the knowledge base.
[0110] The further function descriptions of the above-mentioned various modules and units are the same as those in the corresponding embodiments above, and will not be elaborated here.
[0111] The multi-level large model scheduling and knowledge base adaptive optimization system in this embodiment is presented in the form of functional units. Here, the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and a memory that execute one or more software or fixed programs, and / or other devices that can provide the above functions.
[0112] The embodiment of the present invention also provides a computer device having the above Figure 6 multi-level large model scheduling and knowledge base adaptive optimization system as shown.
[0113] Please refer to Figure 7 , Figure 7 which is a schematic structural diagram of a computer device provided by an optional embodiment of the present invention. As shown in Figure 7 , the computer device includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including a high-speed interface and a low-speed interface. Each component communicates with each other using different buses and can be installed on a common motherboard or installed in other ways as needed. The processor can process instructions executed within the computer device, including instructions stored in the memory or on the memory to display graphical information of the GUI on an external input / output device (such as a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (such as a server array, a set of blade servers, or a multi-processor system). Figure 7 One processor 10 is taken as an example in
[0114] The processor 10 can be a central processing unit, a network processor, or a combination thereof. Among them, the processor 10 can further include a hardware chip. The above hardware chip can be an application specific integrated circuit, a programmable logic device, or a combination thereof. The above programmable logic device can be a complex programmable logic device, a field programmable gate array, a generic array logic, or any combination thereof.
[0115] Among them, the memory 20 stores instructions that can be executed by at least one processor 10, so that the at least one processor 10 executes the method shown in the above embodiments.
[0116] The memory 20 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the computer device, etc. In addition, the memory 20 may include a high-speed random access memory, and may also include a non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some alternative embodiments, the memory 20 may optionally include a memory remotely disposed relative to the processor 10, and these remote memories may be connected to the computer device through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0117] The memory 20 may include a volatile memory, for example, a random access memory; the memory may also include a non-volatile memory, for example, a flash memory, a hard disk, or a solid-state drive; the memory 20 may also include a combination of the above types of memories.
[0118] The computer device further includes a communication interface 30 for the computer device to communicate with other devices or a communication network.
[0119] The embodiments of the present invention also provide a computer-readable storage medium. The method according to the embodiments of the present invention can be implemented in hardware, firmware, or be implemented as computer code that can be recorded on a storage medium, or be implemented as computer code originally stored in a remote storage medium or a non-transitory machine-readable storage medium and to be downloaded through a network and stored in a local storage medium, so that the method described herein can be processed by such software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disc, a read-only memory, a random access memory, a flash memory, a hard disk, or a solid-state drive, etc.; further, the storage medium may also include a combination of the above types of memories. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code, and when the software or computer code is accessed and executed by the computer, the processor, or the hardware, the method shown in the above embodiments is implemented.
[0120] Although the embodiments of the present invention are described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations all fall within the scope defined by the appended claims.
Claims
1. A multi-level large model scheduling and knowledge base adaptive optimization method, characterized in that The method includes: Obtain the task to be processed, preprocess the task to be processed, and calculate the complexity of the task to be processed; Use a small model to process the task to be processed, obtain a test processing result, and calculate the maturity of the test processing result; Determine the scheduling strategy of the large model according to the complexity and the maturity, and use the scheduling strategy to process the task to be processed to obtain a processing result; Optimize the small model or knowledge base related to the scheduling strategy based on the processing result.
2. The method according to claim 1, characterized in that, Preprocess the task to be processed and calculate the complexity of the task to be processed, including: Perform semantic analysis on the task to be processed, perform multi-category classification according to the semantic analysis result, and determine the type of the task to be processed; Obtain a first dynamic factor for calculating complexity, where the first dynamic factor includes: text length, proportion of proper nouns, knowledge base matching degree, number of reasoning steps, context span, data complexity, uncertainty; Analyze the task to be processed based on the type of the task to be processed and the first dynamic factor, and determine the values of the first dynamic factors of the task to be processed; Calculate the complexity of the task to be processed based on the values of the first dynamic factors of the task to be processed and the corresponding preset weights.
3. The method according to claim 1, characterized in that, Use a small model to process the task to be processed, obtain a test processing result, and calculate the maturity of the test processing result, including: Obtain a second dynamic factor for calculating maturity, where the second dynamic factor includes: quality, relevance to the problem, stability, user feedback score, context coherence, information integrity, uncertainty degree; Analyze the test processing result based on the second dynamic factor to obtain the values of the second dynamic factors of the test processing result; Determine the maturity of the test processing result based on the values of the second dynamic factors of the test processing result and the corresponding preset weights.
4. The method according to claim 1, wherein Determine the scheduling strategy of the large model according to the complexity and the maturity, including: If the complexity is less than a preset complexity threshold and the maturity is not less than a preset maturity threshold, select a small model to process the task to be processed; If the complexity is not less than a preset complexity threshold and the maturity is not less than a preset maturity threshold, select a small model + knowledge base to process the task to be processed; If the complexity is not less than a preset complexity threshold and the maturity is less than a preset maturity threshold, select a large model to process the task to be processed.
5. The method according to claim 4, characterized in that, Determine the scheduling strategy of the large model according to the complexity and the maturity, and further include: If the complexity is less than a preset complexity threshold and the maturity is less than a preset maturity threshold, obtain the type of the task to be processed, and determine whether there is a corresponding knowledge base according to the type of the task to be processed; If there is a corresponding knowledge base, select a small model + knowledge base to process the task to be processed; If there is no corresponding knowledge base, select a large model to process the task to be processed.
6. The method according to claim 1, characterized in that, Optimize the small model or knowledge base related to the scheduling strategy based on the processing result, including: From the processing result, screen out the large model processing result obtained by using the large model for processing; Obtain the user feedback and problem-solving rate of the large model processing result, and determine the high-quality processing result according to the user feedback and problem-solving rate; Cache the high-quality processing result or store it in the corresponding knowledge base to optimize the knowledge base.
7. The method according to claim 6, characterized in that Based on the processing result, adaptively optimize the small model or knowledge base related to the scheduling strategy, and further include: If the data in the knowledge base is greater than the preset capacity threshold, use the data in the knowledge base to construct a fine-tuning data set; Use the fine-tuning data set to optimize the corresponding small model.
8. A multi-level large model scheduling and knowledge base adaptive optimization system, characterized in that, The system includes: A complexity calculation module for obtaining a task to be processed, preprocessing the task to be processed, and calculating the complexity of the task to be processed; A maturity calculation module for using a small model to process the task to be processed, obtaining a test processing result, and calculating the maturity of the test processing result; A large model scheduling module for determining the scheduling strategy of the large model according to the complexity and the maturity, and using the scheduling strategy to process the task to be processed to obtain a processing result; A knowledge base optimization module for optimizing the small model or knowledge base related to the scheduling strategy based on the processing result.
9. A computer device, characterized in that, Include: A memory and a processor, which are communicatively connected to each other. The memory stores computer instructions, and the processor executes the computer instructions to execute the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, Computer instructions are stored on the computer-readable storage medium, and the computer instructions are used to cause a computer to execute the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Task management method and related system
CN118656179A
Electric power scene complex task arrangement decision-making method and system based on large semantic model
CN118735222A
Technical avoidance design method based on large language model
CN119398984A
Multi-hardware mixed large model reasoning method, system and related device
CN119539089A
Real-time adaptive decision system and method using predictive modeling
GB201419630D0
Cited By
Instruction identification method, apparatus and device, and computer readable medium
CN120977303A
Large model agent reasoning scheduling method and system based on machine learning
CN122154955A
A neural network model loading implementation method and device and medium
CN122547416A