Model optimization management system combining model distillation and offline caching

By combining the optimization management system of model distillation and offline cache, and adaptive selection of processing modules, the problem of insufficient resource allocation in large language model systems under complex problems is solved, and fast response and efficient performance maintenance are achieved.

CN120276827APending Publication Date: 2025-07-08杭州点存科技有限公司
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510487603.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

When existing large language model systems deal with complex or novel problems, they cannot dynamically adjust resource allocation, resulting in long processing time and degradation of performance, and lack real-time monitoring and optimization mechanisms.

Method used

Combining the model distillation and offline cache model optimization management system, a complexity evaluation model is introduced, and through adaptive selection of processing modules, offline cache or model distillation modules are timely optimized to ensure system performance.

Benefits of technology

Improves system flexibility and adaptability, reduces resource waste, ensures fast and accurate response, and maintains optimal system performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120276827A_ABST
    Figure CN120276827A_ABST
Patent Text Reader

Abstract

The invention discloses a model optimization management system combining model distillation and offline caching, and relates to the technical field of large language models. The system is divided into a plurality of function modules which operate in sequence, and high efficiency and stability of the optimization process are ensured; the system comprises a data preprocessing module, a model distillation module, an offline cache module, a dynamic scheduling module and a monitoring feedback module. According to the technical key points, a complexity evaluation model is introduced, and after two times of judgment processing, a system can adaptively select a processing module according to the complexity of different problems at the moment; by regularly collecting the effect evaluation index and analyzing the change trend of the effect evaluation index, the problem of module performance reduction can be found in time, and when it is detected that the effect evaluation index is in the reduction trend, an optimization adjustment measure is triggered, and targeted optimization is performed on the offline cache module or the model distillation module, so that the system is ensured to always keep the optimal performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of large language models, and specifically to a model optimization management system that combines model distillation and offline caching. Background Technique

[0002] Large language models, abbreviated as LLM, are a type of artificial intelligence model based on deep learning, aiming to process and generate natural language text; by training on large-scale text data, LLM can understand and generate text similar to human language, and perform various natural language processing tasks including text generation, translation, and sentiment analysis; the importance of LLM lies in its wide range of application scenarios, showing its strong practicality and potential from automated customer service to advanced research; LLM usually adopts large-scale neural networks with the number of parameters ranging from millions to billions, enabling it to handle complex language tasks; the training process of LLM is divided into two stages: pre-training and fine-tuning. In the pre-training stage, self-supervised learning is carried out on large-scale unlabeled text data, and in the fine-tuning stage, supervised learning is carried out on labeled data for specific tasks; with the wide application of large language models, inference efficiency and resource consumption have become the main bottlenecks restricting their expansion in practical applications. Especially in complex tasks such as text-to-image prompt generation, optimizing memory usage and inference speed becomes particularly important.

[0003] In the existing document with the application number CN202410057493.6 and the title "An Output Optimization Method for a Large Language Model", it is pointed out that: the method includes: transmitting information between the LMM output generation module and the prompt optimization module through the historical dialogue processing module; generating the first LMM output by the LMM output generation module according to the initial prompt input by the user; determining that the first LMM output needs to be optimized according to the preset evaluation criteria by the prompt optimization module, and obtaining the improved required prompt; generating the second LMM output by the LMM output generation module by combining the first LMM output and the improved required prompt; the present invention realizes the automation of large language model output optimization by combining the LMM output generation module, the prompt optimization module, and the historical dialogue processing module, without the need to repeatedly debug the prompt manually, improving the output efficiency and quality of the LMM output generation model; however, it cannot more accurately and effectively select the current required operation modules, such as the model distillation module and the offline caching module, while optimizing the model.

[0004] Combined with the above document and the prior art: In traditional systems, a fixed resource allocation method is often adopted, which cannot be dynamically adjusted according to the actual situation of the problem. For complex or novel problems, traditional systems may require a long processing time. Even if a model distillation module or an offline cache module is adopted, there will still be problems. For example, it is impossible to accurately and effectively determine which module to use for running processing. Even if there is a determination standard, most of these standards are fixed conventional settings and cannot be adjusted or selected more flexibly. In traditional operating systems, there is often a lack of real-time monitoring and dynamic optimization mechanisms for changes in module performance, resulting in the inability to adjust in time when the system performance deteriorates. Summary of the Invention

[0005] (I) Technical problems to be solved In view of the deficiencies of the prior art, the present invention provides a model optimization management system that combines model distillation and offline caching. By introducing a complexity evaluation model, after two judgment processes, the system can adaptively select a processing module according to the complexity of different problems at this time. It can timely detect problems with deteriorating module performance. When it detects that the effect evaluation index shows a downward trend, it triggers an optimization adjustment measure to specifically optimize the offline cache module or the model distillation module to ensure that the system always maintains the best performance, and solves the problems raised in the background art.

[0006] (II) Technical solutions To achieve the above objectives, the present invention is realized through the following technical solutions: A model optimization management system that combines model distillation and offline caching, the system includes: A data preprocessing module, which collects and organizes text data in the target domain and constructs a dataset for the target domain; A model distillation module, which softens the output of the original model by knowledge distillation, trains a student model, performs inter-layer distillation between the original model and the student model, and determines whether to execute an adaptive distillation strategy according to the performance status of the student model; An offline cache module, which sets a cache data structure and formulates a cache update strategy based on the historical input frequency; A dynamic scheduling module, which determines whether to trigger the cache warm-up mechanism in the cache update strategy. If so, it directly retrieves the offline cache module; otherwise, it determines whether to directly retrieve the model distillation module based on the historical number of questions and the similarity of the question content. If not, it executes the constructed complexity evaluation model and selects to retrieve the offline cache module or the model distillation module according to the output result of the complexity evaluation model; The monitoring and feedback module, when retrieving the offline cache module, regularly collects the cache hit rate and feedback satisfaction, constructs a first linear calculation model, and generates an effect evaluation index; when retrieving the model distillation module, regularly collects the inference speed and feedback satisfaction, constructs a second linear calculation model, and generates an effect evaluation index; according to the time series, obtains the effect evaluation indexes in each period, and analyzes the change trend of the effect evaluation indexes. When there is a downward trend, corresponding optimization and adjustment measures are executed.

[0007] Furthermore, when collecting text data in the target domain, an automated crawler technology is combined with a screening mechanism. Among them, the screening mechanism at least includes: cleaning the text data collected by the crawler. When organizing the text data in the target domain, a data augmentation technology is adopted. Among them, the data augmentation technology includes: synonym replacement and sentence pattern transformation.

[0008] Furthermore, the process of determining whether to execute the adaptive distillation strategy is as follows: Monitor the performance of the student model: Regularly evaluate the performance of the student model during training, that is, the accuracy. Adjust the distillation intensity: According to the change in the accuracy of the student model, adjust the hyperparameters during the distillation process. Continue training: While adjusting the distillation intensity, continue to train the student model until convergence.

[0009] Furthermore, the set cache data structure includes any one of a hash table and a tree structure. The process of formulating a cache update strategy based on the historical input frequency is as follows: Obtain the occurrence frequency of each type of input from the historical data, compare each frequency with the set boundary threshold. When the frequency exceeds the boundary threshold, it is determined that the output of this type of input is a high-frequency input, and the cache preheating mechanism is synchronously triggered; when the frequency does not exceed the boundary threshold, it is determined that the output of this type of input is a low-frequency input. Among them, the cache preheating mechanism: When the system starts or restarts, pre-obtain the output of high-frequency inputs and store them in the cache; when the system receives a user request, read the output of high-frequency inputs from the cache.

[0010] Furthermore, the process of determination based on the historical number of questions and the similarity of question content is as follows: Condition 1: Record and monitor the number of questions of the user. When it is detected that the number of questions of the user is not less than the set first threshold, it is determined not to directly retrieve the model distillation module. Otherwise, condition 2 is executed. Condition 2: Calculate the similarity between the current user's question content and the training data of the model distillation module; if the similarity is higher than the set second threshold, it means that Condition 2 is satisfied; among them, the data in the target domain dataset is the training data; Directly call the model distillation module: When both Condition 1 and Condition 2 are satisfied, call the direct model distillation module.

[0011] Furthermore, the process of running the complexity evaluation model is as follows: Obtain the evaluation data collected when the current user asks a question, and the evaluation data includes the length of the input text and the lexical complexity. Based on the evaluation data after dimensionless processing, establish a weighted formula to calculate the complexity estimate: In the formula, CI represents the complexity estimate, α and β are both weight coefficients, and their value ranges are both [0, 1]; Tl represents the length of the input text, and Vc represents the lexical complexity.

[0012] Furthermore, compare the complexity estimate with the preset evaluation threshold: When the complexity estimate does not exceed the evaluation threshold, select to call the offline cache module; When the complexity estimate exceeds the evaluation threshold, select to call the model distillation module.

[0013] Furthermore, construct the first linear calculation model, and the formula is as follows: In the formula, represents the effect evaluation index in the state of calling the offline cache module, a1 represents the relative importance weight of the cache hit rate and the feedback satisfaction, and a1 ∈ (0, 1), H represents the cache hit rate, represents the feedback satisfaction in the state of the offline cache module, γ represents the first adjustment coefficient, and the value range is [0, 1]; Construct the second linear calculation model, and the formula is as follows: In the formula, represents the effect evaluation index in the state of calling the model distillation module, a2 represents the relative importance weight of the inference speed and the feedback satisfaction, and a2 ∈ (0, 1), P represents the inference speed, represents the feedback satisfaction in the state of the model distillation module, represents the second adjustment coefficient, and the value range is [0, 1]; Among them, before constructing the first linear calculation model, normalize the cache hit rate and the feedback satisfaction, and before constructing the second linear calculation model, normalize the inference speed and the feedback satisfaction.

[0014] Further, the corresponding optimization and adjustment measures are executed as follows: When the effect evaluation index in the state of retrieving the model distillation module shows a downward trend, the training process of the student model is optimized, including adjusting hyperparameters and increasing training data; when the effect evaluation index in the state of retrieving the offline cache module shows a downward trend, optimization is carried out in the offline cache module, including increasing storage capacity and refreshing storage.

[0015] A model optimization management method combining model distillation and offline caching includes the following steps: S1. Collect and organize text data in the target domain to construct a dataset for the target domain; S2. Use knowledge distillation to soften the output of the original model, train to obtain a student model, perform inter-layer distillation between the original model and the student model, and determine whether to execute an adaptive distillation strategy according to the performance state of the student model; S3. Set the cache data structure and formulate a cache update strategy based on the historical input frequency; S4. Determine whether to trigger the cache warm-up mechanism in the cache update strategy. If so, directly retrieve the offline cache module corresponding to S3; otherwise, based on the historical number of questions and the similarity of question content, determine whether to directly retrieve the model distillation module corresponding to S2; if not, execute the constructed complexity evaluation model, and select to retrieve the offline cache module or the model distillation module according to the output result of the complexity evaluation model; S5. In the state of retrieving the offline cache module, regularly collect the cache hit rate and feedback satisfaction, construct a first linear calculation model, and generate an effect evaluation index; in the state of retrieving the model distillation module, regularly collect the inference speed and feedback satisfaction, construct a second linear calculation model, and generate an effect evaluation index; obtain the effect evaluation index in each period according to the time series, and analyze the change trend of the effect evaluation index. When there is a downward trend, execute the corresponding optimization and adjustment measures.

[0016] (III) Beneficial effects The present invention provides a model optimization management system combining model distillation and offline caching, which has the following beneficial effects: (1) For the model output of high-frequency inputs in this solution, it can be preferentially stored in the cache and updated regularly to ensure its effectiveness; at the same time, for the model output of low-frequency inputs, replacement or deletion can be considered when the cache space is insufficient; by adopting the cache warm-up mechanism, the response speed and efficiency of the system in the initial stage of startup are significantly improved; at the same time, since the output of popular inputs has been pre-loaded into the cache, the real-time computing requirements can be further reduced, and the system load can be lowered; (2) This solution can select the most suitable processing module according to the actual situation of the user's question, thus achieving a fast and accurate answer to the question and improving the response efficiency of the system; when the number of user questions is small and the question content is highly similar to the training data of the model distillation module, the model distillation module is directly called, which not only saves cache resources but also ensures the accuracy of the answer; at the same time, for questions with lower complexity, the offline cache module is selected to reduce the overhead of model inference; (3) This solution introduces a complexity evaluation model. After two judgment processes, the system can then adaptively select the processing module according to the complexity of different questions, improving the flexibility and adaptability of the system; (4) This solution constructs a first linear calculation model and a second linear calculation model to quantitatively evaluate the effects of the offline cache module and the model distillation module respectively. These two models consider multiple key indicators such as cache hit rate, feedback satisfaction, and inference speed, and introduce a non-linear relationship to adjust the influence of each indicator, thus achieving a refined evaluation of the module effects; by regularly collecting the effect evaluation index and analyzing its change trend, problems with the decline in module performance can be detected in a timely manner. When it is detected that the effect evaluation index shows a downward trend, optimization and adjustment measures are triggered to perform targeted optimization on the offline cache module or the model distillation module to ensure that the system always maintains the best performance. Brief Description of the Drawings

[0017] Figure 1 It is a schematic diagram of the modular operation of the system in the present invention; Figure 2 It is a schematic diagram of the change of the effect evaluation index over time in the present invention. Detailed Embodiments

[0018] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0019] In the context of the widespread application of current large language models (LLMs), although the models have made significant progress in accuracy and generality, low inference efficiency and high resource consumption have become key problems restricting their further development; especially in scenarios where complex tasks or high real-time requirements are involved, the inference speed of large language models has become a bottleneck; To address this problem, this application proposes a concept of a model optimization management system that combines model distillation and offline cache technologies; By using model distillation technology, a smaller student model is trained to mimic the behavior of the original large language model, thereby improving the inference speed. At the same time, the offline caching technology is used to pre-compute and store the model outputs of common inputs, further reducing the computational requirements during the inference stage. In addition, this application also introduces creative technical points such as an adaptive distillation strategy and a cache warm-up mechanism to optimize the distillation effect and cache performance, so as to comprehensively improve the inference efficiency of the large language model while maintaining its high accuracy and generality. The specific scheme discussion can be referred to the subsequent content description.

[0020] Example 1:

[0021] Please refer to Figure 1 , this embodiment provides a model optimization management system that combines model distillation and offline caching. The system is divided into multiple functional modules and runs sequentially to ensure the efficiency and stability of the optimization process. The system includes a data preprocessing module, a model distillation module, an offline caching module, a dynamic scheduling module, and a monitoring and feedback module. The following is an explanation and elaboration of each module within the system: Data preprocessing module: Collect and organize the text data in the target domain to construct a dataset for the target domain. Among them, data preprocessing is a crucial step in machine learning and natural language processing (NLP) projects, which directly affects the effect and performance of subsequent model training. The technical content adopted by this module is as follows: using automated crawler technology combined with a screening mechanism to ensure the diversity, quality, and scale of the dataset. At the same time, data augmentation techniques such as synonym replacement and sentence pattern transformation are used to enrich the content of the dataset. Automated crawler technology combined with a screening mechanism Automated crawler technology: Purpose: To quickly and massively collect relevant text data from the Internet (i.e., the target domain). Implementation method: Use libraries such as requests, BeautifulSoup, and Scrapy in Python to write crawler programs. Define the crawling strategy, search according to keywords, and crawl the content of specific websites or forums. Set the crawler frequency and polite requests to avoid burdening the target website. Example: Suppose we want to build a dataset on the "artificial intelligence" field, we can write a crawler program to crawl relevant content from technology news websites, academic paper platforms, and AI community forums, etc. Screening mechanism: Purpose: Ensure the quality and relevance of the dataset; Implementation method: Conduct preliminary cleaning on the data collected by the crawler to remove irrelevant, duplicate, and low-quality content; Set screening criteria such as text length, keyword density, content relevance, etc. to achieve further screening work; Example: In the "Artificial Intelligence" dataset, screening can remove articles that only mention AI briefly but lack in-depth content, and retain high-quality texts that deeply explore AI technologies, applications, or trends; Data augmentation techniques Synonym replacement: Purpose: Increase the diversity of text data and improve the model's generalization ability for synonyms; Implementation method: Use resources such as WordNet and synonym dictionaries to find synonyms of words; Randomly select some words in the text for synonym replacement; Example: "Artificial Intelligence" can be replaced with "AI", "Intelligent Technology", etc.

[0022] Sentence transformation: Purpose: Change the expression of the text and increase the richness of the dataset; Implementation method: Use NLP tools (such as NLTK, spaCy) for syntactic analysis; Transform sentences according to the syntactic structure, such as changing active sentences to passive sentences, declarative sentences to interrogative sentences, etc.; Example: "Artificial Intelligence is changing the world." can be transformed into "The world is being changed by Artificial Intelligence." or "Isn't Artificial Intelligence changing the world?"; Other data augmentation techniques: Adding noise: Insert some random characters, words, or phrases into the text to simulate spelling mistakes or colloquial expressions in real scenarios; Back translation: Translate the text into another language and then translate it back to obtain text that is semantically similar but has different expressions; Summary generation: Extract key information from long texts to generate summaries as new training samples.

[0023] Model distillation module: Adopt knowledge distillation to soften the output of the original model, train to obtain a student model, conduct layer-by-layer distillation between the original model and the student model, and determine whether to execute the adaptive distillation strategy according to the performance status of the student model; Among them, the model distillation module is an effective model compression and optimization technology, aiming to train a smaller student model to imitate the behavior of the original large language model (teacher model); Knowledge distillation Principle: Knowledge distillation is a technology that transfers the knowledge of a large and complex model (teacher model) to a small and simple model (student model); Its core idea is to soften the output of the teacher model so that it contains more information, and then use this information as the learning target for the student model; The implementation steps are as follows: Training the teacher model: First, extract a large amount of data (i.e., training data) from the dataset in the target domain to train a teacher model with excellent performance. This model is usually a large deep neural network that can achieve a high accuracy rate on the given task; Softening the output of the teacher model: Use the softmax function in the output layer of the teacher model and soften the output probability distribution by increasing the temperature parameter T, so that the output can contain more information rather than just the final classification result; Training the student model: Use the softened output of the teacher model as the supervision signal to train the student model; The goal of the student model is to minimize the difference between its output and the output of the teacher model. Example: Suppose we have a large BERT model as the teacher model and want to distill its knowledge into a smaller LSTM model as the student model; First, we train the BERT model using a large corpus, and then increase the temperature parameter T of the softmax function in the inference stage to soften the output probability distribution; Next, we use these softened outputs as the supervision signal to train the LSTM model so that it can mimic the behavior of the BERT model. Intermediate layer distillation Principle: Intermediate layer distillation is a technique for establishing a connection between the intermediate layers of the teacher model and the corresponding layers of the student model; By matching the outputs of the intermediate layers of the two models, the distillation efficiency can be further improved and help the student model better learn the internal representation of the teacher model. Implementation steps: Select intermediate layers: Select the corresponding intermediate layers in the teacher model and the student model for distillation; These layers usually contain important feature representations; Calculate the loss function: For the selected intermediate layers, calculate the difference between the outputs of the teacher model and the student model and use it as part of the loss function; Joint training: Combine the loss function of the intermediate layer with the loss function of the output layer to jointly train the student model.

[0024] Example: In the distillation process from BERT to LSTM, we can select several intermediate Transformer layers of BERT and the corresponding layers of the student LSTM model for distillation; For each selected layer, we calculate the mean squared error (MSE) between the output of the BERT layer and the output of the student LSTM layer as part of the loss function; Then, combine these MSE losses with the cross-entropy loss of the output layer to jointly train the LSTM model. Adaptive distillation strategy Principle: The adaptive distillation strategy is a technique for dynamically adjusting the distillation intensity according to the learning progress of the student model; By monitoring the performance change of the student model, the hyperparameters or loss function weights in the distillation process can be adjusted timely to avoid overfitting and improve the distillation effect. Implementation steps: Monitor the performance of the student model: Regularly evaluate the performance of the student model during training, such as accuracy, loss function value, etc. (in this embodiment, accuracy is selected, and it can be known by referring to the following examples); Adjust the distillation intensity: Dynamically adjust the hyperparameters (such as the temperature parameter T) or loss function weights during the distillation process according to the performance changes of the student model; For example, when the performance of the student model improves slowly, the distillation intensity can be appropriately increased; When the student model is approaching overfitting, the distillation intensity can be appropriately reduced; Continue training: While adjusting the distillation intensity, continue to train the student model until convergence; Example: During the distillation process from BERT to LSTM, we can set up a monitoring mechanism to regularly evaluate the performance of the LSTM model; If it is found that the accuracy of the LSTM model improves slowly (that is, if the improvement rate of the accuracy within a unit time is lower than the set threshold for 3 consecutive times, it means slow improvement), we can choose to increase the temperature parameter T (increase it according to the set step value) to soften the output of the BERT model, thereby increasing the amount of information in the distillation process; On the contrary, if it is found that the LSTM model begins to show signs of overfitting (such as the accuracy on the training set increases but the accuracy on the validation set decreases), we can appropriately reduce the temperature parameter T (reduce it according to the set step value) to slow down the distillation intensity; In summary, the model distillation module effectively transfers the knowledge of large and complex models to small and simple models through technical means such as knowledge distillation, inter-layer distillation, and adaptive distillation strategies; These technologies not only reduce the complexity and computational cost of the model, but also improve the generalization ability and practical application effect of the model; In practical applications, appropriate distillation technologies and parameter settings can be selected according to the characteristics of specific tasks and datasets to optimize the model performance.

[0025] Offline cache module: Set the cache data structure and formulate a cache update strategy based on the historical input frequency; Among them, the main function of the offline cache module is to pre-compute and store the model outputs of common inputs, so that the cached results can be directly called during the inference stage, reducing the real-time computing requirements and significantly improving the system response speed and efficiency; The set cache data structure includes any one of a hash table and a tree structure. In this embodiment, the hash table is taken as an example; Hash table: A hash table is an efficient key-value pair storage structure that maps inputs to specific storage locations through a hash function to achieve fast lookup. In the offline cache module, the set regular inputs are used as keys, and the corresponding model outputs are stored in the hash table as values; When inference is required, only the corresponding output needs to be quickly found through the hash table without re-computation; Tree structure: In some cases, to improve the search efficiency or handle complex input relationships, a tree structure can be adopted as the cache data structure; for example, a prefix tree (Trie tree) can be constructed, with the prefixes of the input strings as the nodes of the tree and the corresponding model outputs stored in the leaf nodes; in this way, when searching, one only needs to gradually match the input string along the tree structure to quickly find the corresponding output; The process of formulating the cache update policy based on the historical input frequency is as follows: Obtain the occurrence frequencies of each type of input from the historical data, compare each frequency with the set boundary threshold. When the frequency exceeds the boundary threshold, it is determined that the model output of this type of input is a high-frequency input (i.e., a popular input), and the cache preheating mechanism is triggered synchronously; when the frequency does not exceed the boundary threshold, it is determined that the model output of this type of input is a low-frequency input; Among them, the cache preheating mechanism: when the system starts or restarts, the model is used to pre-compute the outputs of high-frequency inputs and store them in the cache; when the system starts to receive user requests, the outputs of high-frequency inputs can be directly read from the cache without real-time calculation; to reduce the initial response latency; It should be noted that for the model outputs of high-frequency inputs, they can be preferentially stored in the cache and updated regularly to ensure their validity; at the same time, for the model outputs of low-frequency inputs, replacement or deletion can be considered when the cache space is insufficient; by adopting the cache preheating mechanism, the response speed and efficiency of the system in the initial stage are significantly improved; at the same time, since the outputs of popular inputs have been pre-loaded into the cache, the real-time calculation requirements can be further reduced and the system load can be lowered; Dynamic scheduling module: Judge whether to trigger the cache preheating mechanism in the cache update policy. If so, directly call the offline cache module for answering; if not, based on the historical number of questions and the similarity of the question content, determine whether to directly call the model distillation module; if not, execute the constructed complexity evaluation model, and select to call the offline cache module or the model distillation module according to the output result of the complexity evaluation model; among them, if no result can be output by calling any module, replace it with the other module; Among them, the process of determination based on the historical number of questions and the similarity of the question content is as follows: Monitoring of the number of questions: Record and monitor the number of questions of the user. When it is detected that the number of questions of the user is not less than the set first threshold, it is determined as no (i.e., not directly call the model distillation module). When it is detected that the number of questions of the user is lower than the set first threshold, proceed to the next step; Calculation of the similarity of the question content: Calculate the similarity between the current question content of the user and the training data of the model distillation module, which is specifically implemented through text similarity algorithms (such as cosine similarity, Jaccard similarity, etc.); if the similarity is higher than the set second threshold, it indicates that the question content is highly similar to the training data and meets this condition; if the similarity is not higher than the set second threshold, it means that the similarity between the question content and the training data is not high.

[0026] Directly invoke the model distillation module: Under the condition of meeting the above two conditions, directly invoke the model distillation module to answer; there is no need to further consider the complexity of the question, because the model distillation module has been optimized for similar questions.

[0027] Specifically, when the number of questions is small but the question content is highly similar to the training data of the model distillation module, we can directly invoke the model distillation module to answer. This approach is both efficient and accurate, and can avoid the cold start problem and resource waste, which is particularly important for resource-constrained environments or application scenarios that require quick responses.

[0028] The process of running the complexity evaluation model is as follows: Obtain the evaluation data collected when the current user asks a question, and the evaluation data includes the length of the input text and the lexical complexity. Based on the evaluation data after dimensionless processing, establish a weighted formula to calculate the complexity estimate: In the formula, CI represents the complexity estimate, α and β are both weight coefficients, and their value ranges are both [0, 1]; Tl represents the length of the input text, and Vc represents the lexical complexity. For the lexical complexity, it can be measured in various ways, such as the reciprocal of the average word frequency, the lexical diversity index (such as type-token ratio), or the score based on the lexical difficulty level. In this embodiment, the lexical diversity index (such as type-token ratio) is used as the lexical complexity; Compare the complexity estimate with the preset evaluation threshold; When the complexity estimate does not exceed the evaluation threshold, it means that the complexity of the input text is relatively low, and the offline cache module is selected to be invoked; when the complexity estimate exceeds the evaluation threshold, it means that the complexity of the input text is relatively high, and the model distillation module is selected to be invoked; it should be noted that the evaluation threshold is set independently according to the actual situation and is used to judge whether the complexity estimate exceeds a predetermined level; Specifically, the system can intelligently select the most appropriate processing module (offline cache module or model distillation module) according to the actual situation of the user's question (such as the number of questions, content similarity, text complexity), thus achieving a fast and accurate answer to the question and improving the response efficiency of the system; Resource optimization: This solution avoids unnecessary resource waste; when the number of user questions is small and the question content is highly similar to the training data of the model distillation module, the model distillation module is directly invoked, which not only saves cache resources but also ensures the accuracy of answers; at the same time, for questions with low complexity, the offline cache module is selected to reduce the overhead of model inference; Cold start problem mitigation: Through the cache preheating mechanism and the judgment based on the number of questions and content similarity, the system can still give relatively accurate answers when users use it for the first time or ask few questions, thus mitigating the cold start problem; Strong adaptability: The introduction of the complexity evaluation model enables the system to adaptively select processing modules according to the complexity of different questions, improving the flexibility and adaptability of the system; In summary, this technical solution realizes the intelligent allocation and optimal utilization of processing resources, improves the response speed, accuracy and adaptability of the system, and effectively solves the problems of unreasonable resource allocation, slow response speed, cold start problems and inaccurate complexity evaluation existing in traditional systems.

[0029] Monitoring and feedback module: In the state of invoking the offline cache module, regularly collect the cache hit rate and feedback satisfaction, construct the first linear calculation model, and generate the effect evaluation index; in the state of invoking the model distillation module, regularly collect the inference speed and feedback satisfaction, construct the second linear calculation model, and generate the effect evaluation index; according to the time series, obtain the effect evaluation indexes in each period, and analyze the change trend of the effect evaluation indexes. When there is a downward trend, execute the corresponding optimization and adjustment measures; Among them, before constructing the first linear calculation model, normalize the collected cache hit rate and feedback satisfaction so that the values of the two types of data are both between 0 and 1. The formula is as follows: In the formula, represents the effect evaluation index in the state of invoking the offline cache module. The subscript c indicates the state of the offline cache module. a1 represents the relative importance weight of the cache hit rate and feedback satisfaction, and a1 ∈ (0, 1). H represents the cache hit rate, represents the feedback satisfaction in the state of the offline cache module. γ represents the first adjustment coefficient, reflecting the non-linear influence of the cache hit rate on the feedback satisfaction, and the value range is [0, 1]; The logical explanation of the formula is as follows: reflects the linear contribution of the cache hit rate to the effect evaluation index and serves as the linear part; A non - linear relationship is introduced, where γ regulates the impact of cache hit rate on feedback satisfaction; when the cache hit rate H is low, the denominator decreases, thus amplifying the impact of feedback satisfaction; when the cache hit rate is high, the denominator approaches 1, and the impact of feedback satisfaction is relatively weakened; Before constructing the second - order linear calculation model, the collected inference speed and feedback satisfaction are normalized so that the values of both types of data are between 0 and 1. The formula is as follows: In the formula, represents the effect evaluation index in the state of invoking the model distillation module. The subscript m in the lower right indicates the state of the model distillation module. a2 represents the relative importance weight of the inference speed and feedback satisfaction, and a2 ∈ (0, 1). P represents the inference speed, represents the feedback satisfaction in the state of the model distillation module, represents the second - order adjustment coefficient, reflecting the non - linear enhancement or weakening effect of the inference speed on the feedback satisfaction; The logical explanation of the formula is as follows: reflects the linear contribution of the inference speed to the effect evaluation index and serves as the linear part; introduces the non - linear enhancement or weakening effect of the inference speed on the feedback satisfaction; when the inference speed p is higher than 0.5, is positive, enhancing the impact of feedback satisfaction; when the inference speed is lower than 0.5, it weakens the impact of feedback satisfaction; Refer to Figure 2 As shown, regular collection means collecting within the corresponding period K. The value of period K is usually one day or one week (i.e., seven days). The numbers of K are 1, 2, 3, and 4. When analyzing the change trend of the effect evaluation index, the analysis period is at least four periods K. The abscissa in the figure represents the number of periods, and the ordinate represents the value of the effect evaluation index. And the chart is only an example. If it is seen from the chart that the effect evaluation index in the state of invoking the model distillation module shows a downward trend, corresponding optimization and adjustment measures need to be implemented; When analyzing the change trend of the effect evaluation index, data analysis software can be used for self - analysis; Data analysis software Excel: As the most fundamental data analysis tool, Excel provides rich data processing and chart drawing functions, which can be used to initially analyze the changing trends of the effect evaluation index. By drawing line charts, scatter plots, etc., the changes in the index can be visually observed; SPSS: This is a professional statistical analysis software suitable for more complex data analysis requirements. SPSS provides various statistical methods, such as time series analysis, regression analysis, etc., which can help deeply explore the reasons and rules behind the index changes; Python: Python is a powerful programming language with numerous data analysis libraries and tools, such as Pandas, NumPy, Matplotlib, etc.; By writing Python scripts, automated data analysis and chart drawing can be achieved, which is very suitable for processing large-scale data sets and complex analysis tasks. In this embodiment, Excel can be directly used for analysis to obtain the trend results.

[0030] The process of performing the corresponding optimization and adjustment measures is as follows: When the effect evaluation index in the state of invoking the model distillation module shows a downward trend, the training process of the student model is optimized, including adjusting hyperparameters and increasing any one of the training data to further improve the performance of the model; When the effect evaluation index in the state of invoking the offline cache module shows a downward trend, the storage capacity is increased and the storage is refreshed. The former is to store more data to improve the cache hit rate; The latter is to avoid the data in the cache from becoming obsolete, and the cache can be refreshed regularly to ensure that the data in the cache is the latest; The specific adjustment amount is based on the actual requirements and the set amount each time. That is, if the set amount of increased training data in the system's rule engine is 100G each time, then the increased training data amount is 100G.

[0031] Specifically, by constructing the first linear calculation model and the second linear calculation model, the effects of the offline cache module and the model distillation module are quantitatively evaluated respectively. These two models consider multiple key indicators such as cache hit rate, feedback satisfaction, inference speed, etc., and introduce non-linear relationships to adjust the influence of each indicator, thus realizing the refined evaluation of the module effects; Dynamic optimization and adjustment: The solution can timely detect the problem of module performance decline by regularly collecting the effect evaluation index and analyzing its changing trend. When it is detected that the effect evaluation index shows a downward trend, the optimization and adjustment measures are triggered to specifically optimize the offline cache module or the model distillation module to ensure that the system always maintains the best performance; Improve user satisfaction: Through refined evaluation and optimization and adjustment, the system can more accurately meet user needs, improve the cache hit rate and inference speed, thereby enhancing user satisfaction; At the same time, considering the user's feedback satisfaction as one of the evaluation indicators also further reflects the user-centered design concept; In summary, through improvements in aspects such as refined effect evaluation, dynamic optimization adjustment, enhancing user satisfaction, and strengthening system flexibility, this technical solution solves the problems existing in traditional systems, such as inaccurate effect evaluation, untimely optimization adjustment, and low user satisfaction.

[0032] Embodiment 2:

[0033] Based on Embodiment 1, this embodiment also provides a model optimization management method combining model distillation and offline caching, including the following specific steps: S1. Collect and organize text data in the target domain to construct a dataset for the target domain; S2. Use knowledge distillation to soften the output of the original model, train to obtain a student model, perform layer-by-layer distillation between the original model and the student model, and determine whether to execute the adaptive distillation strategy according to the performance status of the student model; S3. Set the cache data structure and formulate a cache update strategy based on the historical input frequency; S4. Determine whether to trigger the cache warm-up mechanism in the cache update strategy. If so, directly call the offline cache module corresponding to S3; otherwise, based on the historical number of questions and the similarity of question content, determine whether to directly call the model distillation module corresponding to S2; if not, execute the constructed complexity evaluation model, and select to call the offline cache module or the model distillation module according to the output result of the complexity evaluation model; S5. In the state of calling the offline cache module, regularly collect the cache hit rate and feedback satisfaction, construct a first linear calculation model, and generate an effect evaluation index; in the state of calling the model distillation module, regularly collect the inference speed and feedback satisfaction, construct a second linear calculation model, and generate an effect evaluation index; obtain the effect evaluation indexes in each period according to the time series, and analyze the change trend of the effect evaluation indexes. When there is a downward trend, execute the corresponding optimization adjustment measures.

[0034] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or in the combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution.

[0035] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, and it may be located in one place or distributed across multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0036] As described above, the foregoing are only specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application.

Claims

1. A model optimization management system combining model distillation and offline caching, characterized in that, The system includes: A data preprocessing module that collects and organizes text data in the target domain and constructs a dataset for the target domain; A model distillation module that softens the output of the original model using knowledge distillation, trains to obtain a student model, performs inter-layer distillation between the original model and the student model, and determines whether to execute an adaptive distillation strategy based on the performance status of the student model; An offline cache module that sets a cache data structure and formulates a cache update strategy based on historical input frequencies; A dynamic scheduling module that determines whether to trigger the cache warm-up mechanism in the cache update strategy. If so, it directly invokes the offline cache module; otherwise, it determines whether to directly invoke the model distillation module based on the historical number of questions and the similarity of the question content. If not, it executes the constructed complexity evaluation model and selects to invoke the offline cache module or the model distillation module according to the output result of the complexity evaluation model; A monitoring and feedback module that, in the state of invoking the offline cache module, regularly collects the cache hit rate and feedback satisfaction, constructs a first linear calculation model, and generates an effect evaluation index; in the state of invoking the model distillation module, regularly collects the inference speed and feedback satisfaction, constructs a second linear calculation model, and generates an effect evaluation index; obtains the effect evaluation indexes in each period according to the time series, analyzes the change trend of the effect evaluation indexes, and executes corresponding optimization and adjustment measures when there is a downward trend.

2. The model optimization management system combining model distillation and offline caching according to claim 1, characterized in that: When collecting text data in the target domain, an automated crawler technology is combined with a screening mechanism; Among them, the screening mechanism at least includes: cleaning the text data collected by the crawler; When organizing the text data in the target domain, a data augmentation technology is used; Among them, the data augmentation technology includes: synonym replacement and sentence pattern transformation.

3. The model optimization management system combining model distillation and offline caching according to claim 1, characterized in that: The process of determining whether to execute the adaptive distillation strategy is as follows: Monitor the performance of the student model: Regularly evaluate the performance of the student model during training, that is, the accuracy; Adjust the distillation intensity: Adjust the hyperparameters in the distillation process according to the change in the accuracy of the student model; Continue training: While adjusting the distillation intensity, continue to train the student model until convergence.

4. The model optimization management system combining model distillation and offline caching according to claim 1, characterized in that: The set cache data structure includes any one of a hash table and a tree structure; The process of formulating a cache update strategy based on historical input frequencies is as follows: Obtain the occurrence frequency of each type of input from historical data, compare each frequency with the set threshold. When the frequency exceeds the threshold, it is determined that the model output of this type of input is a high-frequency input, and the cache warm-up mechanism is synchronously triggered; When the frequency does not exceed the threshold, it is determined that the model output of this type of input is a low-frequency input; Among them, the cache warm-up mechanism: When the system starts or restarts, pre-obtain the output of high-frequency inputs and store them in the cache; When the system receives a user request, read the output of high-frequency inputs from the cache.

5. The model optimization management system combining model distillation and offline caching according to claim 1, wherein: The process of making a determination based on the historical number of questions and the similarity of the question content is as follows: Condition 1: Record and monitor the number of questions of the user. When it is detected that the number of questions of the user is not less than the set first threshold, it is determined not to directly invoke the model distillation module. Otherwise, execute Condition 2; Condition 2: Calculate the similarity between the user's current question content and the training data of the model distillation module; if the similarity is higher than the set second threshold, it means that Condition 2 is satisfied; among them, the data in the dataset of the target domain is the training data; Directly call the model distillation module: When both Condition 1 and Condition 2 are satisfied, call the direct model distillation module.

6. The model optimization management system combining model distillation and offline caching according to claim 1, wherein: The process of running the complexity evaluation model is as follows: Obtain the evaluation data collected when the current user asks a question, and the evaluation data includes the length of the input text and the lexical complexity. Based on the evaluation data after dimensionless processing, establish a weighted formula to calculate the complexity estimate: In the formula, CI represents the complexity estimate, α and β are both weight coefficients, and their value ranges are both [0, 1]; Tl represents the length of the input text, and Vc represents the lexical complexity.

7. The model optimization management system combining model distillation and offline caching according to claim 6, characterized in that: Compare the complexity estimate with the preset evaluation threshold: When the complexity estimate does not exceed the evaluation threshold, select to call the offline cache module; When the complexity estimate exceeds the evaluation threshold, select to call the model distillation module.

8. The model optimization management system combining model distillation and offline caching according to claim 1, characterized in that: Construct the first linear calculation model, and the formula is as follows: In the formula, represents the effect evaluation index in the state of the offline cache module being retrieved. a1 represents the relative importance weight of the cache hit rate and the feedback satisfaction degree, and a1 ∈ (0, 1). H represents the cache hit rate. represents the feedback satisfaction degree in the state of the offline cache module. γ represents the first adjustment coefficient, and its value range is [0, 1]. Construct the second linear calculation model, and the formula is as follows: In the formula, represents the effect evaluation index in the state of invoking the model distillation module, a2 represents the relative importance weight of the inference speed and the feedback satisfaction degree, and a2 ∈ (0, 1), P represents the inference speed, represents the feedback satisfaction degree in the state of the model distillation module, represents the second adjustment coefficient, and the value range is [0, 1]; Among them, before constructing the first linear calculation model, normalize the cache hit rate and feedback satisfaction, and before constructing the second linear calculation model, normalize the inference speed and feedback satisfaction.

9. The model optimization management system combining model distillation and offline caching according to claim 8, wherein: The process of executing the corresponding optimization and adjustment measures is as follows: When the effectiveness evaluation index in the state of calling the model distillation module shows a downward trend, optimize the training process of the student model, including adjusting hyperparameters and increasing training data; When the effectiveness evaluation index in the state of calling the offline cache module shows a downward trend, optimize in the offline cache module, including increasing the storage capacity and refreshing the storage.

Citation Information

Patent Citations

  • An output optimization method for large language models

    CN117573846B