A model output optimization control method and device based on entropy and scoring
Through the entropy and scoring method, the hyperparameters of the large language model are dynamically adjusted, which solves the problem of uncertainty in the output of the large model, and realizes output optimization and control without changing the model structure and parameters.
Patent Information
- Application Number
- CN202510188677.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-02-20
AI Technical Summary
There is uncertainty in the output of large models, and it is difficult to reduce its uncertainty by directly adjusting the model structure and parameters.
The model output optimization control method based on entropy and scoring is adopted. By selecting the appropriate propt prompt word, multiple result texts are generated, the result text is scored using the scoring model, and the entropy value of the highest-scoring result text is calculated. The hyperparameters of the large language model are dynamically adjusted to control the uncertainty of the output results.
Without changing the model structure and parameters, dynamically adjusting the hyperparameters can reduce the uncertainty of the model output and improve the accuracy and reliability of the output.
Smart Images

Figure CN119670898B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of model output control, and in particular to a model output optimization control method and device based on entropy and scoring. Background Art
[0002] In today's era of rapid technological development, with the continuous evolution of artificial intelligence technology, the application of big models has shown an explosive growth trend. Especially in many fields such as writing and question-answering, big models are playing an increasingly important role. In terms of writing, it can assist creators in quickly generating articles of various genres. Whether it is news reports, the framework of academic papers, or the creative conception of novel stories, big models can provide rich ideas and diverse expressions, greatly improving the efficiency of creation. In the question-and-answer scenario, it can quickly give more detailed answers to various questions raised by users with its strong knowledge reserves and semantic understanding capabilities, answering questions for users and becoming a convenient way for people to obtain information.
[0003] However, it cannot be ignored that the output of large models is not perfect and there is significant uncertainty. Due to its large parameter scale and complex internal structure, it is extremely difficult to modify the model itself. From a technical perspective, the training process of large models involves highly complex algorithms and massive data processing. Its internal parameters are interrelated and influence each other, and one move can affect the whole body. For ordinary users or general developers, it is almost an impossible task to train large models on their own. This is not only because the training process requires the collection, organization and processing of a large amount of data, but also because the requirements for computing power are extremely demanding, often requiring the strong support of specialized high-performance computing clusters or cloud computing resources, which undoubtedly poses a huge obstacle in terms of time, cost and technical capabilities.
[0004] In view of this, in order to improve the reliability and effectiveness of large models in practical applications, it is urgent to explore other innovative ways to optimize the control of model output to reduce the uncertainty of model output. Summary of the invention
[0005] In view of this, an embodiment of the present invention provides a model output optimization control method and device based on entropy and scoring, so as to solve the problem of uncertainty in the output of existing large models.
[0006] The technical solution adopted by the present invention is:
[0007] In a first aspect, the present invention provides a model output optimization control method based on entropy and scoring, comprising:
[0008] Select corresponding prompt words according to the application scenario, input the prompt words into the large language model, and control the large language model to generate N result texts through preset model parameters;
[0009] The scoring model is used to score the N result texts respectively to obtain N result text scores;
[0010] Select the result text with the highest score from the N result text scores, and calculate the entropy value of the result text with the highest score;
[0011] Based on the entropy value of the result text with the highest score, the hyperparameters of the large language model output are adjusted to control the uncertainty of the output result of the large language model; the hyperparameters of the large language model output include temperature and sampling strategy top_k.
[0012] Furthermore, the method of selecting a corresponding prompt word according to the application scenario, inputting the prompt word into the large language model, and controlling the large language model to generate N result texts through preset model parameters includes:
[0013] According to the application scenario of the large language model, select a prompt word that matches the application scenario from the preset prompt word group;
[0014] The model parameters of the large language model are preconfigured, and the selected prompt word is input into the large language model. The large language model is controlled by the model parameters to generate N different result texts, and the N different result texts are stored in the database according to the preset text format.
[0015] Furthermore, the scoring model is used to score the N result texts respectively to obtain N result text scores, including:
[0016] Select the open source language model as the scoring model, and use a 100-point scoring system as the scoring system;
[0017] The prompt word and the corresponding N result texts are input into the scoring model, and the result text scores corresponding to the N result texts are generated according to the scoring system.
[0018] Furthermore, the scoring model can be fine-tuned based on the open source model:
[0019] Collecting business question-answer pairs and historical question-answer pairs of a large language model from a third-party knowledge base, and constructing a first data set based on the business question-answer pairs and the historical question-answer pairs; the first data set includes question texts and original answer texts;
[0020] Input the question text in the first data set into the open source language model to generate multiple answer texts;
[0021] The answer text is combined with the original answer text, and the answer text and the original answer text are separated by a separator, and the answer text and the original answer text are combined with the question text to generate a second data set;
[0022] Manually scoring the answer texts in the second data set according to a preset scoring standard, and constructing a third data set based on the manual scoring results and the second data set;
[0023] Taking the question text and answer text in the third data set as input and the manual scoring results as output, the large language model is fine-tuned, and the model parameters of the large language model are optimized using the cross entropy loss function to minimize the scoring error rate of the large language model and obtain a fine-tuned scoring model.
[0024] Furthermore, the step of selecting the result text with the highest score from the N result text scores and calculating the entropy value of the result text with the highest score comprises:
[0025] Sort the scores of N result texts and select the result text with the highest score;
[0026] Assume that the probability of any word in the highest-scoring result text appearing in the large language model is , calculate the entropy value of each word token in the result text with the highest score:
[0027] ;
[0028] in, is the word token; The number of all possible values of the word token; The final output is The probability of a possible value; is the log function;
[0029] Multiply the entropy values of all word tokens to get the entropy value of the result text with the highest score E :
[0030] ;
[0031] in, is the number of word tokens in the result text with the highest score; Indicates The entropy of a word token.
[0032] Furthermore, the hyperparameters output by the large language model are adjusted based on the entropy value of the result text with the highest score to control the uncertainty of the output result of the large language model, including:
[0033] Determine whether the entropy value of the result text with the highest score exceeds a preset entropy value, and if the entropy value of the result text does not exceed the preset entropy value, keep the temperature and sampling strategy top_k output by the large language model unchanged;
[0034] If the entropy value of the result text exceeds the preset entropy value, the temperature and sampling strategy top_k of the large language model output are increased.
[0035] In a second aspect, the present invention provides a model output optimization control device based on entropy and scoring, comprising:
[0036] A text generation module, used to select corresponding prompt words according to the application scenario, input the prompt words into the large language model, and control the large language model to generate N result texts through preset model parameters;
[0037] A text scoring module is used to score the N result texts respectively using a scoring model to obtain N result text scores;
[0038] An entropy value calculation module is used to select the result text with the highest score from the N result text scores, and calculate the entropy value of the result text with the highest score;
[0039] The output optimization module is used to adjust the hyperparameters of the large language model output based on the entropy value of the result text with the highest score to control the uncertainty of the output result of the large language model; the hyperparameters of the large language model output include temperature and sampling strategy top_k.
[0040] In summary, the beneficial effects of the present invention are as follows:
[0041] The present invention provides a model output optimization control method based on entropy and scoring. The method scores the result text output by a large language model and calculates the entropy value of the result text with the highest score. Without changing the structure and parameters of the model itself, the hyperparameters of the model are dynamically adjusted to optimize and control the model output, thereby reducing the uncertainty of the model output. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] In order to more clearly illustrate the technical solution of the embodiment of the present invention, the following is a brief introduction to the drawings required for use in the embodiment of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work, and these are all within the protection scope of the present invention.
[0043] Figure 1 A flow chart of a model output optimization control method based on entropy and scoring according to the present invention;
[0044] Figure 2This is a functional module diagram of a model output optimization control device based on entropy and scoring according to the present invention. DETAILED DESCRIPTION
[0045] In order to make the purpose, technical solution and advantages of the embodiment of the present invention clearer, the technical solution in the embodiment of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiment of the present invention. It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. If there is no conflict, the various features of the present invention and the embodiments can be combined with each other, all within the scope of protection of the present invention.
[0046] In order to improve the reliability and effectiveness of large models in practical applications, it is urgent to explore other innovative ways to optimize and control model output. It cannot rely solely on direct adjustment of model structure and parameters, but should be based on the post-processing of model output results, multi-model fusion strategy, rule-based constraint guidance and other perspectives to develop a series of targeted optimization methods. For example, a set of perfect result evaluation system can be established to perform multi-dimensional quality evaluation on the output generated by the large model, and then a specific algorithm can be used to screen, correct and improve it according to the evaluation results, so as to effectively reduce the uncertainty of the output without changing the core architecture and parameters of the model itself, improve the accuracy, logic and practicality of the output, so as to better meet the diverse needs of users in different application scenarios, and further expand the deep application and sustainable development of large models in various fields. Therefore, the present invention provides a model output optimization control method based on entropy and scoring, which is customized for professional fields and scenarios, and by scoring the model output and calculating the entropy value, the hyperparameters of the model are dynamically adjusted, and the optimization and control of the model output are realized without changing the structure and parameters of the model itself, so as to improve the effect of the model.
[0047] The detailed implementation process of the present invention is shown in the following embodiments.
[0048] Example 1: Reference Figure 1 As shown, Figure 1 This is a flow chart of a model output optimization control method based on entropy and scoring of the present invention. Figure 1 As shown, the model output optimization control method based on entropy and scoring of the present invention includes:
[0049] Select corresponding prompt words according to the application scenario, input the prompt words into the large language model, and control the large language model to generate N result texts through preset model parameters;
[0050] The scoring model is used to score the N result texts respectively to obtain N result text scores;
[0051] Select the result text with the highest score from the N result text scores, and calculate the entropy value of the result text with the highest score;
[0052] Based on the entropy value of the result text with the highest score, the hyperparameters of the large language model output are adjusted to control the uncertainty of the output result of the large language model; the hyperparameters of the large language model output include temperature and sampling strategy top_k.
[0053] Furthermore, in an embodiment of the present invention, a corresponding prompt word is selected according to an application scenario, and the prompt word is input into a large language model, and the large language model is controlled by preset model parameters to generate N result texts, including:
[0054] According to the application scenario of the large language model, select a prompt word that matches the application scenario from the preset prompt word group;
[0055] The model parameters of the large language model are preconfigured, and the selected prompt word is input into the large language model. The large language model is controlled by the model parameters to generate N different result texts, and the N different result texts are stored in the database according to the preset text format.
[0056] Specifically, in the embodiment of the present invention, a suitable prompt is selected according to the usage scenario, which can be a short sentence, a question or an instruction, to guide the large language model to generate the content you expect. At the same time, the parameters of the large language model are set, such as temperature (to control the randomness of the generated results), top_k, maximum number of generation times (i.e. The prompt is input into the large language model and the generation process is run. Based on the set parameters, the model will generate different results. The results are collected and saved in an appropriate format (such as text file, database, etc.) for subsequent analysis and use.
[0057] The prompt is designed according to the application scenario. Different application scenarios use different prompts. For example, for the writing scenario, the prompt is designed as "Write an argumentative essay on the topic 'Why is the sky blue' in a formal and serious style, no less than 800 words"; for the question-and-answer scenario, the prompt is designed as "Why is the sky blue? Please answer the above questions in concise and rigorous language."
[0058] When generating results based on the same prompt multiple times, the following situations usually occur:
[0059] 1. Similar but with slight differences:
[0060] For example, if the prompt "Describe the scenery of the park in spring" is input into the large language model, the generated results may mention common elements such as flowers blooming in the park, green grass, and birds singing on the branches. However, there will be differences in the details such as the sentence expression and the order of the specific types of flowers listed. One time it may say "Pink peach blossoms and white apricot blossoms are in full bloom", and another time it may be "White apricot blossoms and pink peach blossoms make the park particularly beautiful."
[0061] 2. Slightly different angles:
[0062] Similarly, for the above prompt, some generated results may focus on describing the vibrant and colorful feeling of the park's spring scenery from an overall perspective, such as "the park in spring is like a colorful painting, brimming with vitality everywhere, the grass is like a green carpet covering the earth, and the flowers are in full bloom, which is beautiful beyond words"; while some results may focus on the perspective of tourists' experience, such as "walking into the spring park, the breeze blows, bringing bursts of floral fragrance, and the ears are filled with the cheerful singing of birds. People stroll leisurely among the flowers and on the paths, enjoying this beautiful spring time."
[0063] 3. Occasionally, there are large differences:
[0064] Since the results generated by the large language model have certain random factors, in a few cases, the content generated based on the same prompt may be quite different. For example, sometimes it may focus more on the lake in the park and the poetic picture of the willows by the lake reflected in the water; but another time it may focus on describing the lively scene of children playing in the children's playground in the park in spring. Although both are about the scenery of the park in spring, the specific sections they focus on are obviously different.
[0065] Furthermore, in the embodiment of the present invention, the scoring model is used to score the N result texts respectively to obtain the N result text scores, including:
[0066] Select the open source language model as the scoring model, and use a 100-point scoring system as the scoring system;
[0067] The prompt word and the corresponding N result texts are input into the scoring model, and the result text scores corresponding to the N result texts are generated according to the scoring system.
[0068] Specifically, the scoring model can use open source models such as llama and qwen. Specifically, the question and the corresponding different results are input into the scoring model so that it can generate the corresponding score. The scoring system adopts a 100-point system, and the evaluation dimensions cover multiple aspects such as scientific accuracy, completeness, logic, and information richness. For example, in terms of accuracy, check whether the scientific knowledge and data references in the results are accurate; in terms of completeness, examine whether the results fully cover the key points involved in the problem; in terms of logic, pay attention to whether the reasoning process of the results is reasonable and coherent; in terms of information richness, mainly evaluate the breadth and depth of the information contained in the results. If you want to improve the adaptability of the scoring model to your specific usage scenarios, you can implement fine-tuning operations based on the open source model.
[0069] Therefore, the method in the embodiment of the present invention further includes fine-tuning the large language model based on the open source model, and the process includes:
[0070] Collecting business question-answer pairs and historical question-answer pairs of a large language model from a third-party knowledge base, and constructing a first data set based on the business question-answer pairs and the historical question-answer pairs; the first data set includes question texts and original answer texts;
[0071] Input the question text in the first data set into the open source language model to generate multiple answer texts;
[0072] The answer text is combined with the original answer text, and the answer text and the original answer text are separated by a separator, and the answer text and the original answer text are combined with the question text to generate a second data set;
[0073] Manually scoring the answer texts in the second data set according to a preset scoring standard, and constructing a third data set based on the manual scoring results and the second data set;
[0074] Taking the question text and answer text in the third data set as input and the manual scoring results as output, the large language model is fine-tuned, and the model parameters of the large language model are optimized using the cross entropy loss function to minimize the scoring error rate of the large language model and obtain a fine-tuned scoring model.
[0075] Specifically, taking the question-answering scenario as an example, the fine-tuning process of the large language model includes:
[0076] (1) Collect data. On the one hand, you can collect existing question-answer pairs, which can come from authoritative knowledge bases in related fields, business question-answer records accumulated in the past, etc. On the other hand, you can generate question-answer pairs based on your own data, and construct a first data set that meets the requirements through in-depth mining and analysis of the data. Next, input the question into the open source model to generate multiple answers, and then combine these generated answers with the original answers to form the second data set. Special symbols can be used to separate different answers, such as "[SEP]". Finally, arrange for professionals to score the generated answers according to the established scoring criteria, and integrate the manual scoring results with the second data set to form the third data set.
[0077] (2) Fine-tuning. Fine-tune the model using the constructed third dataset. In this process, the input is the question-answer pair and the output is the score. The cross-entropy loss function is used to optimize the model parameters, with the core goal of minimizing the scoring error rate as much as possible. During the fine-tuning process, the model will continuously adjust its internal parameter settings based on the input question-answer pairs and the corresponding manual scores, so that the model's scoring results for the question-answer pairs are more accurate and reliable, so as to better adapt to specific usage scenarios and needs, improve the evaluation efficiency and accuracy of the model in this scenario, and provide a better quality and accurate scoring basis for subsequent result screening.
[0078] In the embodiment of the present invention, after the generated result text is scored and the result text with the highest score is selected, the entropy of the result text with the highest score is calculated. Specifically, the entropy of each token in the result text needs to be calculated separately, and then the entropies of multiple tokens are multiplied to finally obtain the entropy value of the entire result text.
[0079] In information theory, entropy can accurately quantify the uncertainty and randomness of probability distribution. For language models, entropy is mainly used to measure the degree of uncertainty when predicting the probability distribution of the next word. Generally speaking, if the entropy value is low, it means that the language model has a high degree of certainty when predicting the next word, that is, the model can know more clearly what kind of word will appear next with a high probability; on the contrary, if the entropy value is high, it means that the model has a greater uncertainty in predicting the next word, and the range of possible words is wider and the probability distribution is more dispersed.
[0080] Further, in the embodiment of the present invention, the result text with the highest score is selected from the N result text scores, and the entropy value of the result text with the highest score is calculated and generated, including:
[0081] Sort the scores of N result texts and select the result text with the highest score;
[0082] Assume that the probability of any word in the highest-scoring result text appearing in the large language model is , calculate the entropy value of each word token in the result text with the highest score:
[0083] ;
[0084] in, is the word token; The number of all possible values of the word token; The final output is The probability of a possible value; log is the log function; the summation symbol Indicates that all possible values of the token are accumulated and calculated.
[0085] Multiply the entropy values of all word tokens to get the entropy value of the result text with the highest score E :
[0086] ;
[0087] in, is the number of word tokens in the result text with the highest score; Indicates The entropy of a word token.
[0088] Furthermore, in the embodiment of the present invention, the hyperparameters output by the large language model are adjusted based on the entropy value of the result text with the highest score to control the uncertainty of the output result of the large language model, including:
[0089] Determine whether the entropy value of the result text with the highest score exceeds a preset entropy value, and if the entropy value of the result text does not exceed the preset entropy value, keep the temperature and sampling strategy top_k output by the large language model unchanged;
[0090] If the entropy value of the result text exceeds the preset entropy value, the temperature and sampling strategy top_k of the large language model output are increased.
[0091] Specifically, the hyperparameters of the model output, including temperature, top-k, etc., are adjusted according to entropy, so as to adjust the output of the large model without changing the results and parameters of the large model, encourage the large model to perform deeper reasoning and reduce hallucinations. In short, when the model is certain about the output, keep the original hyperparameters unchanged; when the model is uncertain about the output, try to explore more options to generate more creative results.
[0092] The low entropy of the result text indicates that the model is very sure about the output, so this output should be kept because it is the best result, so the values of temperature and top-k remain unchanged.
[0093] A high entropy of the result text indicates that the model is uncertain about the output result, and some tokens require a wider range of values to obtain. In this case, the randomness of the model output should be increased, that is, the temperature and top_k should be increased.
[0094] Specifically, the preset entropy value is used to judge the high and low thresholds of entropy, which can be set according to the situation, such as , The number of tokens in the result text.
[0095] Specifically, temperature is a hyperparameter used to adjust the creativity and diversity of the model when generating text. It is a value greater than 0, usually between 0 and 1, and can also be greater than 1. Temperature affects the probability distribution of sampled predicted vocabulary when the model generates text. When the temperature of the model is high (such as 0.8, 1 or higher), the model will be more inclined to choose from a more diverse and different vocabulary, which makes the generated text riskier and more creative, but may also produce more errors and incoherence. When the temperature is lower (such as 0.2, 0.3, etc.), the model will mainly choose from vocabulary with higher probability, resulting in smoother and more coherent text. But at this time, the generated text may appear too conservative and repetitive. Therefore, in practical applications, it is necessary to weigh and choose the appropriate temperature value according to specific needs.
[0096] The formula for calculating the temperature increase is:
[0097] ;
[0098] is the original temperature, is the adjusted temperature, To adjust the coefficient, it can be adjusted according to the actual situation. The larger it is, the faster the temperature changes.
[0099] The top_k is a sampling strategy that samples from the top k tokens for each token, allowing other tokens with higher scores or probabilities to have a chance to be selected. In many cases, the randomness brought by this sampling helps improve the quality of generation while reducing the computational workload without sacrificing too much model response diversity.
[0100] The calculation formula for increasing top_k is:
[0101] ;
[0102] is the original top_k, is the adjusted top_k, To adjust the coefficient, it can be adjusted according to the actual situation. The larger it is, the faster the temperature changes.
[0103] In the embodiment of the present invention, the technical advantages of the model output optimization control method based on entropy and scoring include:
[0104] The model structure itself does not need to be changed, and there is no need to adjust the model parameters through fine-tuning, which reduces the dependence on professional knowledge and computing resources, allowing more people in non-professional fields who have relevant needs to conveniently use this method to optimize the model output;
[0105] To achieve optimal control of model output, without changing the structure and parameters of the model itself, the results generated by the model can be deeply analyzed and controlled through a unique entropy calculation and scoring mechanism, thereby improving the model effect.
[0106] Example 2: Reference Figure 2 As shown, the present invention provides a model output optimization control device based on entropy and scoring, comprising:
[0107] A text generation module, used to select corresponding prompt words according to the application scenario, input the prompt words into the large language model, and control the large language model to generate N result texts through preset model parameters;
[0108] A text scoring module is used to score the N result texts respectively using a scoring model to obtain N result text scores;
[0109] An entropy value calculation module is used to select the result text with the highest score from the N result text scores, and calculate the entropy value of the result text with the highest score;
[0110] The output optimization module is used to adjust the hyperparameters of the large language model output based on the entropy value of the result text with the highest score to control the uncertainty of the output result of the large language model; the hyperparameters of the large language model output include temperature and sampling strategy top_k.
[0111] In the embodiment of the present invention, by scoring the model output and calculating the entropy value, the hyperparameters of the model are dynamically adjusted, and the model output is optimized and controlled without changing the structure and parameters of the model itself, thereby improving the effect of the model.
[0112] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A model output optimization control method based on entropy and scoring, characterized in that: include: Select corresponding prompt words according to the application scenario, input the prompt words into the large language model, and control the large language model to generate N result texts through preset model parameters; Using a scoring model to score the N result texts respectively, to obtain N result text scores; the scoring model is an open source language model; Select the result text with the highest score from the N result text scores, and calculate the entropy value of the result text with the highest score; Based on the entropy value of the result text with the highest score, the hyperparameters output by the large language model are adjusted to control the uncertainty of the output result of the large language model; the hyperparameters output by the large language model include temperature and sampling strategy top_k; The hyperparameters output by the large language model are adjusted based on the entropy value of the result text with the highest score to control the uncertainty of the output result of the large language model, including: Determine whether the entropy value of the result text with the highest score exceeds a preset entropy value, and if the entropy value of the result text does not exceed the preset entropy value, keep the temperature and sampling strategy top_k output by the large language model unchanged; If the entropy value of the result text exceeds the preset entropy value, the temperature and sampling strategy top_k of the large language model output are increased.
2. The model output optimization control method based on entropy and scoring according to claim 1 is characterized in that: The method of selecting a corresponding prompt word according to the application scenario, inputting the prompt word into a large language model, and controlling the large language model to generate N result texts through preset model parameters includes: According to the application scenario of the large language model, select a prompt word that matches the application scenario from the preset prompt word group; The model parameters of the large language model are preconfigured, and the selected prompt word is input into the large language model. The large language model is controlled by the model parameters to generate N different result texts, and the N different result texts are stored in the database according to the preset text format.
3. The model output optimization control method based on entropy and scoring according to claim 1 is characterized in that: The scoring model is used to score the N result texts respectively to obtain the N result text scores, including: Select the open source language model as the scoring model, and use a 100-point scoring system as the scoring system; The prompt word and the corresponding N result texts are input into the scoring model, and the result text scores corresponding to the N result texts are generated according to the scoring system.
4. The model output optimization control method based on entropy and scoring according to claim 1 is characterized in that: The scoring model can be fine-tuned based on the open source model, specifically: Collecting business question-answer pairs and historical question-answer pairs of a large language model from a third-party knowledge base, and constructing a first data set based on the business question-answer pairs and the historical question-answer pairs; the first data set includes question texts and original answer texts; Input the question text in the first data set into the open source language model to generate multiple answer texts; The answer text is combined with the original answer text, and the answer text and the original answer text are separated by a separator, and the answer text and the original answer text are combined with the question text to generate a second data set; Manually scoring the answer texts in the second data set according to a preset scoring standard, and constructing a third data set based on the manual scoring results and the second data set; Taking the question text and answer text in the third data set as input and the manual scoring results as output, the large language model is fine-tuned, and the model parameters of the large language model are optimized using the cross entropy loss function to minimize the scoring error rate of the large language model and obtain a fine-tuned scoring model.
5. The model output optimization control method based on entropy and scoring according to claim 1 is characterized in that: The step of selecting the result text with the highest score from the N result text scores and calculating the entropy value of the result text with the highest score comprises: Sort the scores of N result texts and select the result text with the highest score; Assume that the probability of any word in the highest-scoring result text appearing in the large language model is , calculate the entropy value of each word token in the result text with the highest score: ; in, is the word token; The number of all possible values of the word token; The final output is The probability of a possible value; is the log function; Multiply the entropy values of all word tokens to get the entropy value of the result text with the highest score E : ; in, is the number of word tokens in the result text with the highest score; Indicates The entropy of a word token.
6. A model output optimization control device based on entropy and scoring, characterized in that: include: A text generation module, used to select corresponding prompt words according to the application scenario, input the prompt words into the large language model, and control the large language model to generate N result texts through preset model parameters; A text scoring module, used to score N result texts respectively using a scoring model to obtain N result text scores; the scoring model is an open source language model; An entropy value calculation module is used to select the result text with the highest score from the N result text scores, and calculate the entropy value of the result text with the highest score; An output optimization module, used to adjust the hyperparameters of the large language model output based on the entropy value of the result text with the highest score, so as to control the uncertainty of the output result of the large language model; the hyperparameters of the large language model output include temperature and sampling strategy top_k; The hyperparameters output by the large language model are adjusted based on the entropy value of the result text with the highest score to control the uncertainty of the output result of the large language model, including: Determine whether the entropy value of the result text with the highest score exceeds a preset entropy value, and if the entropy value of the result text does not exceed the preset entropy value, keep the temperature and sampling strategy top_k output by the large language model unchanged; If the entropy value of the result text exceeds the preset entropy value, the temperature and sampling strategy top_k of the large language model output are increased.
Citation Information
Patent Citations
Open source large language model fine tuning optimization method based on transfer learning
CN118095441A
Text information extraction method based on large language model and efficient parameter fine tuning
CN118132674A