Large language model updating method and question answering method
Through self-game technology combining evaluation model and optimization model, the model preset prompt items of the large language model are updated, which solves the dependence on manual annotation and computing resources in traditional methods, and improves the update efficiency and flexibility of the large language model.
Patent Information
- Application Number
- CN202510251302.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-06-20
AI Technical Summary
The traditional large language model optimization method relies on a large amount of manual annotation and expensive computing resources, resulting in limited scalability and flexibility of the model, and the training resources are consumed greatly, and the efficiency of updating large language models is low.
Through self-game technology combining evaluation models and optimization models, the large language model is used to process the problem data according to the model preset prompt items, and the answer data and evaluation data are obtained. The model preset prompt items are updated based on these data until the target large language model that meets the model update conditions are obtained.
Reliance on manual annotation is reduced, the consumption of training resources is reduced, the update efficiency and update flexibility of large language models are improved, and the resource-intensive and inefficient problems in traditional methods are avoided.
Smart Images

Figure CN120181231A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this specification relate to the field of computer technology, and particularly to a method for updating large language models and a question and answer method. Background Art
[0002] With the rapid progress and continuous innovation of artificial intelligence technology, large language models have achieved remarkable achievements. They not only demonstrate excellent instruction-following capabilities, being able to accurately understand and execute various complex instructions, but also possess powerful problem-solving capabilities, being able to quickly give reasonable solutions when faced with diverse problems. Large language models have also opened a new chapter in open-ended conversations, being able to communicate smoothly and naturally with users, providing users with rich and valuable information and interactive experiences.
[0003] However, traditional large model optimization methods highly rely on a large amount of manual annotation and expensive computing resources, resulting in limited scalability and flexibility of the models. To solve this problem, self-play technology has emerged. It significantly reduces the dependence on manual annotation by self-labeling data and performing iterative training. Nevertheless, self-play technology still faces the problems of huge consumption of training resources and low update efficiency of large language models. Therefore, there is an urgent need for a more effective method for updating large language models to solve the above problems. Summary of the Invention
[0004] In view of this, the embodiments of this specification provide a method for updating large language models. One or more embodiments of this specification also relate to a device for updating large language models, a question and answer method, a question and answer device, another question and answer method, another question and answer device, a computing device, a computer-readable storage medium, and a computer program product to solve the technical defects existing in the prior art.
[0005] According to the first aspect of the embodiments of this specification, a method for updating large language models is provided, including:
[0006] Using a large language model to process question data according to the model preset prompt items to obtain answer data, where the model preset prompt items are used to guide the prediction behavior of the large language model;
[0007] Inputting the question data and the answer data into an evaluation model for processing to obtain the answer evaluation data corresponding to the large language model;
[0008] Inputting the question data, the answer data, and the answer evaluation data into an optimization model for processing to obtain the prompt update guiding information of the model preset prompt items;
[0009] Updating the model preset prompt items of the large language model based on the above-mentioned prompt update guidance information until a target large language model that meets the model update conditions is obtained.
[0010] According to a second aspect of the embodiments of the present specification, there is provided a large language model updating device, including:
[0011] A large language model processing module, configured to process question data using a large language model according to model preset prompt items to obtain answer data, where the model preset prompt items are used to guide the prediction behavior of the large language model;
[0012] An evaluation model processing module, configured to input the question data and the answer data into an evaluation model for processing to obtain answer evaluation data corresponding to the large language model;
[0013] An optimization model processing module, configured to input the question data, the answer data, and the answer evaluation data into an optimization model for processing to obtain prompt update guidance information for the model preset prompt items;
[0014] An update module, configured to update the model preset prompt items of the large language model based on the prompt update guidance information until a target large language model that meets the model update conditions is obtained.
[0015] According to a third aspect of the embodiments of the present specification, there is provided a question and answer method, including:
[0016] Determining target question data;
[0017] Inputting the target question data into a target large language model for processing to obtain target answer data corresponding to the target question data.
[0018] According to a fourth aspect of the embodiments of the present specification, there is provided a question and answer device, including:
[0019] A determination module, configured to determine target question data;
[0020] A processing module, configured to input the target question data into a target large language model for processing to obtain target answer data corresponding to the target question data.
[0021] According to a fifth aspect of the embodiments of the present specification, there is provided another question and answer method, including:
[0022] Receiving question data to be answered submitted by a user;
[0023] The large language model is utilized to process the problem data to be answered according to the preset prompt items of the model, obtain the target answer data, and send the target answer data to the user. The preset prompt items of the model are used to guide the prediction behavior of the large language model;
[0024] The problem data to be answered and the target answer data are input into an evaluation model for processing to obtain the answer evaluation data corresponding to the large language model;
[0025] The problem data to be answered, the target answer data, and the answer evaluation data are input into an optimization model for processing to obtain the prompt update guiding information of the preset prompt items of the model;
[0026] Based on the prompt update guiding information, the preset prompt items of the large language model are updated until a target large language model that meets the model update conditions is obtained.
[0027] According to the sixth aspect of the embodiments of this specification, another question and answer device is provided, including:
[0028] A receiving module, configured to receive the problem data to be answered submitted by the user;
[0029] A sending module, configured to utilize the large language model to process the problem data to be answered according to the preset prompt items of the model, obtain the target answer data, and send the target answer data to the user. The preset prompt items of the model are used to guide the prediction behavior of the large language model;
[0030] An input module, configured to input the problem data to be answered and the target answer data into an evaluation model for processing to obtain the answer evaluation data corresponding to the large language model;
[0031] A processing module, configured to input the problem data to be answered, the target answer data, and the answer evaluation data into an optimization model for processing to obtain the prompt update guiding information of the preset prompt items of the model;
[0032] An update module, configured to update the preset prompt items of the large language model based on the prompt update guiding information until a target large language model that meets the model update conditions is obtained.
[0033] According to the seventh aspect of the embodiments of this specification, a computing device is provided, including:
[0034] A memory and a processor;
[0035] The memory is used to store computer programs or instructions, and the processor is used to execute the computer programs or instructions. When the computer programs or instructions are executed by the processor, the steps of the above method are implemented.
[0036] According to the eighth aspect of the embodiments of this specification, a computer-readable storage medium is provided, which stores a computer program or instructions. When the computer program or instructions are executed by a processor, the steps of the above method are implemented.
[0037] According to the ninth aspect of the embodiments of this specification, a computer program product is provided, including a computer program or instructions. When the computer program or instructions are executed by a processor, the steps of the above method are implemented.
[0038] The large language model update method provided by an embodiment of this specification uses a large language model to process problem data according to model preset prompt items to obtain answer data. The model preset prompt items are used to guide the prediction behavior of the large language model. According to the model preset prompt items, system instructions can be provided when the large language model processes problem data, providing the large language model with thinking methods, knowledge bases, and problem processing capabilities that can be used. The problem data and answer data are input into an evaluation model for processing to obtain answer evaluation data corresponding to the large language model, realizing the evaluation and criticism of the answer data and the answer process of the large language model. The problem data, answer data, and answer evaluation data are input into an optimization model for processing to obtain prompt update guiding information for the model preset prompt items. Based on the prompt update guiding information, the model preset prompt items of the large language model are updated until a target large language model that meets the model update conditions is obtained. The prompt update guiding information can be used to optimize the model preset prompt items of the large language model to achieve the purpose of updating the large language model. By combining the evaluation model and the optimization model, the large language model optimizes the model preset prompt items in a self-playing manner, and the update of the large language model can be completed without using the model update method of model training, which can reduce the consumption of training resources and improve the update efficiency and update flexibility of the large language model. Description of the Drawings
[0039] Figure 1 is a schematic diagram of a large language model update method provided by an embodiment of this specification;
[0040] Figure 2 is a flowchart of a large language model update method provided by an embodiment of this specification;
[0041] Figure 3 is a processing process flowchart of a large language model update method provided by an embodiment of this specification;
[0042] Figure 4 is a schematic structural diagram of a large language model update device provided by an embodiment of this specification;
[0043] Figure 5It is a flowchart of a question-and-answer method provided by an embodiment of this specification;
[0044] Figure 6 It is a schematic structural diagram of a question-and-answer device provided by an embodiment of this specification;
[0045] Figure 7 It is a flowchart of another question-and-answer method provided by an embodiment of this specification;
[0046] Figure 8 It is a schematic structural diagram of another question-and-answer device provided by an embodiment of this specification;
[0047] Figure 9 It is a structural block diagram of a computing device provided by an embodiment of this specification. Detailed implementation manners
[0048] In the following description, numerous specific details are set forth in order to provide a thorough understanding of this specification. However, this specification can be implemented in many other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the connotation of this specification. Therefore, this specification is not limited by the specific implementations disclosed below.
[0049] The terms used in one or more embodiments of this specification are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of this specification. The singular forms "a", "the", and "said" used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more of the associated listed items.
[0050] It should be understood that although the terms first, second, etc. may be used in one or more embodiments of this specification to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein can be interpreted as "when" or "while" or "in response to determining".
[0051] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data that have been authorized by the user or fully authorized by all parties. Moreover, the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or reject.
[0052] In one or more embodiments of this specification, a large model refers to a deep learning model with a large number of model parameters, usually including hundreds of millions, tens of billions, hundreds of billions, trillions, or even more than one quadrillion model parameters. A large model can also be referred to as a Foundation Model. Through pre-training of the large model with a large amount of unlabeled corpus, a pre-trained model with more than one billion parameters is produced. This model can adapt to a wide range of downstream tasks and has good generalization ability. For example, Large Language Model (LLM), multi-modal pre-training model, etc.
[0053] When a large model is actually applied, it only needs to fine-tune the pre-trained model with a small number of samples to be applied to different tasks. Large models can be widely applied in fields such as Natural Language Processing (NLP), computer vision, etc. Specifically, they can be applied to tasks in the field of computer vision such as Visual Question Answering (VQA), Image Caption (IC), image generation, etc., and tasks in the field of natural language processing such as text-based sentiment classification, text summary generation, machine translation, etc. The main application scenarios of large models include digital assistants, intelligent robots, search, online education, office software, e-commerce, intelligent design, etc.
[0054] First, the noun terms involved in one or more embodiments of this specification are explained.
[0055] Monte Carlo sampling: A random simulation method that approximates the solution of problems such as expectation, mean, area, integral, etc. by sampling a large number of particles from a specific distribution.
[0056] Decoding-based method: Refers to Decoding-based Learning, which is a machine learning method. It maps the input data to a set of latent feature spaces and then decodes it into output data to achieve a better representation of the original data.
[0057] Self-Play: A training method in the field of reinforcement learning. During the process of an agent playing against different versions or copies of itself, it realizes self-evolution and optimization by continuously learning and adjusting its strategy. When the agent plays against itself, it can simulate different opponents and scenarios, thereby continuously adjusting and optimizing its strategy. This self-confrontational approach helps the agent learn more robust and efficient strategies in complex environments.
[0058] Prompt: In the field of natural language processing (NLP), it generally refers to an input text or instruction used to guide the model to generate a specific type of output. It can be a question, request, or command input by the user, aiming to obtain a specific response from the model or complete a specific task. Prompts play a crucial role in NLP tasks because they determine how the model will understand and respond to the input.
[0059] Prompt Engineering: An engineering method based on natural language processing technology. By formulating a series of principles and iterative processes, it transforms text inputs into prompts with specific semantics to guide machine learning models to produce the required outputs. Prompt Engineering utilizes the model's sensitivity to inputs. By designing appropriate prompts, it influences the model's internal state, thereby guiding the model to generate outputs that meet specific requirements.
[0060] System Prompt: In conversational AI applications, it refers to a set of initial instructions or background information provided to the AI to guide its behavior and response patterns. It helps set the AI's role, tone, knowledge scope, etc., ensuring that the AI can interact with users as expected. System prompts are usually set by developers or administrators and are not directly shown to end users. It can be very specific, such as specifying the rules that the AI should follow, or relatively broad, such as defining the AI's personality traits.
[0061] Self-Improvement: It refers to the process by which a model continuously improves its own performance through self-generating data, self-evaluating, and self-training, etc., without external supervision or with only a small amount of external supervision.
[0062] SFT (Scaling Fine-Tuning): A fine-tuning technique for large language models (LLMs). In natural language processing (NLP), fine-tuning is a commonly used technique that allows a model to be further trained on a specific task or dataset to improve its performance. The SFT technique achieves customized optimization of the model by first pre-training on large-scale data and then fine-tuning on a dataset for a specific task. This method can significantly improve the model's performance on specific tasks and make it more adaptable to actual application scenarios.
[0063] RLHF (Reinforcement Learning with Human Feedback): An AI training method that combines reinforcement learning and human feedback. In this method, the AI model first learns by self-generating and evaluating answers, and then humans rank or score these answers to provide feedback. The model further learns and adjusts based on this feedback to generate answers that better meet human expectations. The core of the RLHF method lies in using human wisdom and judgment to guide the learning process of the AI model, making the answers it generates more accurate, natural, and useful. This method is of great significance in enhancing the dialogue ability and user experience of AI models.
[0064] In the field of artificial intelligence, large language models have demonstrated excellent capabilities in following instructions, solving problems, and engaging in open conversations. However, the optimization of traditional large models highly relies on manual annotation and high computing power, which limits their scalability and flexibility. To address this issue, self-play techniques significantly reduce the dependence on manual annotation by automatically annotating data and iteratively training, but still face the challenge of high training resource consumption. In recent years, researchers have explored test-time enhancement algorithms, such as Monte Carlo sampling and decoding-based methods, aiming to find optimal solutions or achieve model alignment during the test phase, further reducing the dependence on training. The capabilities of large language models are multi-faceted, such as direct problem-solving ability, reflection ability, critical ability, etc. Empirically, the effects of various self-play techniques in the industry have confirmed that self-improvement of model performance can be further achieved by mining and integrating different aspects of the model's capabilities.
[0065] The potential of large language models needs to be induced and enhanced through prompts, that is, prompt engineering. System prompts are an important type in prompt engineering. An embodiment of this specification provides a method for updating a large language model that focuses on iteratively optimizing system prompts through self-play, leveraging the direct problem-solving ability, reflection ability, critical ability, etc. of the large language model to achieve a self-improving model effect.
[0066] Figure 1 FIG. shows a schematic diagram of a method for updating a large language model according to an embodiment of this specification. The model preset prompt item is the system prompt of the large language model. The large language model processes the problem data according to the model preset prompt item to obtain the response data. According to the model preset prompt item, it can provide system instructions when the large language model processes the problem data, providing the large language model with a thinking mode, knowledge base, and problem-solving ability that can be used. The problem data and the response data are input into an evaluation model for processing to obtain the response evaluation data corresponding to the large language model, realizing the evaluation and criticism of the response data and the response process of the large language model.
[0067] The problem data, the response data, and the response evaluation data are input into an optimization model for processing to obtain the prompt update guidance information for the model preset prompt item. Based on the prompt update guidance information, the model preset prompt item of the large language model is updated until a target large language model that meets the model update conditions is obtained. The prompt update guidance information can be used to optimize the model preset prompt item of the large language model to achieve the purpose of updating the large language model. By combining the evaluation model and the optimization model, the large language model optimizes the model preset prompt item in a self-play manner, and can complete the update of the large language model without using the model update method of model training, which can reduce the consumption of training resources and improve the update efficiency and flexibility of the large language model. Through self-play, the system prompt is iteratively optimized by leveraging the direct problem-solving ability, reflection ability, critical ability, etc. of the large language model to achieve a self-improving model effect.
[0068] In this specification, a method for updating a large language model is provided. This specification also relates to a device for updating a large language model, a question-answering method, a question-answering device, another question-answering method, another question-answering device, a computing device, and a computer-readable storage medium, which will be described in detail one by one in the following embodiments.
[0069] See Figure 2 ,Figure 2 The flowchart of a large language model update method provided according to an embodiment of this specification is shown, specifically including the following steps.
[0070] Step 202: Use the large language model to process the question data according to the model preset prompt items to obtain the answer data, and the model preset prompt items are used to guide the prediction behavior of the large language model.
[0071] Specifically, the large language model is a large model with multi-faceted capabilities. A large model refers to a deep learning model with a large number of model parameters. The large language model usually has various capabilities such as direct problem-solving ability, reflection ability, and critical ability. The large language model can handle problems in various fields such as text generation, language translation, dialogue systems, information retrieval and summarization, and natural language understanding. The model preset prompt items are system instructions for the large language model, used to guide the prediction behavior or data processing behavior of the large language model. The model preset prompt items can ensure that the large language model can perform problem reasoning and answer in the expected way by presetting prompt items such as roles, capabilities, thinking methods, knowledge bases, text styles, tones, and knowledge scopes. The model preset prompt items can be determined by manual writing according to requirements.
[0072] Among them, the identity represents the role that the large language model needs to play, such as helpful assitant, outstanding physicist, etc. The ability represents the fields and capabilities that the large language model needs to be good at. If the question data is a math problem, the large language model needs to have excellent math reasoning ability and corresponding subject knowledge ability; if it is a kitchen knowledge question, the large language model needs to have sophisticated cooking skills, etc.; the thinking method represents the thinking methods that the large language model can use in the process of answering questions, mainly including step-by-step cot (chain of thought), reflection, backtracking, self-correction, etc.; the knowledge base can autonomously supplement the knowledge points required by the question data; the question instruction constraint means abstracting the instruction constraints involved in the question data into the model preset prompt items of the large language model to enhance the large language model's compliance ability with this instruction. The question data is the data that can be input into the large language model for processing. The answer data is the answer obtained by the large language model combining the model preset prompt items of the large language model to reason and answer the question data.
[0073] Based on this, after determining the question data, the question data can be input into the large language model for question understanding, analysis, reasoning, and answering. Use the large language model and combine the model preset prompt items preset by the large language model to process the question data to obtain the answer data corresponding to the question data.
[0074] In practical applications, the training of large language models includes: selecting sample question data from a training sample set and inputting it into an initial large language model to obtain predicted answer data; selecting sample answer data corresponding to the sample question data in the training sample set, and calculating the large language model loss value based on the large language model loss function, the sample answer data, and the predicted answer data; adjusting the parameters of the initial large language model based on the large language model loss value until the large language model that meets the large language model training stop condition is obtained.
[0075] Specifically, when implemented, the large language model can be a machine learning model that has completed model pre-training and has certain problem understanding and reasoning abilities. The large language model usually has the ability to directly solve problems and is used to directly answer questions raised by users. The large language model can be pre-trained using supervised training. Before pre-training the large language model, determine the training sample set and the large language model loss function. The large language model loss function can select various loss functions such as cross-entropy loss function and logarithmic loss function. When pre-training the large language model, the large language model training stop condition can be reaching a preset number of training rounds, the prediction accuracy of the model reaching an accuracy threshold, reaching a preset training time length, etc.
[0076] Furthermore, considering that multiple sub-prompt items are preset in the model preset prompt items, and each sub-prompt item corresponds to a different prompt setting, in order to be able to understand and analyze the question data more accurately, a prompt sub-item that matches the question attribute of the question data can be selected to process the question data. The specific implementation is as follows:
[0077] Input the question data into the large language model, and determine the prompt sub-item in the model preset prompt items according to the question attribute of the question data; use the large language model to process the question data according to the prompt sub-item to obtain the answer data.
[0078] Specifically, the question attribute of the question data can include information such as question type, question domain, and question constraints. The model preset prompt items contain at least one sub-prompt item, and the sub-prompt item can be items such as identity, ability, role, thinking mode, knowledge base, text style, tone, knowledge scope, and question instruction constraints. The prompt sub-item is at least one sub-prompt item that matches the question attribute and can be used to process the question data.
[0079] Based on this, the problem data is input into the large language model, and at least one hint sub-item that matches the problem attribute is selected from the sub-hint items included in the model preset hint items according to the problem attribute of the problem data. The large language model is used to process the problem data according to at least one hint sub-item to obtain the answer data. When using the large language model to process the problem data, the problem data can be processed in combination with hint items such as the identity (physicist, assistant, mathematician), ability (problem reasoning ability), knowledge base, etc. corresponding to at least one hint sub-item to obtain accurate answer data.
[0080] For example, in the case where the problem data belongs to a problem in the information retrieval field, the problem data can be input into the large language model for problem understanding, reasoning, and answer prediction. The problem data can be "Please provide information on the development of artificial intelligence and generate a 200-word summary". Therefore, it is determined that the problem attributes of the problem data are: information retrieval field, text processing type, and the problem constraint is a 200-word summary. Inputting the problem data into the large language model, hint sub-items that match the problem attributes can be determined from the model preset hint items of the large language model: identity, ability, role, thinking mode, knowledge base, text style, tone, knowledge scope, question instruction constraints, etc.: identity, ability, and question instruction constraints. Using the large language model to understand, analyze, reason, and answer the problem data according to the selected hint sub-items, and then obtain the answer data "Artificial intelligence (AI), as an important branch of computer science, has experienced a leapfrog development from theoretical exploration to technological implementation since the mid-20th century. Its core goal is to simulate human intelligence and achieve autonomous learning, reasoning, and decision-making through algorithms and big data. In the early 21st century, the breakthrough of machine learning (especially deep learning) pushed AI into a new stage. In recent years, AI technology has accelerated its penetration into various fields: assisting disease diagnosis and drug research and development in the medical field; optimizing risk assessment and quantitative trading in the financial field; improving traffic efficiency in autonomous driving; and supporting personalized learning in the education field."
[0081] In summary, when using the large language model to process the problem data, the problem data is processed in combination with at least one hint sub-item to obtain accurate answer data and improve the practicality of the answer data.
[0082] Step 204: Input the problem data and the answer data into an evaluation model for processing to obtain the answer evaluation data corresponding to the large language model.
[0083] Specifically, after processing the question data according to the model preset prompt items using the large language model and obtaining the answer data, the question data and the answer data can be input into the evaluation model for processing to obtain the answer evaluation data corresponding to the large language model. Among them, the evaluation model can be a derivative model of the large language model. The evaluation model has a critical ability and is used to comment on the answer data output by the large language model and the answering process of the large language model for answering the question data, and output the shortcomings and suggestions of the large language model for this question answer. The evaluation model can evaluate the answer data output by the large language model in multiple evaluation dimensions, evaluate the advantages and disadvantages of the answer data, and give answer suggestions. The answer evaluation data can include the optimization suggestions of the large language model, the shortcomings, advantages and optimization suggestions of the answer data.
[0084] Based on this, after processing the question data according to the model preset prompt items using the large language model and obtaining the answer data, the question data and the answer data are input into the evaluation model for processing. The evaluation model is used to evaluate the answering process and the answer data of the large language model, give the advantages, disadvantages and optimization suggestions of the answer data, and obtain the answer evaluation data corresponding to the large language model. The evaluation model can objectively evaluate the answer data output by the large language model to promote the further improvement and optimization of the large language model. The evaluation model can evaluate the answer data in dimensions such as security, practicality, logic and language expression dimension.
[0085] In practical applications, the training of the evaluation model includes: selecting the sample data to be evaluated from the evaluation model training sample set and inputting it into the initial evaluation model to obtain the answer evaluation data; selecting the sample answer evaluation data corresponding to the sample data to be evaluated from the evaluation model training sample set, and calculating the evaluation model loss value based on the evaluation model loss function, the sample data to be evaluated and the answer evaluation data; adjusting the parameters of the initial evaluation model based on the evaluation model loss value until the evaluation model that meets the evaluation model training stop condition is obtained.
[0086] In specific implementation, the evaluation model can be a derivative model of a large language model. After the initial large language model completes pre-training to obtain the large language model, data can be collected during the use of the large language model, collecting the input data and output data of the large language model, and generating evaluation model training samples for training the initial evaluation model. For the input data and output data of the large language model, output evaluation data is labeled. The initial evaluation model training samples are composed of the input data, output data, and output evaluation data, and the initial evaluation model is trained to obtain an evaluation model that can evaluate the large language model. The evaluation model has certain problem understanding, reasoning, and evaluation capabilities. The initial evaluation model can be trained using a supervised training method. Before training the initial evaluation model, the evaluation model training samples and the evaluation model loss function are determined. The evaluation model loss function can select various loss functions such as the cross-entropy loss function and the logarithmic loss function. When training the initial evaluation model, the evaluation model training stop condition can be reaching a preset number of training rounds, the prediction accuracy of the model reaching an accuracy threshold, reaching a preset training time length, etc.
[0087] Furthermore, considering the answer data output by the large language model, which represents the direct problem-solving ability of the large language model, when using the evaluation model to evaluate the answer data output by the large language model in response to problem data, the large language model can be evaluated in multiple evaluation dimensions, and the specific implementation is as follows:
[0088] The problem data and the answer data are input into the evaluation model, and the corresponding logical evaluation dimension, information security evaluation dimension, practicality evaluation dimension, and language expression evaluation dimension of the evaluation model are determined; the evaluation model processes the problem data and the answer data according to the logical evaluation dimension, the information security evaluation dimension, the practicality evaluation dimension, and / or the language expression evaluation dimension to obtain the answer evaluation data corresponding to the large language model.
[0089] Specifically, the logical evaluation dimension is used to evaluate the logical coherence and logical rigor of the answer data, as well as the accuracy of the answer; the information security evaluation dimension is used to evaluate whether the answer data involves information security network problems, that is, information confidentiality and integrity; the practicality evaluation dimension is used to evaluate whether the answer data has practicality, that is, whether the answer data is accurate and comprehensive, and whether the answer data can be effectively applied in a real environment; the language expression evaluation dimension is used to evaluate whether the language expression of the answer data is accurate, and whether the language style meets the user's needs, and whether there are any inappropriate words in the language expression.
[0090] Based on this, the question data and answer data are input into an evaluation model for processing, and the corresponding logical evaluation dimension, information security evaluation dimension, practicality evaluation dimension, and language expression evaluation dimension of the evaluation model are determined. When using the evaluation model to evaluate the large language model based on the answer data and question data, at least one evaluation dimension can be selected from the evaluation dimensions such as the logical evaluation dimension, information security evaluation dimension, practicality evaluation dimension, and language expression evaluation dimension, so as to evaluate the large language model based on the question data and answer data and obtain the corresponding answer evaluation data of the large language model.
[0091] Continuing with the above example, after using the large language model to process the question data and obtaining the answer data given by the large language model, the evaluation model can be used to evaluate the answer process and answer data of the large language model based on the question data and answer data. Input the question data into the evaluation model, and select at least one evaluation dimension from multiple evaluation dimensions such as the logical evaluation dimension, information security evaluation dimension, practicality evaluation dimension, and language expression evaluation dimension provided by the evaluation model to evaluate the answer data of the large language model for the question data. The logical evaluation dimension, practicality evaluation dimension, and language expression evaluation dimension can be selected to evaluate the answer data and obtain the answer evaluation data: Logical evaluation dimension: The answer in the answer data is logically clear and systematically elaborates from the definition, historical development to the current applications of artificial intelligence. The connection between each part is natural and conforms to the logical order. However, a transitional sentence can be considered to be added between "Machine learning drives AI into a new stage" and "AI technology accelerates its penetration into various fields" to make the logic smoother. Practicality evaluation dimension: The answer data lists the applications of AI technology in fields such as healthcare, finance, transportation, and education, and these applications have high practicality. However, to enhance the practicality, cases on how AI technology specifically solves industry pain points or improves efficiency can be further supplemented. Language expression evaluation dimension: The language expression of the answer is accurate and concise. However, to enhance the vividness and attractiveness of the expression, metaphors or specific cases can be appropriately used to enrich the content, and the writing style can be changed to a humorous and interesting type to attract readers to read.
[0092] To sum up, select at least one evaluation dimension from the evaluation dimensions such as the logical evaluation dimension, information security evaluation dimension, practicality evaluation dimension, and language expression evaluation dimension, objectively evaluate the answer data output by the large language model in combination with the question data, and give optimization suggestions. Thereby further improving the data processing ability of the large language model.
[0093] Step 206: Input the question data, the answer data, and the answer evaluation data into an optimization model for processing to obtain the prompt update guiding information of the preset prompt items of the model.
[0094] Specifically, after inputting the question data and the answer data into the evaluation model for processing to obtain the answer evaluation data corresponding to the large language model, the question data, the answer data, and the answer evaluation data can be input into the optimization model for processing to obtain the hint update guiding information for the model preset hint items. Among them, the optimization model can be a derivative model of the large language model, with the ability of planning and reflection, and is used to guide the update of the model preset hint items of the large language model. The hint update guiding information is used to guide the update of the model preset hint items, and the hint update guiding information can be an update suggestion for updating the model preset hint items. The hint update guiding information can be used to update (delete or modify) the sub-hint items in the model preset hint items, or to add new sub-hint items to the model preset hint items. That is, based on the hint update guiding information, an adjustment update for deleting or modifying the model preset hint items can be performed, or an addition update for adding new items to the model preset hint items can be performed.
[0095] Based on this, after inputting the question data and the answer data into the evaluation model for processing to obtain the answer evaluation data corresponding to the large language model, the question data, the answer data, and the answer evaluation data are input into the optimization model for processing to obtain the hint update guiding information for the model preset hint items. Subsequently, an adjustment update for directly modifying or deleting the model preset hint items can be performed based on the hint update guiding information, or an addition update for adding new sub-hint items to the model preset hint items can be performed. Thereby, the data processing ability of the large language model can be further improved.
[0096] In practical applications, the training of the optimization model includes: selecting the sample data to be optimized from the optimization model training sample set and inputting it into the initial optimization model to obtain the hint optimization data. The sample data to be optimized includes the question sample, the answer sample corresponding to the question sample, and the answer evaluation sample corresponding to the answer sample; selecting the sample hint optimization data corresponding to the sample data to be optimized from the optimization model training sample set, and calculating the optimization model loss value based on the optimization model loss function, the sample data to be optimized, and the hint optimization data; adjusting the parameters of the initial optimization model based on the optimization model loss value until the optimization model that meets the optimization model training stop condition is obtained.
[0097] In specific implementation, the optimization model can be a derivative model of a large language model. After the initial large language model completes pre-training to obtain the large language model, data can be collected during the use of the large language model, including the input data and output data of the large language model. Combining with the evaluation model to evaluate the input data and output data to generate output evaluation data, and generating optimization model training samples for training the initial optimization model. For the input data and output data of the large language model, as well as the output evaluation data generated by the evaluation model based on the input data and output data of the large language model, optimize the data. The initial optimization model training samples are composed of input data, output data, output evaluation data, and optimized data, and the initial optimization model is trained to obtain an optimization model that can optimize the large language model. The optimization model has certain planning and reflection capabilities. The initial optimization model can be trained using supervised training. Before training the initial optimization model, determine the optimization model training samples and the optimization model loss function. The optimization model loss function can select various loss functions such as cross-entropy loss function and logarithmic loss function. When training the optimization model, the optimization model training stop condition can be reaching the preset number of training rounds, the prediction accuracy of the model reaching the accuracy threshold, reaching the preset training time length, etc.
[0098] Furthermore, when the optimization model evaluates the answer data, it needs to refer to the question data and the answer evaluation data. The specific implementation is as follows:
[0099] Use the optimization model to compare the answer data and the answer evaluation data based on the question data, and determine the update guidance information of the sub-prompt items corresponding to the preset sub-prompt items in the model preset prompt items, as well as the update guidance information corresponding to the model preset prompt items; use the sub-prompt item update guidance information and the update guidance information as the prompt update guidance information.
[0100] Specifically, the preset sub-prompt item is the model prediction prompt item preset in the model preset prompt item, which is used to provide system instructions when the large language model predicts the question data. The sub-prompt item update guidance information is used to guide the update of the preset sub-prompt item, and can delete and modify the preset sub-prompt item; the update guidance information is used to guide the update of the model preset prompt item, and can add new sub-prompt items based on the existing preset sub-prompt items in the model preset prompt item.
[0101] Based on this, the optimization model is used to compare the answer data and answer evaluation data based on the question data, determine the preset sub - prompt items that need to be updated in the model preset prompt items according to the comparison results, and determine the prompt sub - item update guidance information corresponding to the preset sub - prompt items for updating the preset sub - prompt items. It is also possible to determine the update guidance information corresponding to the model preset prompt items according to the comparison results, and the update guidance information is used to add sub - prompt items to the model preset prompt items. The prompt sub - item update guidance information and the update guidance information are used as the prompt update guidance information for subsequent updating of the model preset prompt items, so as to achieve the purpose of updating the large - language model.
[0102] Continuing with the above example, the question data, answer data, and answer evaluation data are input into the optimization model together to update the model preset prompt items of the large - language model. The answer evaluation data involves three evaluation dimensions: the logical evaluation dimension, the practicality evaluation dimension, and the language expression evaluation dimension. The evaluation data of each evaluation dimension can be compared with the answer data to determine whether there are advantages and disadvantages mentioned in the answer evaluation data in the answer data, and optimization suggestions are given for the disadvantages of the answer data to generate the prompt update guidance information. The answer evaluation data mentions that in the logical evaluation dimension, there is a lack of connection content between sentences in the answer data; in the practicality evaluation dimension, there is a lack of cases to solve industry pain points or improve efficiency in the answer data; in the language expression evaluation dimension, the writing style of the answer data needs to be adjusted, and metaphors or specific cases should be appropriately used to enrich the answer data. Based on the optimization suggestions for each dimension, case - based prompt items for adding sub - prompt items can be generated for the model preset prompt items to obtain the update guidance information; for the preset sub - prompt items in the model preset prompt items: in the topic instruction constraints, add constraints on the writing style to obtain the prompt update guidance information; and add prompt items for logical connection to the thinking mode prompt items in the model preset prompt items to obtain the update guidance information.
[0103] In summary, the prompt sub - item update guidance information and the update guidance information are used as the prompt update guidance information for subsequent updating of the model preset prompt items to achieve a comprehensive update of the large - language model.
[0104] Step 208: Update the model preset prompt items of the large - language model based on the prompt update guidance information until a target large - language model that meets the model update conditions is obtained.
[0105] Specifically, after inputting the question data, answer data, and answer evaluation data into the optimization model for processing to obtain the prompt update guidance information of the model preset prompts, the model preset prompts of the large language model can be updated based on the prompt update guidance information until a target large language model that meets the model update conditions is obtained. Among them, the model update conditions can be preset model detection conditions, that is, after updating the model preset prompts of the large language model to obtain the target large language model, the question data can be input into the target large language model for processing to obtain the target answer data, and it is detected whether the target answer data matches the answer evaluation data. If it matches, it is determined that the target large language model meets the model update conditions.
[0106] Based on this, after inputting the question data, answer data, and answer evaluation data into the optimization model for processing to obtain the prompt update guidance information of the model preset prompts, the model preset prompts of the large language model are updated based on the prompt update guidance information, that is, the sub-prompts in the model preset prompts are updated based on the update guidance of the prompt update guidance information to obtain the target large language model. By detecting the target large language model based on the question data, it is determined whether the target large language model meets the model update conditions. Until a target large language model that meets the model update conditions is obtained.
[0107] Furthermore, when updating the model preset prompts of the large language model, multi-round iterative updates can be performed based on the question set, answer set, and answer evaluation set by using the optimization model. The specific implementation is as follows:
[0108] In the case where the question data is a question set, the answer data is an answer set, and the answer evaluation data is an answer evaluation set, determine the iterative update round corresponding to the large language model; according to the iterative update round, generate a prompt item update sample set based on the question set, the answer set, and the answer evaluation set; divide the prompt item update sample set into at least one prompt item update sample group based on a preset number of groups, and determine the group index corresponding to the at least one prompt item update sample group; use the optimization model to perform iterative processing on the prompt item updates included in each prompt item update sample group based on the at least one prompt item update sample group and the group index corresponding to the at least one prompt item update sample group to obtain the prompt update guidance information of the model preset prompts.
[0109] Specifically, the question set contains multiple question data. The multiple answer data contained in the answer set corresponds to the multiple question data in the question set, that is, the answer data is the answer to the question data. The answer evaluation set contains multiple answer evaluation data. The answer evaluation data in the answer evaluation set corresponds to the question data in the question set and the answer data in the answer set. The iteration update round refers to the number of iterations for updating the model preset prompt items of the large language model. The prompt item update sample set is a sample set composed of the data in the question set, the answer set, and the answer evaluation set, and is used to update the model preset prompt items of the large language model. A piece of sample data in the prompt item update sample set contains question data, answer data, and answer evaluation data, and a piece of sample data in the prompt item update sample set is the input content for optimizing the model. The prompt item update sample group is a sample group obtained by grouping the sample data in the prompt item update sample set. The group index corresponding to the prompt item update sample group is used to record the number of prompt item update sample groups.
[0110] Based on this, determine the iteration update round corresponding to the large language model. Multiple rounds of iterative updates can be performed on the model preset prompt items based on the iteration update round. According to the iteration update round, generate a prompt item update sample set based on the question set, the answer set, and the answer evaluation set. Determine the preset number of groups, and divide the prompt item update sample set into at least one prompt item update sample group based on the preset number of groups. Each prompt item update sample group can contain the same number of prompt update samples. Determine the group index corresponding to at least one prompt item update sample group. Combine the prompt item update sample group and the group index, and use the optimization model to perform iterative processing on each prompt item update sample in the prompt item update sample set in turn, that is, perform iterative processing on the prompt item update samples contained in each prompt item update sample group, and obtain the prompt update guiding information of the model preset prompt items.
[0111] Continuing with the above example, set the number of iterative update rounds to 10; based on the updated sample set N (including 12 updated samples of prompt items) determined by the prompt items, determine the preset number of groups to be 4, and the corresponding group indices for each group to be 0 - 3. In the first round of iteration, select the group of updated samples of the prompt item corresponding to index 0, and based on the 4 updated samples of the prompt item included in the group of updated samples corresponding to index 0, use the optimization model to update the model preset prompt items of the large language model, that is, generate prompt update guiding information that can update the model preset prompt items. This is done until the use of 4 groups of updated samples of prompt items is completed based on group indices 0 - 3, that is, generate prompt update guiding information based on the 4 groups of updated samples of prompt items and the optimization model in sequence. And so on, until 10 rounds of iterative updates are completed, and obtain the prompt update guiding information for updating the model preset prompt items. Based on the optimization suggestions in each dimension, case prompt items for adding sub-prompt items can be generated for the model preset prompt items to obtain the update guiding information; for the preset sub-prompt items in the model preset prompt items: in the topic instruction constraints, add a constraint item for the writing style to obtain the prompt update guiding information; and add a prompt item for logical connection to the thinking mode prompt items in the model preset prompt items to obtain the prompt update guiding information. Update the model preset prompt items based on the update guiding information and / or the prompt update guiding information, and the target large language model can be obtained.
[0112] In summary, iterative processing is performed on the updated samples of the prompt items included in each group of updated samples of prompt items to obtain the prompt update guiding information for the model preset prompt items. By setting the number of iterative update rounds and updating the model preset prompt items of the large language model in an iterative update manner, the direct problem-solving ability, reflection ability, critical ability, and planning ability of the large language model are fully explored and integrated.
[0113] The large language model update method provided by an embodiment of this specification processes problem data using a large language model according to pre-set model prompts to obtain answer data. The pre-set model prompts are used to guide the prediction behavior of the large language model. The pre-set model prompts can provide system instructions when the large language model processes problem data, providing the large language model with thinking methods, knowledge bases, and problem processing capabilities that can be used. The problem data and answer data are input into an evaluation model for processing to obtain answer evaluation data corresponding to the large language model, realizing the evaluation and criticism of the answer data and the answer process of the large language model. The problem data, answer data, and answer evaluation data are input into an optimization model for processing to obtain prompt update guiding information for the pre-set model prompts. Based on the prompt update guiding information, the pre-set model prompts of the large language model are updated until a target large language model that meets the model update conditions is obtained. The prompt update guiding information can be used to optimize the pre-set model prompts of the large language model, achieving the purpose of updating the large language model. By combining the evaluation model and the optimization model, the large language model optimizes the pre-set model prompts in a self-playing manner, and the update of the large language model can be completed without using the model update method of model training, which can reduce the consumption of training resources and improve the update efficiency and flexibility of the large language model.
[0114] The following combines the attached Figure 3 , taking the application of the large language model update method provided by this specification in the update of the large language model in the scenario of solving math problems as an example, to further illustrate the large language model update method. Among them, Figure 3 shows the processing process flow chart of a large language model update method provided by an embodiment of this specification, specifically including the following steps.
[0115] Step 302: Input the problem data into the large language model, and determine the prompt sub-items in the pre-set model prompts according to the problem attributes of the problem data.
[0116] In practical applications, the large language model update method can be applied to the update of the large language model in the scenario of solving math problems. When the problem data is a math application problem, the math application problem is input into the large language model. The prompt sub-items related to the solution of the math application problem are determined in the system instructions (pre-set model prompts) of the large language model, that is, identity: mathematician; ability: mathematical reasoning ability and the ability to handle knowledge in the subject of mathematics; thinking methods: step-by-step thinking, reflection, backtracking, and self-correction, etc.
[0117] Step 304: Use the large language model to process the problem data according to the prompt sub-items to obtain answer data.
[0118] The large language model will solve the math application problem according to the prompt sub-items to obtain answer data.
[0119] Step 306: Input the question data and the answer data into the evaluation model, and determine the logical evaluation dimension, information security evaluation dimension, practicality evaluation dimension, and language expression evaluation dimension corresponding to the evaluation model.
[0120] The evaluation model evaluates the current answer of the large language model, and gives the shortcomings and suggestions of the answer result.
[0121] Step 308: Use the evaluation model to process the question data and the answer data according to the logical evaluation dimension, information security evaluation dimension, practicality evaluation dimension, and / or language expression evaluation dimension, and obtain the answer evaluation data corresponding to the large language model.
[0122] In specific implementation, the evaluation model can evaluate the answer of the large language model from multiple dimensions such as the logical evaluation dimension, information security evaluation dimension, practicality evaluation dimension, and language expression evaluation dimension.
[0123] When the large language model answers a math application problem and only outputs the answer to the math application problem, the evaluation model can evaluate the current answer from the evaluation dimension of problem-solving practicality. Since the key in the solution process of a math application problem lies in the derivation process and the problem-solving steps. Therefore, in the practicality evaluation dimension, the practicality of the answer data given by the large language model is weak; in the logical evaluation dimension, the answer data given by the large language model does not have the reasoning logic that a math application problem should have; in the language expression evaluation dimension, due to the lack of analysis of the application problem, the language expression of the answer result is relatively weak.
[0124] Step 310: Input the question data, the answer data, and the answer evaluation data into the optimization model for processing, and obtain the prompt update guiding information of the model preset prompt items.
[0125] Furthermore, input the math application problem, the math application problem answer, and the answer evaluation data into the optimization model, analyze the application problem answer data and the answer evaluation data, and use its planning and reflection ability to give the prompt update guiding information that can update the system instructions of the large language model, which is used to guide the optimization of the system instructions of the large language model. Considering that when answering a math application problem based on the system instructions of the large language model, only the answer to the math application problem is given, therefore, the prompt update guiding information can be: add the analysis of the application problem stem and give the detailed solution steps.
[0126] Step 312: Update the model preset prompt items of the large language model based on the prompt update guiding information until the target large language model that meets the model update conditions is obtained.
[0127] Update the system instructions of the large language model based on the prompts given by the optimization model to update the guiding information, and improve the ability of the large language model to solve math application problems.
[0128] In practical applications, the system instructions of the large language model can be updated by iterative update. In the preparation stage, construct a validation set val containing math application problems, and determine the large language model π, the evaluation model πc, the optimization model πp, and initialize the system instructions x of the large language model.
[0129] for iter = 1, 2,... do (iter represents the iteration round)
[0130] Generate the answer set of the validation set: {π(vali), i = 0, 1, 2,..., M}
[0131] Generate the evaluation information set (criticles) of the above validation set and answer set: {πc(vali, π(vali)), i = 0, 1, 2,..., M}
[0132] for index, batch in M (M represents the number of samples composed of the validation set, answer set and evaluation information set, M is greater than or equal to 1; batch represents the number of sample groups obtained by grouping M, and the value of batch can be set according to needs; index represents the index of batch)
[0133] i = index * batch
[0134] j = (index + 1) * batch - 1
[0135] Update the system instructions x batch by batch x <-- πp(x, (vali, π(vali), πc(vali, π(vali))),...(valj, π(valj), πc(valj, π(valj))))
[0136] Obtain the updated system instructions x, and finally obtain the updated large language model, that is, the target large language model.
[0137] Step 314: Use the target large language model to process the target problem data according to the updated model preset prompt items to obtain the target answer data corresponding to the target problem data.
[0138] After updating the system instructions of the large language model, other problem inferences and answers can be made based on the updated system instructions.
[0139] Step 316: Integrate the target problem data and the updated model preset prompt items into target model preset prompt items.
[0140] Step 318: Use the target large language model to process the target problem data according to the target model preset prompt items to obtain target answer data corresponding to the target problem data.
[0141] After updating the system instructions of the large language model, it is also possible to integrate the questions to be answered and the updated system instructions. After the integration is completed, use the large language model to perform reasoning and answering of other questions.
[0142] In summary, the update process of the system instructions of the large language model does not involve model training (SFT or RLHF). In the iterative update process of the system instructions, an interactive method of self-play is used. By iteratively optimizing the system instructions, the self-enhancing effect is achieved. The iterative update process of the system instructions can fully explore and integrate the direct problem-solving ability, reflection ability, critical ability, and planning ability of the model, and better induce the ability of the large language model itself.
[0143] Corresponding to the above method embodiments, this specification also provides embodiments of a large language model update device. Figure 4 The structural schematic diagram of a large language model update device provided by an embodiment of this specification is shown. As Figure 4 shown, the device includes:
[0144] A large language model processing module 402, configured to use the large language model to process the problem data according to the model preset prompt items to obtain answer data, and the model preset prompt items are used to guide the prediction behavior of the large language model;
[0145] An evaluation model processing module 404, configured to input the problem data and the answer data into an evaluation model for processing to obtain answer evaluation data corresponding to the large language model;
[0146] An optimization model processing module 406, configured to input the problem data, the answer data, and the answer evaluation data into an optimization model for processing to obtain prompt update guiding information of the model preset prompt items;
[0147] An update module 408, configured to update the model preset prompt items of the large language model based on the prompt update guiding information until a target large language model that meets the model update conditions is obtained.
[0148] In an optional embodiment, the large language model processing module 402 is further configured to:
[0149] Input the problem data into the large language model, and determine a prompt sub-item from the model preset prompt items according to the problem attributes of the problem data;
[0150] Use the large language model to process the problem data according to the prompt sub-item to obtain the answer data.
[0151] In an optional embodiment, the evaluation model processing module 404 is further configured to:
[0152] Input the problem data and the answer data into an evaluation model, and determine the corresponding logical evaluation dimension, information security evaluation dimension, practicality evaluation dimension, and language expression evaluation dimension of the evaluation model;
[0153] Use the evaluation model to process the problem data and the answer data according to the logical evaluation dimension, the information security evaluation dimension, the practicality evaluation dimension, and / or the language expression evaluation dimension to obtain the answer evaluation data corresponding to the large language model.
[0154] In an optional embodiment, the optimization model processing module 406 is further configured to:
[0155] Use the optimization model to compare the answer data and the answer evaluation data based on the problem data, and determine the prompt sub-item update guidance information corresponding to the preset sub-prompt item in the model preset prompt items, and the update guidance information corresponding to the model preset prompt items;
[0156] Use the prompt sub-item update guidance information and the update guidance information as the prompt update guidance information.
[0157] In an optional embodiment, the update module 408 is further configured to:
[0158] Determine the iterative update round corresponding to the large language model;
[0159] According to the iterative update round, generate a prompt item update sample set based on the problem set, the answer set, and the answer evaluation set;
[0160] Divide the prompt item update sample set into at least one prompt item update sample group based on a preset number of groups, and determine the group index corresponding to the at least one prompt item update sample group;
[0161] Use the optimization model to perform iterative processing on the prompt item updates included in each prompt item update sample group based on the at least one prompt item update sample group and the group index corresponding to the at least one prompt item update sample group to obtain the prompt update guidance information of the model preset prompt items.
[0162] An optionally implemented example, the large language model processing module 402 is further configured to:
[0163] Select sample problem data from the training sample set and input it into the initial large language model to obtain predicted answer data;
[0164] Select sample answer data corresponding to the sample problem data from the training sample set, and calculate the large language model loss value based on the large language model loss function, the sample answer data, and the predicted answer data;
[0165] Adjust the parameters of the initial large language model based on the large language model loss value until the large language model that meets the large language model training stop condition is obtained.
[0166] An optionally implemented example, the evaluation model processing module 404 is further configured to:
[0167] Select the sample data to be evaluated from the evaluation model training sample set and input it into the initial evaluation model to obtain answer evaluation data;
[0168] Select the sample answer evaluation data corresponding to the sample data to be evaluated from the evaluation model training sample set, and calculate the evaluation model loss value based on the evaluation model loss function, the sample data to be evaluated, and the answer evaluation data;
[0169] Adjust the parameters of the initial evaluation model based on the evaluation model loss value until the evaluation model that meets the evaluation model training stop condition is obtained.
[0170] An optionally implemented example, the optimization model processing module 406 is further configured to:
[0171] Select the sample data to be optimized from the optimization model training sample set and input it into the initial optimization model to obtain prompt optimization data, where the sample data to be optimized includes a problem sample, the answer sample corresponding to the problem sample, and the answer evaluation sample corresponding to the answer sample;
[0172] Select the sample prompt optimization data corresponding to the sample data to be optimized from the optimization model training sample set, and calculate the optimization model loss value based on the optimization model loss function, the sample data to be optimized, and the prompt optimization data;
[0173] Adjust the parameters of the initial optimization model based on the optimization model loss value until the optimization model that meets the optimization model training stop condition is obtained.
[0174] The large language model update device provided by an embodiment of this specification processes problem data using a large language model according to model preset prompt items to obtain answer data. The model preset prompt items are used to guide the prediction behavior of the large language model. According to the model preset prompt items, system instructions can be provided when the large language model processes problem data, providing the large language model with available thinking methods, knowledge bases, and problem processing capabilities. Input the problem data and answer data into an evaluation model for processing to obtain answer evaluation data corresponding to the large language model, realizing the evaluation and criticism of the answer data and the answer process of the large language model. Input the problem data, answer data, and answer evaluation data into an optimization model for processing to obtain prompt update guidance information for the model preset prompt items. Update the model preset prompt items of the large language model based on the prompt update guidance information until a target large language model that meets the model update conditions is obtained. The prompt update guidance information can be used to optimize the model preset prompt items of the large language model to achieve the purpose of updating the large language model. Combining the evaluation model and the optimization model enables the large language model to optimize the model preset prompt items in a self-playing manner, and the large language model can be updated without using the model update method of model training, which can reduce the consumption of training resources and improve the update efficiency and flexibility of the large language model.
[0175] The above is a schematic solution of a large language model update device according to this embodiment. It should be noted that the technical solution of this large language model update device and the technical solution of the above large language model update method belong to the same concept. For the details not described in the technical solution of the large language model update device, reference can be made to the description of the technical solution of the above large language model update method.
[0176] See Figure 5 , Figure 5 shows a flowchart of a question-and-answer method provided according to an embodiment of this specification, specifically including the following steps.
[0177] Step 502: Determine the target problem data;
[0178] Step 504: Input the target problem data into a target large language model for processing to obtain target answer data corresponding to the target problem data.
[0179] In practical applications, after updating the model preset prompt items of the large language model and obtaining the target large language model, the target large language model can be used to process any target problem data. Input the target problem data into the target large language model, and use the target large language model to answer the target problem data according to the updated model preset prompt items to obtain the target answer data.
[0180] For example, in the scenario of text generation and creation, large language models have the ability to generate and create text. The target question data of "Please write an article describing the scenery" can be input into the large language model. After understanding the question, the large language model can generate an article about the scenery.
[0181] Furthermore, after determining the updated model preset prompt items, the updated model preset prompt items can be used as the system settings of the target large language model to continue processing other question data. The specific implementation is as follows:
[0182] Use the target large language model to process the target question data according to the updated model preset prompt items to obtain the target answer data corresponding to the target question data.
[0183] Based on this, take the updated model preset prompt items as the system instructions of the target large language model. Use the target large language model to process the target question data according to the updated model preset prompt items. Under the guidance of the updated model preset prompt items, generate the target answer data corresponding to the target question data.
[0184] Continuing with the above example, set the updated model preset prompt items as the system settings of the target large language model and perform normal reasoning on the question. System settings: "You are a science fiction writer, good at creating stories about future technology and social change."; Question: "Please write a short story about future city life."; Reasoning process and answer: In this reasoning mode, the large language model will use the system settings as its creation background. When receiving the question, the model will conceive and generate a short story about future city life based on the identity of "science fiction writer" and the setting of "good at creating stories about future technology and social change". The story may include elements of future technology such as intelligent buildings, self-driving cars, virtual reality, etc., and explore how these technologies affect city life and people's social relationships.
[0185] To sum up, using the target large language model to process the target question data according to the updated model preset prompt items to obtain the target answer data corresponding to the target question data makes the target answer data have higher accuracy and comprehensiveness.
[0186] Furthermore, after determining the updated model preset prompt items, by integrating the updated model preset prompt items and the next question data to be processed, the integrated content can be used as the system settings of the target large language model to continue processing the next question data. The specific implementation is as follows:
[0187] Integrate the target problem data and the updated model preset prompt items into target model preset prompt items; use the target large language model to process the target problem data according to the target model preset prompt items to obtain target answer data corresponding to the target problem data.
[0188] Based on this, use the updated model preset prompt items as the system instructions for the target large language model. Integrate the target problem data and the updated model preset prompt items into target model preset prompt items. Use the target large language model to process the target problem data according to the target model preset prompt items, and under the guidance of the target model preset prompt items, generate target answer data corresponding to the target problem data.
[0189] Continuing with the above example, connect or integrate the updated model preset prompt items with the question raised by the user, and then send it to the target large language model for reasoning. System settings of the target large language model: "You are a historical novelist, good at integrating real events into novel plots."; Question: "Please write a story about a small person during the war."; Reasoning process and answer: In this reasoning method, we will splice the system settings (updated model preset prompt items) and the user's question together to form a new input: "You are a historical novelist, good at integrating real historical events into novel plots.\n\nPlease write a story about a small person during the war." Then, send this spliced input to the target large language model for reasoning. When the target large language model receives this spliced input, it will first understand and absorb the system settings, that is, the identity of "historical novelist" and the specialty of "good at integrating real historical events into novel plots". Then, the target large language model will conceive and generate a story according to the specific requirements of the question, that is, "write a story about a small person during the war". This story may revolve around an ordinary soldier, civilian or refugee, and through their perspectives, show the cruelty of the war and the glory of human nature, while integrating real historical events as background or plot elements.
[0190] To sum up, determine the target problem data, input the target problem data into the target large language model for processing to obtain target answer data corresponding to the target problem data. During the process of processing the target problem data, answer the target problem data according to the updated model preset prompt items, so as to improve the practicability, accuracy and comprehensiveness of the target answer data.
[0191] Corresponding to the above method embodiments, this specification also provides embodiments of a large language model update device. Figure 6 Shows a schematic structural diagram of a question-answering device provided by an embodiment of this specification. As Figure 6 shown, the device includes:
[0192] A determination module 602, configured to determine target problem data;
[0193] A processing module 604, configured to input the target problem data into a target large language model for processing to obtain target answer data corresponding to the target problem data.
[0194] In an optional embodiment, the processing module 604 is further configured to:
[0195] Use the target large language model to process the target problem data according to the updated model preset prompt items to obtain target answer data corresponding to the target problem data.
[0196] In an optional embodiment, the processing module 604 is further configured to:
[0197] Integrate the target problem data and the updated model preset prompt items into a target model preset prompt item;
[0198] Use the target large language model to process the target problem data according to the target model preset prompt items to obtain target answer data corresponding to the target problem data.
[0199] In summary, determine the target problem data, input the target problem data into the target large language model for processing, and obtain the target answer data corresponding to the target problem data. During the process of processing the target problem data, answer the target problem data according to the updated model preset prompt items, so as to improve the practicability, accuracy, and comprehensiveness of the target answer data.
[0200] The above is a schematic solution of a large language model update device according to this embodiment. It should be noted that the technical solution of this large language model update device and the technical solution of the above large language model update method belong to the same concept. For the details not described in the technical solution of the large language model update device, reference can be made to the description of the technical solution of the above large language model update method.
[0201] See Figure 7 , Figure 7 shows a flowchart of another question - answering method provided according to an embodiment of this specification, which specifically includes the following steps.
[0202] Step 702: Receive the question data to be answered submitted by the user;
[0203] Step 704: Use the large language model to process the data of the question to be answered according to the pre-set prompt items of the model, obtain the target answer data, and send the target answer data to the user. The pre-set prompt items of the model are used to guide the prediction behavior of the large language model;
[0204] Step 706: Input the data of the question to be answered and the target answer data into an evaluation model for processing to obtain the answer evaluation data corresponding to the large language model;
[0205] Step 708: Input the data of the question to be answered, the target answer data, and the answer evaluation data into an optimization model for processing to obtain the prompt update guiding information for the pre-set prompt items of the model;
[0206] Step 710: Update the pre-set prompt items of the large language model based on the prompt update guiding information until a target large language model that meets the model update conditions is obtained.
[0207] In practical applications, while the large language model answers the data of the question to be answered submitted by the user, the pre-set prompt items of the large language model can be updated based on the data of the question to be answered and the target answer data output by the large language model. After receiving the data of the question to be answered submitted by the user, input the data of the question to be answered into the large language model. Use the large language model to process the data of the question to be answered according to the pre-set prompt items of the model to obtain the target answer data. Send the target answer data as the answer to the data of the question to be answered to the user. At the same time, after obtaining the target answer data, the data of the question to be answered and the target answer data can be input into an evaluation model for processing together. Use the evaluation model to evaluate the answering process and the answering result of the large language model for the data of the question to be answered, and give the shortcomings and improvements of this answering. The evaluation model gives the answer evaluation data for the current answering of the large language model.
[0208] Furthermore, input the data of the question to be answered, the target answer data, and the answer evaluation data into an optimization model for processing. Use the planning and reflection capabilities of the optimization model to analyze the data of the question to be answered, the target answer data, and the answer evaluation data to obtain the prompt update guiding information for the pre-set prompt items of the model. The prompt update guiding information is used to guide the update of the pre-set prompt items of the large language model. Update the pre-set prompt items of the large language model based on the prompt update guiding information until a target large language model that meets the model update conditions is obtained.
[0209] For example, large language models can be applied to the field of language translation. The user submits the data of the question to be answered, that is, the text content to be translated "Artificial intelligence is transforming the way we live and work, offering unprecedented opportunities for innovation and efficiency". The text content to be translated is input into the large language model, and the large language model can give the target answer data, that is, the translation result "Artificial intelligence is changing the way we live and work, providing unprecedented opportunities for innovation and efficiency". After obtaining the translation result, the text content to be translated and the translation result can be input into the evaluation model to evaluate the data processing of the large language model this time. The evaluation model can give the answer evaluation data "After the translation is completed, proofreading is required to ensure that there are no grammar errors; the translation should be targeted at readers in the technology field, and the language style should be professional and easy to understand". The text content to be translated, the translation result, and the answer evaluation data that the evaluation model can give are input into the optimization model to update the model preset prompt items (system instructions) of the large language model. Set the language style and text grammar proofreading as prompt sub-items and add them to the model preset prompt items to achieve the update of the large language model.
[0210] In summary, the prompt update guidance information can be used to optimize the model preset prompt items of the large language model, so as to achieve the purpose of updating the large language model. Combining the evaluation model and the optimization model enables the large language model to optimize the model preset prompt items in a self-playing manner, and the update of the large language model can be completed without using the model update method of model training, which can reduce the consumption of training resources and improve the update efficiency and update flexibility of the large language model. While providing question-and-answer services for users, the model preset prompt items of the large language model are updated based on the data of the question to be answered submitted by the user, continuously improving the question-and-answer ability of the large language model.
[0211] Corresponding to the above method embodiments, this specification also provides embodiments of a large language model update device. Figure 8 It shows the structural schematic diagram of another question-and-answer device provided by an embodiment of this specification. As Figure 8 shown, the device includes:
[0212] A receiving module 802, configured to receive the data of the question to be answered submitted by the user;
[0213] A sending module 804, configured to process the question data to be answered by using a large language model according to the pre-set prompt items of the model to obtain target answer data, and send the target answer data to the user, where the pre-set prompt items of the model are used to guide the prediction behavior of the large language model;
[0214] An input module 806, configured to input the question data to be answered and the target answer data into an evaluation model for processing to obtain answer evaluation data corresponding to the large language model;
[0215] A processing module 808, configured to input the question data to be answered, the target answer data, and the answer evaluation data into an optimization model for processing to obtain prompt update guiding information for the pre-set prompt items of the model;
[0216] An update module 810, configured to update the pre-set prompt items of the large language model based on the prompt update guiding information until a target large language model that meets the model update condition is obtained.
[0217] The question-answering device provided by an embodiment of this specification, while providing a question-answering service for a user, inputs the question data to be answered, the target answer data, and the answer evaluation data into an optimization model for processing. By using the planning and reflection capabilities of the optimization model, the question data to be answered, the target answer data, and the answer evaluation data are analyzed to obtain prompt update guiding information for the pre-set prompt items of the model. The prompt update guiding information is used to guide the update of the pre-set prompt items of the large language model. The pre-set prompt items of the large language model are updated based on the prompt update guiding information until a target large language model that meets the model update condition is obtained. The prompt update guiding information can be used to optimize the pre-set prompt items of the large language model, achieving the purpose of updating the large language model. By combining the evaluation model and the optimization model, the large language model optimizes the pre-set prompt items in a self-playing manner, and the update of the large language model can be completed without using the model update method of model training, which can reduce the consumption of training resources and improve the update efficiency and update flexibility of the large language model. While providing a question-answering service for a user, the pre-set prompt items of the large language model are updated based on the question data to be answered proposed by the user, continuously improving the question-answering ability of the large language model.
[0218] The above is a schematic solution of a large language model update device according to this embodiment. It should be noted that the technical solution of this large language model update device and the technical solution of the above large language model update method belong to the same concept. For the details not described in the technical solution of the large language model update device, reference can be made to the description of the technical solution of the above large language model update method.
[0219] Figure 9FIG. 0 shows a structural block diagram of a computing device 900 provided according to an embodiment of this specification. The components of the computing device 900 include, but are not limited to, a memory 910 and a processor 920. The processor 920 is connected to the memory 910 via a bus 930, and a database 950 is used to store data.
[0220] The computing device 900 further includes an access device 940, which enables the computing device 900 to communicate via one or more networks 960. Examples of such networks include the Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or a combination of communication networks such as the Internet. The access device 940 may include one or more of any type of wired or wireless network interfaces (e.g., a network interface card (NIC)), such as an IEEE 802.11 Wireless Local Area Network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, or a Near Field Communication (NFC) interface.
[0221] In an embodiment of this specification, the above components of the computing device 900 and Figure 9 other components not shown therein may also be connected to each other, e.g., via a bus. It should be understood that Figure 9 the shown structural block diagram of the computing device is for illustrative purposes only and is not a limitation on the scope of this specification. Those skilled in the art may add or replace other components as needed.
[0222] The computing device 900 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or personal computers (PCs). The computing device 900 can also be a mobile or stationary server.
[0223] Wherein, the processor 920 is configured to execute the following computer program or instructions, and when the computer program or instructions are executed by the processor, the steps of the above method are implemented.
[0224] Each embodiment in this specification is described in a progressive manner. For the parts that are the same or similar among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the embodiment of the computing device, since it is basically similar to the embodiments of the large language model update method and the question and answer method, the description is relatively simple. For the relevant parts, reference can be made to the partial descriptions of the embodiments of the large language model update method and the question and answer method.
[0225] An embodiment of this specification also provides a computer-readable storage medium, which stores a computer program or instructions, and when the computer program or instructions are executed by a processor, the steps of the above method are implemented.
[0226] Each embodiment in this specification is described in a progressive manner. For the parts that are the same or similar among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the embodiment of the computer-readable storage medium, since it is basically similar to the embodiments of the large language model update method and the question and answer method, the description is relatively simple. For the relevant parts, reference can be made to the partial descriptions of the embodiments of the large language model update method and the question and answer method.
[0227] An embodiment of this specification also provides a computer program product, including a computer program or instructions, and when the computer program or instructions are executed by a processor, the steps of the above method are implemented.
[0228] The above is a schematic solution of a computer program product of this embodiment. It should be noted that the technical solution of this computer program product and the technical solutions of the above large language model update method and question and answer method belong to the same concept. For the details not described in detail in the technical solution of the computer program product, reference can be made to the descriptions of the technical solutions of the above large language model update method and question and answer method.
[0229] The above describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired results. Additionally, the processes depicted in the drawings do not necessarily require the particular order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0230] The computer instructions include computer program code, which may be in source code form, object code form, executable file, or some intermediate form, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, removable hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of patent practice. For example, in some regions, according to patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0231] It should be noted that for the foregoing method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should be aware that the embodiments of this specification are not limited by the described order of actions, because according to the embodiments of this specification, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the embodiments of this specification.
[0232] In the above embodiments, the descriptions of the various embodiments have their own focuses. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0233] The preferred embodiments of this specification disclosed above are only used to help explain this specification. The alternative embodiments do not elaborate on all the details, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and variations can be made according to the content of the embodiments of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the embodiments of this specification, so that those skilled in the art can well understand and utilize this specification. This specification is only limited by the claims and their full scope and equivalents.
Claims
1. A large language model updating method, characterized in that: include: Processing the question data according to the preset prompt items of the model using the large language model to obtain answer data, wherein the preset prompt items of the model are used to guide the prediction behavior of the large language model; Inputting the question data and the answer data into an evaluation model for processing to obtain answer evaluation data corresponding to the large language model; Inputting the question data, the answer data and the answer evaluation data into the optimization model for processing, and obtaining prompt update guidance information of the preset prompt items of the model; The model preset prompt items of the large language model are updated based on the prompt update guidance information until a target large language model that meets the model update condition is obtained.
2. The large language model updating method according to claim 1, characterized in that: The method of using the large language model to process the question data according to the preset prompt items of the model to obtain the answer data includes: Inputting the question data into the large language model, and determining prompt sub-items in the model preset prompt items according to the question attributes of the question data; The question data is processed according to the prompt sub-items using the large language model to obtain the answer data.
3. The large language model updating method according to claim 1, characterized in that: The step of inputting the question data and the answer data into an evaluation model for processing to obtain answer evaluation data corresponding to the large language model includes: Inputting the question data and the answer data into an evaluation model, and determining a logic evaluation dimension, an information security evaluation dimension, a practicality evaluation dimension, and a language expression evaluation dimension corresponding to the evaluation model; The evaluation model is used to process the question data and the answer data according to the logic evaluation dimension, the information security evaluation dimension, the practicality evaluation dimension and / or the language expression evaluation dimension to obtain the answer evaluation data corresponding to the large language model.
4. The large language model updating method according to claim 1, characterized in that: The step of inputting the question data, the answer data and the answer evaluation data into the optimization model for processing to obtain prompt update guidance information of the preset prompt items of the model includes: Using the optimization model to compare the answer data and the answer evaluation data based on the question data, and determining prompt sub-item update guidance information corresponding to the preset sub-prompt item in the model preset prompt item and update guidance information corresponding to the model preset prompt item according to the comparison result; The prompt sub-item update guide information and the update guide information are used as the prompt update guide information.
5. The large language model updating method according to claim 1, characterized in that: In the case where the question data is a question set, the answer data is an answer set, and the answer evaluation data is an answer evaluation set, the question data, the answer data, and the answer evaluation data are input into an optimization model for processing to obtain prompt update guidance information of the model preset prompt items, including: Determining an iterative update round corresponding to the large language model; According to the iterative update round, generating a prompt item update sample set based on the question set, the answer set and the answer evaluation set; Dividing the prompt item update sample set into at least one prompt item update sample group based on a preset number of groups, and determining a group index corresponding to the at least one prompt item update sample group; The optimization model is used to iteratively process the prompt item update samples contained in each prompt item update sample group based on the at least one prompt item update sample group and the group index corresponding to the at least one prompt item update sample group to obtain the prompt update guidance information of the model preset prompt item.
6. The large language model updating method according to claim 1, characterized in that: The training of the large language model includes: Select sample question data from the training sample set and input it into the initial large language model to obtain predicted answer data; Selecting sample answer data corresponding to the sample question data in the training sample set, and calculating a large language model loss value based on a large language model loss function, the sample answer data, and the predicted answer data; The initial large language model is adjusted based on the large language model loss value until the large language model that meets the large language model training stop condition is obtained.
7. The large language model updating method according to claim 1, characterized in that: The training of the evaluation model includes: Select sample data to be evaluated from the evaluation model training sample set and input it into the initial evaluation model to obtain answer evaluation data; Selecting sample answer evaluation data corresponding to the sample data to be evaluated from the evaluation model training sample set, and calculating the evaluation model loss value based on the evaluation model loss function, the sample data to be evaluated and the answer evaluation data; The initial evaluation model is adjusted based on the evaluation model loss value until the evaluation model that meets the evaluation model training stop condition is obtained.
8. The large language model updating method according to claim 1, characterized in that: The training of the optimization model includes: Select sample data to be optimized from the optimization model training sample set and input it into the initial optimization model to obtain prompt optimization data, wherein the sample data to be optimized includes a question sample, an answer sample corresponding to the question sample, and an answer evaluation sample corresponding to the answer sample; Selecting sample prompt optimization data corresponding to the sample data to be optimized from the optimization model training sample set, and calculating the optimization model loss value based on the optimization model loss function, the sample data to be optimized and the prompt optimization data; The initial optimization model is adjusted based on the optimization model loss value until the optimization model that meets the optimization model training stop condition is obtained.
9. A question-answering method, characterized in that: Said, including: Determine target problem data; The target question data is input into the target large language model as described in any one of claims 1 to 5 for processing to obtain target answer data corresponding to the target question data.
10. The large language model updating method according to claim 9, characterized in that: The step of inputting the target question data into the target large language model as described in any one of claims 1 to 8 for processing to obtain target answer data corresponding to the target question data comprises: The target large language model is used to process the target question data according to the updated model preset prompt items to obtain target answer data corresponding to the target question data.
11. The large language model updating method according to claim 10, characterized in that: The step of inputting the target question data into the target large language model as described in any one of claims 1 to 8 for processing to obtain target answer data corresponding to the target question data comprises: Integrating the target question data and the updated model preset prompt item into a target model preset prompt item; The target large language model is used to process the target question data according to preset prompt items of the target model to obtain target answer data corresponding to the target question data.
12. A question-answering method, characterized in that: Said, including: Receive unanswered question data submitted by users; Using the large language model to process the question data to be answered according to the model preset prompt items, obtaining target answer data, and sending the target answer data to the user, wherein the model preset prompt items are used to guide the prediction behavior of the large language model; Inputting the to-be-answered question data and the target answer data into an evaluation model for processing to obtain answer evaluation data corresponding to the large language model; Inputting the unanswered question data, the target answer data and the answer evaluation data into the optimization model for processing, and obtaining prompt update guidance information of the model preset prompt items; The model preset prompt items of the large language model are updated based on the prompt update guidance information until a target large language model that meets the model update condition is obtained.
13. A computing device comprising: Memory and processor; The memory is used to store computer programs or instructions, and the processor is used to execute the computer programs or instructions. When the computer program or instructions are executed by the processor, the steps of the method described in any one of claims 1 to 12 are implemented.
14. A computer-readable storage medium storing a computer program or instruction, wherein the computer program or instruction, when executed by a processor, implements the steps of the method according to any one of claims 1 to 12.
15. A computer program product, comprising a computer program or instructions, which implement the steps of the method according to any one of claims 1 to 12 when executed by a processor.