Model optimization method and device

By using the method of scoring the model to be optimized and generating enhanced instructions, the problem that model optimization training in the existing technology depends on manual annotation is solved, the training efficiency and instruction quality are improved, labor costs are reduced, and the independent evolution of the model is realized.

CN120145133APending Publication Date: 2025-06-13BEIJING WODONG TIANJUN INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510140517.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-08
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The prior art relies on human expert experience and manual annotation in model optimization training, resulting in low quality of instruction results, low training efficiency and high labor cost.

Method used

By using the model to be optimized to score the answers of the sample instructions, the category of the sample instructions is determined, and the seed instructions are determined based on the target category, and the context information of the sample instructions and the seed instructions are generated to optimize the model to be optimized.

Benefits of technology

It improves the efficiency of model optimization training and the diversity and pertinence of instructions, reduces the dependence on manual design and annotation, reduces labor costs, and realizes the independent evolution and capability iteration of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120145133A_ABST
    Figure CN120145133A_ABST
Patent Text Reader

Abstract

The invention provides a model optimization method and device, and the method comprises the steps: carrying out the scoring of a to-be-optimized model for a sample instruction through a to-be-optimized model, so as to determine the type of the sample instruction; in response to the fact that the category of the sample instruction is a target category, determining a seed instruction corresponding to the sample instruction; and based on the context information of the sample instruction and the seed instruction, generating an enhancement instruction corresponding to the sample instruction, so as to optimize the to-be-optimized model based on the enhancement instruction. According to the embodiment, the diversity and pertinence of the instructions for model optimization can be improved, the closed-loop self-evolution of the to-be-optimized model is realized, the optimization training efficiency of the to-be-optimized model is improved, and the labor cost is saved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and more particularly, to a model optimization method and device. Background Art

[0002] Generally speaking, the training of large language models is divided into two stages. First, it is pre-trained using a large amount of unlabeled corpus to enable the model to possess a large amount of language knowledge. Subsequently is the alignment optimization stage, that is, through supervised training or reinforcement learning to make the model align with human preferences, which can activate the internal knowledge of the pre-trained large language model and make it reply and perform operations according to human needs. In order to make the model better follow instructions and comprehensively master various skills, the alignment optimization stage usually requires high-quality and wide-coverage labeled data to optimize and train the model.

[0003] However, currently this process mainly relies on expert experience and manual trial and error, that is, designing an appropriate instruction distribution through expert experience, and then obtaining result construction optimization data by manually annotating instruction results; this leads to the instruction results of the annotation and the instructions used for optimizing and training the model being limited by human levels, and there may be a situation of low quality, which not only makes the model optimization training inefficient, but also consumes a large amount of human costs. Summary of the Invention

[0004] In view of this, the embodiments of the present disclosure at least provide a model optimization method, device, electronic device and storage medium, which can improve the diversity and pertinence of the instructions for model optimization, while realizing the closed-loop self-evolution of the model to be optimized, improve the efficiency of optimizing and training the model to be optimized, and save human costs.

[0005] In a first aspect, the embodiments of the present disclosure provide a model optimization method, including:

[0006] Using the model to be optimized to score the answers of the model to be optimized for sample instructions to determine the category of the sample instructions;

[0007] In response to the category of the sample instructions being the target category, determining the seed instructions corresponding to the sample instructions;

[0008] Generating enhanced instructions corresponding to the sample instructions based on the context information of the sample instructions and the seed instructions, so as to optimize the model to be optimized based on the enhanced instructions.

[0009] Optionally, using the model to be optimized to score the answers of the model to be optimized for sample instructions includes:

[0010] Inputting the sample instructions into the model to be optimized to obtain multiple answer results of the model to be optimized for the sample instructions for multiple times;

[0011] Input multiple answer results and a sample instruction into the model to be optimized, so that the model to be optimized outputs scores for the multiple answer results.

[0012] Optionally, determine the category of the sample instruction, including:

[0013] Compare the scores of the multiple answer results with the classification threshold corresponding to the sample instruction; among them, the classification threshold corresponding to the sample instruction is preset, and different sample instructions have different corresponding classification thresholds;

[0014] In response to there being an answer result among the multiple answer results whose corresponding score is less than the classification threshold, determine the category of the sample instruction as the target category.

[0015] Optionally, determine the seed instruction corresponding to the sample instruction, including:

[0016] Convert the sample instruction into a vector representation;

[0017] Based on the vector representation of the sample instruction, determine the seed instruction corresponding to the sample instruction in the pre-constructed instruction library; the similarity between the seed instruction and the sample instruction exceeds the preset similarity threshold.

[0018] Optionally, optimize the model to be optimized based on the enhanced instruction, including:

[0019] Input the enhanced instruction into the model to be optimized, so that the model to be optimized outputs multiple answer results for multiple answers to the enhanced instruction;

[0020] Input the enhanced instruction and the multiple answer results corresponding to the enhanced instruction into the model to be optimized, so that the model to be optimized scores the multiple answer results corresponding to the enhanced instruction;

[0021] In response to there being an answer result among the multiple answer results corresponding to the enhanced instruction whose score does not exceed the preset threshold, generate a new enhanced instruction, and repeat the above steps until the scores of the multiple answer results corresponding to the enhanced instruction all exceed the preset threshold, obtaining the optimized model to be optimized.

[0022] Optionally, the method further includes:

[0023] In response to the category of the sample instruction being a non-target category, select the answer result with the highest answer score and the answer result with the lowest score from the multiple answer results corresponding to the sample instruction, and construct a triple training set corresponding to the sample instruction based on the answer result with the highest score, the answer result with the lowest score, and the sample instruction;

[0024] Optimize and train the model to be optimized based on the triple training set.

[0025] Second aspect, embodiments of the present disclosure provide a model optimization device, including:

[0026] A scoring module, configured to score the response of the model to be optimized for a sample instruction by using the model to be optimized, so as to determine the category of the sample instruction;

[0027] A determination module, configured to determine a seed instruction corresponding to the sample instruction in response to the category of the sample instruction being the target category;

[0028] An optimization module, configured to generate an enhanced instruction corresponding to the sample instruction based on the context information of the sample instruction and the seed instruction, so as to optimize the model to be optimized based on the enhanced instruction.

[0029] Third aspect, embodiments of the present disclosure further provide an electronic device, including: a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the computer device runs, the processor communicates with the memory through the bus. When the machine-readable instructions are executed by the processor, the steps in the first aspect, or any optional implementation manner in the first aspect are executed.

[0030] Fourth aspect, embodiments of the present disclosure further provide a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, the steps in the first aspect, or any optional implementation manner in the first aspect are executed.

[0031] Fifth aspect, embodiments of the present disclosure further provide a computer program product, including a computer program, which implements the method of any of the above embodiments when executed by a processor.

[0032] In any of the above aspects or any implementation manner of any aspect, by constructing a closed-loop self-evolution mechanism, efficient optimization training of the model to be optimized is achieved, and at the same time, the instruction diversity and pertinence are improved. First, the model to be optimized itself is used to score its response to the sample instruction, so as to screen out the sample instructions that need to be optimized and trained. Then, based on the sample instructions of the target category, the associated seed instructions are further determined. As the core reference for optimization, the seed instructions can ensure that the generated enhanced instructions have high relevance and pertinence. Next, by combining the context information of the sample instructions and the seed instructions, enhanced instructions are generated, making the instruction set richer in semantics and content, thereby improving the model's adaptability to different scenarios and requirements. This generation method not only expands the instruction diversity but also accurately captures the weaknesses of the model to be optimized, ensuring that the key direction of the optimization training is clear. By recycling the model itself for instruction generation and optimization training, the dependence on manual design and annotation is reduced, the labor cost is significantly reduced, and while improving the optimization efficiency, the self-evolution and ability iteration of the model to be optimized are realized.

[0033] For the effects of the above model optimization device, electronic device and storage medium, refer to the description of the above model optimization method, which will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the accompanying drawings required for use in the embodiments will be briefly introduced below. The accompanying drawings herein are incorporated into the specification and form a part of this specification. These drawings show embodiments consistent with the present disclosure and, together with the specification, are used to illustrate the technical solutions of the present disclosure. It should be understood that the following drawings only show some embodiments of the present disclosure and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0035] Figure 1 FIG. shows a flowchart of a model optimization method provided by an embodiment of the present disclosure;

[0036] Figure 2 FIG. shows a schematic flowchart of a model optimization method provided by an embodiment of the present disclosure;

[0037] Figure 3 FIG. shows a schematic diagram of a model optimization device provided by an embodiment of the present disclosure;

[0038] Figure 4 FIG. shows an exemplary system architecture to which the embodiments of the present disclosure can be applied;

[0039] Figure 5 FIG. shows a schematic diagram of the structure of a computer system of a terminal device or a server for implementing the embodiments of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0040] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only some of the embodiments of the present disclosure, rather than all of them. The components of the embodiments of the present disclosure usually described and illustrated in the accompanying drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present disclosure provided in the drawings is not intended to limit the scope of the present disclosure to be protected, but merely represents selected embodiments of the present disclosure. Based on the embodiments of the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present disclosure.

[0041] It should be noted that in the technical solution of the present invention, the processing of the collection, use, storage, sharing, transfer, etc. of the user's personal information complies with the provisions of relevant laws and regulations, and the user needs to be informed and obtain the consent or authorization of the user. When applicable, technical processing such as de-identification and / or anonymization and / or encryption is performed on the user's personal information.

[0042] The above problems and solutions are all the results obtained by the inventors after practice and careful research. The discovery process of the above problems and the solutions proposed for the above problems should be the contributions made by the inventors to this disclosure during the disclosure process.

[0043] Next, the technical solutions in this disclosure will be clearly and completely described in conjunction with the accompanying drawings in this disclosure. Obviously, the described embodiments are only a part of the embodiments of this disclosure, rather than all of the embodiments. Usually, the components of this disclosure described and illustrated in the accompanying drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this disclosure provided in the accompanying drawings is not intended to limit the scope of the claimed disclosure, but merely represents the selected embodiments of this disclosure. Based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of this disclosure.

[0044] It should be noted that similar reference numerals and letters indicate similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.

[0045] For ease of understanding of this embodiment, first, a model optimization method disclosed in the embodiments of this disclosure will be introduced in detail. The execution subject of the model optimization method provided in the embodiments of this disclosure is generally a computer device with certain computing capabilities. Such a computer device includes, for example: a terminal device, a server, or other processing devices. The terminal device can be a user equipment (UE), a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (PDA), a handheld device, a computing device, a vehicle-mounted device, a wearable device, etc. In some possible implementation manners, the model optimization method can be implemented by a processor calling computer-readable instructions stored in a memory.

[0046] See Figure 1 As shown, it is a flowchart of the model optimization method provided in the embodiments of this disclosure. The method includes S101 to S103, where:

[0047] S101: Score the responses of the model to be optimized for sample instructions using the model to be optimized, so as to determine the category of the sample instructions.

[0048] In the embodiments of the present disclosure, before scoring the responses of the model to be optimized for sample instructions using the model to be optimized, a sample instruction set may be constructed first. The sample instruction set may include various types of sample instructions, and each type of sample instruction may also have different forms of expression. Exemplarily, the sample instruction set may include information query instructions, such as "What is the highest mountain on Earth?", "What is the speed of light?", etc.; it may include content generation instructions, such as "Write a short article about the impact of artificial intelligence on society", "Write a poem about spring", etc.; it may also include explanatory instructions, such as "Why is the sky blue?", "Describe the discovery process and significance of gravitational waves", etc.; it may also include dialogue instructions, such as "What's your opinion on today's weather?", "Pretend to be an ancient philosopher and answer my questions", etc. It should be noted that the above examples of sample instructions are only for auxiliary explanation as a possible implementation manner in the embodiments of the present disclosure, and do not constitute an improper limitation of the present invention. The embodiments of the present disclosure do not specifically limit the types and specific contents of the sample instructions. In practical applications, the sample instruction set may also include guiding instructions, analytical reasoning instructions, translation instructions, emotional expression instructions, etc.

[0049] In the embodiments of the present disclosure, scoring the responses of the model to be optimized for sample instructions using the model to be optimized includes: inputting the sample instructions into the model to be optimized, and obtaining multiple response results of the model to be optimized for the sample instructions for multiple times; inputting the multiple response results and the sample instructions into the model to be optimized, so that the model to be optimized outputs scores for the multiple response results. After constructing the sample instruction set, it can be input into the model to be optimized. For each sample instruction in the sample instruction set, the steps described below can be performed to achieve the purpose of finally optimizing the model to be optimized. Only one sample instruction is used as an example for illustration below.

[0050] In specific implementation, the sample instructions can be input into the model to be optimized, so that the model to be optimized answers the sample instructions multiple times and outputs multiple response results. As Figure 2 shown, after inputting the sample instructions into the model to be optimized, k response results for sample instruction i can be generated . Then, the multiple response results and the sample instructions can be input into the model to be optimized again, so that the model to be optimized scores each response result, and k scoring results for the k response results are obtained It should be noted that the number of times the model to be optimized answers a sample instruction and the number of answer results obtained can be preset according to the application scenario and requirements. That is, k in the above text can be preset, and the embodiments of the present disclosure do not make specific limitations on this.

[0051] Exemplarily, assume that the sample instruction is "Describe the overfitting phenomenon in machine learning and give examples". It is preset that the model to be optimized needs to answer 3 times for the sample instruction and output 3 answer results. After inputting the above sample instruction into the model to be optimized, 3 answer results can be obtained, which are: Answer result 1: "Overfitting means that the model performs well on the training data but poorly on the test data. This is usually because the model is too complex and remembers the details in the training data without learning the general patterns. For example, if training a model to predict housing prices and there is a house in the training data with a much higher price than others, the model may over-rely on this data point and thus make mistakes in the test data"; Answer result 2: "Overfitting refers to the situation where a machine learning model learns the noise or unimportant details in the data during training, resulting in a decline in performance on new data. For example, when training an image classifier, if it remembers the background noise instead of the shape of the object, overfitting will occur"; and Answer result 3: "Overfitting is a phenomenon where the model performs excellently on the training set but poorly on new data. This often occurs when the amount of data is small or the model is too complex. For example, a neural network may completely remember the labels of each sample in the training set but perform poorly on new samples". Then, the above sample instruction and the 3 generated answer results can be re-input into the model to be optimized to let it score the quality of each answer. Here, the criteria for the model to be optimized to score the multiple generated answer results can be set according to the actual situation and requirements. For example, it can be scored based on dimensions such as semantic consistency, logic, and accuracy. The embodiments of the present disclosure do not make specific limitations on this, as long as its function can be realized.

[0052] In another possible implementation, in addition to re-inputting the multiple answer results of the model to be optimized for a sample instruction and the sample instruction into the model to be optimized for scoring, the multiple answer results for the sample instruction and the sample instruction can also be input into other already mature models, so that these relatively mature models score the multiple answer results of the model to be optimized. It should be noted that the above methods for scoring the multiple answer results of the model to be optimized for a sample instruction are only examples of feasible implementation manners in the embodiments of the present disclosure and do not constitute improper limitations on the present invention. In actual applications, it can be set according to actual needs and situations. The embodiments of the present disclosure do not make specific limitations on this, as long as its function can be realized.

[0053] In the embodiments of the present disclosure, as described above, when constructing a sample instruction set, a corresponding classification threshold may be set for each sample instruction to determine the category of the sample instruction; among them, different types of sample instructions may correspond to different classification thresholds, and different sample instructions of the same type may also correspond to different classification thresholds. The embodiments of the present disclosure do not make specific limitations in this regard. Specifically, determining the category of a sample instruction includes: comparing the scores of multiple answer results with the classification threshold corresponding to the sample instruction; where the classification threshold corresponding to the sample instruction is preset, and different sample instructions correspond to different classification thresholds; in response to the existence of an answer result with a corresponding score less than the classification threshold among the multiple answer results, the category of the sample instruction is determined as the target category.

[0054] In a specific implementation, assume that the classification threshold of the current sample instruction is 8 points, that is, if there is an answer result with a score exceeding 8 points among the multiple answer results for the sample instruction, then the category of the current sample instruction is a non-target category; if there is an answer result less than 8 points among the multiple answer results for the sample instruction, then the category of the current sample instruction is the target category. The purpose of this step is to select, based on the scores of the multiple answer results for the sample instruction, the sample instructions that the model to be optimized currently cannot accurately answer and process. If there are still unsatisfactory answer results among the multiple answer results of the model to be optimized for a certain sample instruction, it indicates that the model to be optimized cannot handle such sample instructions well at the current stage. Therefore, it is necessary to optimize and train the model to be optimized based on such sample instructions. If the multiple answer results of the model to be optimized for a certain sample instruction all exceed the corresponding classification threshold, it indicates that the model to be optimized can handle such problems well at the current stage. Therefore, it can be regarded as a non-target category. The steps for optimizing and training the model to be optimized for non-target category sample instructions are described in detail in S103 and will not be elaborated here.

[0055] Exemplarily, continuing with the previous example, assume that the scores of the model to be optimized for answer results 1 to 3 are 8.5, 7.0, and 9.0 respectively, and the classification threshold of the sample instruction is 8 points. Then, there is an answer result 2 whose score is less than the classification threshold corresponding to the sample instruction. At this time, the category of the sample instruction can be determined as the target category. Alternatively, the determination criterion for the target category can also be set as follows: if there is no answer result with a score greater than the classification threshold among the scores of multiple answer results, then the category of the current sample instruction is determined as the target category. Under this determination criterion, the category of the sample instruction corresponding to the above example can be determined as the non-target category; assume that the classification threshold is 9.5 points, and there is no answer result with a score greater than 9.5 points among the current 3 answer results, then the type of the corresponding sample instruction can be determined as the target type. The determination criterion of the present disclosure embodiment for whether the sample instruction is the target type is not specifically limited, and can be set according to actual needs in practical applications, as long as its function can be achieved.

[0056] S102: In response to the category of the sample instruction being the target category, determine the seed instruction corresponding to the sample instruction.

[0057] In the embodiments of the present disclosure, determining the seed instruction corresponding to the sample instruction includes: converting the sample instruction into a vector representation; based on the vector representation of the sample instruction, determining the seed instruction corresponding to the sample instruction in the pre-constructed instruction library; the similarity between the seed instruction and the sample instruction exceeds a preset similarity threshold.

[0058] In specific implementation, since the sample instruction is usually in the form of natural language text, a pre-trained sentence encoding model can be used to encode the sample instruction to obtain a fixed-length vector representation. The pre-constructed instruction library can contain a large number of encoded instruction sets, and each instruction is also represented as a vector. After obtaining the vector representation corresponding to the sample instruction, based on cosine similarity or Euclidean distance, etc., several instructions closest to the vector representation corresponding to the sample instruction in the instruction library can be found as the seed instructions. The number of seed instructions determined from the instruction library in the embodiments of the present disclosure is not specifically limited, and can be set according to actual needs in practical applications.

[0059] Another possible implementation is to expand or refine the sample instruction to directly generate multiple seed instructions related to the sample instruction. For example, if the sample instruction is "Classify objects as red, green, or blue"; then seed instructions such as "Classify items by color", "Perform color recognition on objects in an image", "Group objects according to color characteristics", etc. can be generated. The embodiments of the present disclosure do not specifically limit the method for determining the seed instructions corresponding to the sample instruction. The above two methods are only used as examples of possible implementations. In actual applications, other methods can also be used to determine the seed instructions corresponding to the sample instruction according to the actual situation, as long as its function can be achieved.

[0060] In the embodiments of the present disclosure, each seed instruction is usually a natural language description, which can provide rich context information and templates for the subsequently generated enhanced instructions and provide inspiration for generating enhanced instructions. After obtaining the seed instructions, an enhancement model can be used to generate enhanced instructions corresponding to the sample instruction. In a specific implementation, the sample instruction can be used as the input of the enhancement model to clarify the target direction for generating enhanced instructions. Multiple seed instructions can be sorted according to semantic similarity to ensure that the seed instructions more relevant to the sample instruction are prioritized. Then, the goal of generating enhanced instructions can be clarified, such as "Generate instructions related to the sample instruction but clearer according to the following seed instructions". The enhancement model can be used to analyze the semantic relationship between the input sample instruction and the seed instructions, weight the seed instructions through the self-attention mechanism, and extract the content most inspiring to the sample instruction to ensure that the generated enhanced instructions are both related to the sample instruction and diverse. Specifically, the enhancement model can generate more refined instructions from the general descriptions extracted from the seed instructions. For example, "Classify objects as warm or cool colors" can be expanded to "Classify red objects into different shades". In addition, the semantic clarity or format of sample instructions with unclear language structures can be adjusted. For example, "Classify as red, green, or blue" can be optimized to "Divide the objects in the image into three categories: red, green, and blue"; the same sample instruction can also be described in different ways to cover multiple possible understanding angles of the sample instruction. For example, for the sample instruction "Classify objects as red, green, or blue", it can be expanded to "Identify and group objects by color (red, green, blue)", "Label each object as a red, green, or blue category", "Classify and count objects of the three colors red, green, and blue", etc. In addition, when generating enhanced instructions corresponding to the sample instruction, a control signal can also be introduced to control the style and domain of the enhanced instructions. By generating enhanced instructions corresponding to the sample instruction, the diversity and pertinence of the instructions used for model optimization can be improved.

[0061] It should be noted that the above content is only an example of a feasible implementation manner for generating an enhanced instruction corresponding to a sample instruction in the embodiments of the present disclosure, and does not constitute an improper limitation of the present invention. In actual applications, any method can be selected according to actual needs and circumstances to generate an enhanced instruction corresponding to the sample instruction. The embodiments of the present disclosure do not make specific limitations on this, and it is subject to being able to achieve its function.

[0062] In the embodiments of the present disclosure, after obtaining the enhanced instruction corresponding to the sample instruction, the enhanced instruction can be used to optimize and train the model to be optimized. Specifically, optimizing the model to be optimized based on the enhanced instruction includes: inputting the enhanced instruction into the model to be optimized so that the model to be optimized outputs multiple answer results for answering the enhanced instruction multiple times; inputting the enhanced instruction and the multiple answer results corresponding to the enhanced instruction into the model to be optimized so that the model to be optimized scores the multiple answer results corresponding to the enhanced instruction; in response to the existence of an answer result whose score does not exceed the preset threshold among the multiple answer results corresponding to the enhanced instruction, generating a new enhanced instruction, and repeating the above steps until the scores of all the multiple answer results corresponding to the enhanced instruction exceed the preset threshold, and obtaining the optimized model to be optimized.

[0063] In a specific implementation, after obtaining the enhanced instruction corresponding to the sample instruction, the enhanced instruction can be input into the model to be optimized again, so that the model to be optimized answers each enhanced instruction multiple times and outputs multiple answer results. Then, the steps of scoring the multiple answer results described above can be repeated. The method of scoring the multiple answer results is similar to the method described above and will not be elaborated here. Then, the answer results can be judged according to the preset threshold corresponding to different enhanced instructions. The judgment method is the same as the method of judging the category of the sample instruction described above and will not be elaborated here. The preset threshold corresponding to the enhanced instruction can follow the classification threshold of the sample instruction corresponding to the threshold, or can be reset. The embodiments of the present disclosure do not make specific limitations on this. If there are still answer results that do not meet the threshold conditions among the multiple answer results corresponding to the enhanced instruction, the corresponding enhanced instruction can be used as a new sample instruction to generate a new enhanced instruction to optimize and train the model to be optimized until the answer results output by the model to be optimized meet the threshold conditions.

[0064] In the embodiments of the present disclosure, the method further includes: in response to the category of the sample instruction being a non-target category, selecting the answer result with the highest answer score and the answer result with the lowest score from the multiple answer results corresponding to the sample instruction, and constructing a triple training group corresponding to the sample instruction based on the answer result with the highest score, the answer result with the lowest score, and the sample instruction; optimizing and training the model to be optimized based on the triple training group.

[0065] In a specific implementation, when classifying a sample instruction, if the classification result indicates that the sample instruction is a target category, the model to be optimized can be optimized and trained for the sample instruction according to the steps described above. If the classification result indicates that the sample instruction is a non-target category, the n lowest-scoring answer results and the n highest-scoring answer results can be selected from multiple answer results, and together with the sample instruction, they form a triple training set to optimize and train the model to be optimized for the sample instruction.

[0066] Exemplarily, assume that the sample instruction is "Describe the planetary composition of the solar system", and the corresponding multiple answer results may include "There are eight major planets in the solar system", "The solar system includes planets such as Mercury, Venus, and Earth", and "There are many stars in the solar system". After scoring these three answer results, it is indicated that "The solar system includes planets such as Mercury, Venus, and Earth" is the best answer, and "There are many stars in the solar system" is the worst answer. At this time, a triple training set can be formed based on the best answer, the worst answer, and the sample instruction. When optimizing and training the model to be optimized based on the triple training set, the training objective can be to maximize the similarity between the model answer and the best answer and to minimize the similarity between the model answer and the worst answer, so that the answer result output by the model to be optimized can be biased towards the best answer.

[0067] In another possible implementation manner, the non-target type of sample instruction can also be input into a mature model to obtain a template answer that meets the training requirements, and the model to be optimized can be optimized and trained based on this template answer and the sample instruction. It should be noted that the above content is only a feasible implementation manner for optimizing and training the model to be optimized based on the non-target type of sample instruction in the embodiments of the present disclosure, and does not constitute an improper limitation of the present invention. In actual applications, the model to be optimized can be optimized and trained based on the non-target type of sample instruction according to actual needs and situations. The embodiments of the present disclosure do not make specific limitations on this, as long as its functions can be realized.

[0068] According to the second aspect of the embodiments of the present disclosure, as Figure 3 shown, a model optimization device 300 is provided, including:

[0069] A scoring module 301, configured to score the answer of the model to be optimized for the sample instruction by using the model to be optimized, so as to determine the category of the sample instruction;

[0070] A determination module 302, configured to determine a seed instruction corresponding to the sample instruction in response to the category of the sample instruction being a target category;

[0071] An optimization module 303, configured to generate an enhanced instruction corresponding to the sample instruction based on the context information of the sample instruction and the seed instruction, so as to optimize the model to be optimized based on the enhanced instruction.

[0072] Optionally, the scoring module 301 is specifically configured to:

[0073] Input the sample instruction into the model to be optimized, and obtain multiple response results of the model to be optimized for multiple responses to the sample instruction;

[0074] Input the multiple response results and the sample instruction into the model to be optimized, so that the model to be optimized outputs a score for the multiple response results.

[0075] Optionally, the scoring module 301 is specifically configured to:

[0076] Compare the scores of the multiple response results with the classification threshold corresponding to the sample instruction; wherein, the classification threshold corresponding to the sample instruction is preset, and different sample instructions have different corresponding classification thresholds;

[0077] In response to the existence of a response result whose corresponding score in the multiple response results is less than the classification threshold, determine the category of the sample instruction as the target category.

[0078] Optionally, the determination module 302 is specifically configured to:

[0079] Convert the sample instruction into a vector representation;

[0080] Based on the vector representation of the sample instruction, determine the seed instruction corresponding to the sample instruction in the pre-constructed instruction library; the similarity between the seed instruction and the sample instruction exceeds the preset similarity threshold.

[0081] Optionally, the optimization module 303 is specifically configured to:

[0082] Input the enhanced instruction into the model to be optimized, so that the model to be optimized outputs multiple response results for multiple responses to the enhanced instruction;

[0083] Input the enhanced instruction and the multiple response results corresponding to the enhanced instruction into the model to be optimized, so that the model to be optimized scores the multiple response results corresponding to the enhanced instruction;

[0084] In response to the existence of a response result whose score in the multiple response results corresponding to the enhanced instruction does not exceed the preset threshold, generate a new enhanced instruction, and repeat the above steps until the scores of the multiple response results corresponding to the enhanced instruction all exceed the preset threshold, and obtain the model to be optimized that has completed optimization.

[0085] Optionally, the optimization module 303 is further configured to:

[0086] In response to the category of the sample instruction being a non-target category, select the response result with the highest response score and the response result with the lowest score from the multiple response results corresponding to the sample instruction, and construct a triple training set corresponding to the sample instruction based on the response result with the highest score, the response result with the lowest score, and the sample instruction;

[0087] Based on the triple training set, perform optimization training on the model to be optimized.

[0088] According to the third aspect of the embodiments of the present disclosure, there is provided an electronic device for model optimization, including: one or more processors; a storage device for storing one or more programs, which when executed by the one or more processors, cause the one or more processors to implement the method provided in the first aspect of the embodiments of the present invention.

[0089] According to the fourth aspect of the embodiments of the present disclosure, there is provided a computer-readable medium having a computer program stored thereon, and when the program is executed by a processor, the method provided in the first aspect of the embodiments of the present invention is implemented.

[0090] According to the fifth aspect of the embodiments of the present invention, there is provided a computer program product including a computer program, and when the computer program is executed by a processor, the method of any of the above embodiments is implemented.

[0091] Figure 4 An exemplary system architecture 400 to which the model optimization method or model optimization device according to the present disclosure can be applied is shown.

[0092] As Figure 4 shown, the system architecture 400 may include terminal devices 401, 402, 403, a network 404, and a server 405. The network 404 is used to provide a medium for a communication link between the terminal devices 401, 402, 403 and the server 405. The network 404 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.

[0093] Users may use the terminal devices 401, 402, 403 to interact with the server 405 through the network 404 to receive or send messages, etc. Various communication client applications may be installed on the terminal devices 401, 402, 403, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only as examples).

[0094] The terminal devices 401, 402, 403 may be various electronic devices having a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop portable computers, and desktop computers, etc.

[0095] The server 405 may be a server that provides various services, such as a backend management server (for example only) that provides support for shopping websites browsed by users using the terminal devices 401, 402, and 403. The backend management server may process the received model optimization request and feed back the processing result (for example only) to the terminal device.

[0096] It should be noted that the model optimization method provided in the embodiment of the present invention is generally executed by the server 405, and accordingly, the model optimization device is generally set in the server 405. The model optimization method provided in the embodiment of the present invention can also be executed by the terminal devices 401, 402, 403, and accordingly, the model optimization device can be set in the terminal devices 401, 402, 403.

[0097] It should be understood that Figure 4 The number of terminal devices, networks and servers in the embodiment is only for illustration. Any number of terminal devices, networks and servers may be provided according to implementation requirements.

[0098] Reference below Figure 5 , which shows a schematic diagram of the structure of a computer system 500 of a terminal device suitable for implementing an embodiment of the present invention. Figure 5 The terminal device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present invention.

[0099] like Figure 5 As shown, the computer system 500 includes a central processing unit (CPU) 501, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 502 or a program loaded from a storage part 508 into a random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of the system 500 are also stored. The CPU 501, the ROM 502, and the RAM 503 are connected to each other through a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.

[0100] The following components are connected to the I / O interface 505: an input section 506 including a keyboard, a mouse, etc.; an output section 507 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 508 including a hard disk, etc.; and a communication section 509 including a network interface card such as a LAN card, a modem, etc. The communication section 509 performs communication processing via a network such as the Internet. A drive 510 is also connected to the I / O interface 505 as needed. A removable medium 511, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 510 as needed, so that a computer program read therefrom is installed into the storage section 508 as needed.

[0101] In particular, according to the embodiments disclosed in the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present invention include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 509, and / or installed from the removable medium 511. When the computer program is executed by the central processing unit (CPU) 501, the above-mentioned functions defined in the system of the present invention are executed.

[0102] It should be noted that the computer-readable medium shown in the present invention can be a computer-readable signal medium or a computer-readable storage medium or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, the computer-readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or in combination with an instruction execution system, apparatus, or device. In the present invention, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries the computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and the computer-readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any suitable medium, including but not limited to: wireless, wire, optical fiber, RF, etc., or any suitable combination of the above.

[0103] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, as well as combinations of blocks in the block diagram or flowchart, may be implemented by a dedicated hardware-based system that performs the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.

[0104] The modules described in the embodiments of the present invention can be implemented in software or in hardware. The described modules can also be provided in a processor. For example, a processor includes a scoring module, a determination module, and an optimization module. In some cases, the names of these modules do not constitute a limitation on the module itself. For example, the scoring module can also be described as "a module that scores the response of the model to be optimized for a sample instruction using the model to be optimized to determine the category of the sample instruction".

[0105] As another aspect, the present invention also provides a computer-readable medium, which may be included in the device described in the above embodiments; or may exist separately and not be assembled into the device. The above computer-readable medium carries one or more programs. When the one or more programs are executed by the device, the device implements the following method: scoring the response of the model to be optimized for a sample instruction using the model to be optimized to determine the category of the sample instruction; in response to the category of the sample instruction being the target category, determining a seed instruction corresponding to the sample instruction; generating an enhanced instruction corresponding to the sample instruction based on the context information of the sample instruction and the seed instruction, so as to optimize the model to be optimized based on the enhanced instruction.

[0106] Finally, it should be noted that the above embodiments are only specific implementation manners of the present disclosure, which are used to illustrate the technical solutions of the present disclosure, rather than limiting them. The protection scope of the present disclosure is not limited thereto. Although the present disclosure has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that any person skilled in the art within the technical scope disclosed by the present disclosure can still modify the technical solutions recorded in the foregoing embodiments, or can easily think of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present disclosure, and should all be covered by the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure shall be subject to the protection scope of the claims.

Claims

1. A model optimization method, characterized in that: include: Using the model to be optimized, scoring the answer of the model to be optimized to the sample instruction to determine the category of the sample instruction; In response to the category of the sample instruction being a target category, determining a seed instruction corresponding to the sample instruction; Based on the context information of the sample instruction and the seed instruction, an enhanced instruction corresponding to the sample instruction is generated to optimize the model to be optimized based on the enhanced instruction.

2. The method according to claim 1, characterized in that Scoring the answer of the model to be optimized to the sample instruction by using the model to be optimized, including: Inputting the sample instruction into the model to be optimized, and obtaining multiple answer results of the model to be optimized for the sample instruction; The plurality of answer results and the sample instructions are input into the model to be optimized, so that the model to be optimized outputs scores for the plurality of answer results.

3. The method according to claim 2, characterized in that Determine the category of the sample instruction, including: Comparing the scores of the plurality of answer results with the classification threshold corresponding to the sample instruction; wherein the classification threshold corresponding to the sample instruction is pre-set, and different sample instructions correspond to different classification thresholds; In response to the presence of an answer result with a corresponding score less than the classification threshold among the plurality of answer results, the category of the sample instruction is determined as a target category.

4. The method according to claim 1, characterized in that: Determining a seed instruction corresponding to the sample instruction includes: Converting the sample instruction into a vector representation; Based on the vector representation of the sample instruction, a seed instruction corresponding to the sample instruction is determined in a pre-built instruction library; and the similarity between the seed instruction and the sample instruction exceeds a preset similarity threshold.

5. The method according to claim 1, characterized in that Optimizing the model to be optimized based on the enhanced instruction includes: Inputting the enhancement instruction into the model to be optimized, so that the model to be optimized outputs multiple answer results for the enhancement instruction; Inputting the enhancement instruction and a plurality of answer results corresponding to the enhancement instruction into the model to be optimized, so that the model to be optimized scores the plurality of answer results corresponding to the enhancement instruction; In response to the presence of an answer result whose score does not exceed a preset threshold among the multiple answer results corresponding to the enhancement instruction, a new enhancement instruction is generated, and the above steps are repeated until the scores of the multiple answer results corresponding to the enhancement instruction all exceed the preset threshold, thereby obtaining a model to be optimized that has been optimized.

6. The method according to claim 2, characterized in that The method further comprises: In response to the category of the sample instruction being a non-target category, selecting an answer result with a highest answer score and an answer result with a lowest answer score from a plurality of answer results corresponding to the sample instruction, and constructing a ternary training group corresponding to the sample instruction based on the answer result with the highest answer score, the answer result with the lowest answer score and the sample instruction; Based on the ternary training group, the model to be optimized is optimized and trained.

7. A model optimization device, characterized in that: include: A scoring module, used for scoring the answer of the model to be optimized to the sample instruction by using the model to be optimized, so as to determine the category of the sample instruction; a determination module, configured to determine a seed instruction corresponding to the sample instruction in response to the category of the sample instruction being a target category; The optimization module is used to generate an enhanced instruction corresponding to the sample instruction based on context information of the sample instruction and the seed instruction, so as to optimize the model to be optimized based on the enhanced instruction.

8. An electronic device, characterized in that: include: one or more processors; a storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 6.

9. A computer readable medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.

10. A computer program product, comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 6.