Method for evaluating capability of large language model, method for aligning large language model, related device and computer program product

By evaluating the ability of large language models and adjusting the model using alignment strategies, the quality improvement problem in the training process of large language models is solved, and more efficient model alignment and overall training effect is achieved.

CN120579577AActive Publication Date: 2025-09-02BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
CN202510736388.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-04
Publication Date
2025-09-02
Estimated Expiration
2045-06-04

AI Technical Summary

Technical Problem

How to better complete the overall training process of large language models, especially to improve the performance and quality of the model in the collaborative work of pre-training and post-training stages.

Method used

By evaluating the ability method of the large language model, the target ability evaluation value is generated, the initial model is adjusted using the alignment strategy, and the high-quality intermediate model is selected to finally obtain the target large language model.

Benefits of technology

A more comprehensive and efficient model capability evaluation is achieved, and the alignment quality and overall training effect of large language models are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120579577A_ABST
    Figure CN120579577A_ABST
Patent Text Reader

Abstract

The invention provides a method for evaluating the ability of a large language model, a method for aligning the large language model, a related device and a computer program product, and relates to the technical field of artificial intelligence such as large language model alignment, model ability evaluation and deep learning. A specific embodiment of the method for evaluating the ability of the large language model comprises the steps of processing a sample question by using a to-be-evaluated large language model to obtain at least two to-be-evaluated answers; determining a correct answer set from the at least two to-be-evaluated answers by using sample answers corresponding to the sample questions; in response to at least two correct answers included in the correct answer set, generating a first ability evaluation value based on a similarity comparison result between the correct answers, and generating a second ability evaluation value based on a quantitative relationship between the correct answers in the correct answer set and the to-be-evaluated answers; and based on the first capability evaluation value and the second capability evaluation value, generating a target capability evaluation value for evaluating the model capability of the to-be-evaluated large language model. Therefore, the model capability of the large language model can be evaluated more comprehensively, qualitatively and efficiently.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technology, specifically to the field of artificial intelligence technologies such as large language model alignment, model capability evaluation, and deep learning, and especially to methods, devices, electronic devices, computer-readable storage media, and computer program products for evaluating the capabilities of large language models and aligning large language models. Background Art

[0002] Large Language Models (LLMs) are deep learning-based AI models that are trained on large amounts of text data to understand, generate, and reason about natural language. Their primary goal is to generate reasonable, grammatically correct output given a given input.

[0003] Because large language models are often massive and have a wide range of applications, they are often trained in two phases: pre-training and post-training. These two phases each have distinct goals and tasks, but working together enables the model to excel in a variety of natural language processing tasks.

[0004] In this context, how to better complete the overall training process of large language models is worthy of attention and an urgent need. Summary of the Invention

[0005] The embodiments of the present disclosure provide a method, apparatus, electronic device, computer-readable storage medium, and computer program product for evaluating the capabilities of a large language model and aligning a large language model.

[0006] In a first aspect, an embodiment of the present disclosure proposes a method for evaluating the capability of a large language model, comprising: processing a sample question using a large language model to be evaluated to obtain at least two answers to be evaluated; determining a correct answer set from at least two answers to be evaluated using sample answers corresponding to the sample question; in response to the correct answer set including at least two correct answers, generating a first capability evaluation value based on a similarity comparison result between the correct answers, and generating a second capability evaluation value based on a quantitative relationship between the correct answers and the answers to be evaluated in the correct answer set; generating a target capability evaluation value for evaluating the model capability of the large language model to be evaluated based on the first capability evaluation value and the second capability evaluation value.

[0007] In the second aspect, an embodiment of the present disclosure proposes a device for evaluating the capability of a large language model, comprising: an answer generation unit to be evaluated, configured to process a sample question using the large language model to be evaluated to obtain at least two answers to be evaluated; a correct answer determination unit, configured to determine a correct answer set from at least two answers to be evaluated using sample answers corresponding to the sample question; a sub-ability evaluation value generation unit, configured to generate a first capability evaluation value based on a similarity comparison result between the correct answers in response to the correct answer set including at least two correct answers, and to generate a second capability evaluation value based on a quantitative relationship between the correct answers and the answers to be evaluated in the correct answer set; and a total capability evaluation value generation unit, which generates a target capability evaluation value for evaluating the model capability of the large language model to be evaluated based on the first capability evaluation value and the second capability evaluation value.

[0008] In a third aspect, an embodiment of the present disclosure proposes a method for aligning large language models, including: adjusting an initial large language model using a first alignment strategy to obtain at least two large language models to be evaluated; generating a first target capability evaluation value corresponding to the large language model to be evaluated, wherein the first target capability evaluation value is generated based on the method for evaluating the capability of a large language model in the first aspect above; selecting a first intermediate large language model from the at least two large language models to be evaluated based on a ranking result of the first target capability evaluation value; and adjusting the first intermediate large language model using a second alignment strategy to obtain a target large language model.

[0009] In a fourth aspect, an embodiment of the present disclosure proposes an apparatus for aligning large language models, comprising: a first model adjustment unit, configured to adjust an initial large language model using a first alignment strategy to obtain at least two large language models to be evaluated; a first model evaluation unit, configured to generate a first target capability evaluation value corresponding to the large language model to be evaluated, wherein the first target capability evaluation value is generated based on the apparatus for evaluating the capability of a large language model according to the third aspect; an intermediate model selection unit, configured to select a first intermediate large language model from the at least two large language models to be evaluated based on a ranking result of the first target capability evaluation value; and a second model adjustment unit, configured to adjust the first intermediate large language model using a second alignment strategy to obtain a target large language model.

[0010] In a fifth aspect, an embodiment of the present disclosure provides an electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that when the at least one processor executes, it is able to implement the method for evaluating the capability of a large language model as described in any implementation manner in the first aspect and / or implement the method for aligning a large language model as described in any implementation manner in the third aspect.

[0011] In a sixth aspect, an embodiment of the present disclosure provides a non-transitory computer-readable storage medium storing computer instructions, which are used to enable a computer to implement the method for evaluating the capabilities of a large language model as described in any implementation manner in the first aspect and / or the method for aligning a large language model as described in any implementation manner in the third aspect when executed.

[0012] In a seventh aspect, an embodiment of the present disclosure provides a computer program product comprising a computer program, which, when executed by a processor, can implement the method for evaluating the capabilities of a large language model as described in any implementation manner in the first aspect and / or implement the method for aligning a large language model as described in any implementation manner in the third aspect.

[0013] The methods, devices, electronic devices, computer-readable storage media, and computer program products for evaluating the capabilities of a large language model provided by the embodiments of the present disclosure first process a sample question using the large language model to be evaluated to obtain at least two answers to be evaluated; then, a set of correct answers is determined from the at least two answers to be evaluated using the sample answers corresponding to the sample question; next, if the correct answer set includes at least two correct answers, a first capability evaluation value is generated based on a similarity comparison result between the correct answers, and a second capability evaluation value is generated based on a quantitative relationship between the correct answers and the answers to be evaluated in the correct answer set; finally, a target capability evaluation value for evaluating the model capability of the large language model to be evaluated is generated based on the first capability evaluation value and the second capability evaluation value.

[0014] Therefore, the present disclosure can evaluate the model capabilities of large language models in a more comprehensive, high-quality and efficient manner.

[0015] The methods, apparatuses, electronic devices, computer-readable storage media, and computer program products for aligning large language models provided by the embodiments of the present disclosure first adjust an initial large language model using a first alignment strategy to obtain at least two large language models to be evaluated. Then, the method for evaluating the capabilities of large language models described above is used to generate a first target capability evaluation value corresponding to each of the large language models to be evaluated. Next, based on the ranking results of the first target capability evaluation values, a first intermediate large language model is selected from the at least two large language models to be evaluated. Finally, the first intermediate large language model is adjusted using a second alignment strategy to obtain a target large language model.

[0016] Therefore, the present disclosure can achieve higher-quality alignment of large language models by combining at least two alignment strategies, thereby improving alignment quality. Furthermore, in this process, to further improve the alignment effect of each alignment strategy, the above-mentioned more comprehensive and efficient evaluation method can be used to generate multiple candidates and select the best one to determine the large language model that will enter the next alignment stage and alignment strategy. This effectively improves the alignment quality of large language models.

[0017] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Other features, objects and advantages of the present disclosure will become more apparent from a reading of the detailed description of non-limiting embodiments made with reference to the following drawings: Figure 1 is an exemplary system architecture in which the present disclosure may be applied; Figure 2 A flowchart of a process for evaluating the capabilities of a large language model provided in an embodiment of the present disclosure; Figure 3 A flowchart of another process for evaluating the capabilities of a large language model provided in an embodiment of the present disclosure; Figure 4 A flowchart of another process for evaluating the capabilities of a large language model provided in an embodiment of the present disclosure; Figure 5 A flowchart of a process for aligning a large language model provided in an embodiment of the present disclosure; Figure 6 A flowchart illustrating two processes, including evaluating the capabilities of a large language model and aligning a large language model, in an application scenario provided by an embodiment of the present disclosure; Figure 7 A structural block diagram of a device for evaluating the capabilities of a large language model provided by an embodiment of the present disclosure; Figure 8 A structural block diagram of a device for aligning a large language model provided by an embodiment of the present disclosure; Figure 9 A schematic structural diagram of an electronic device suitable for executing a method for evaluating the capability of a large language model and / or a method for aligning a large language model, provided in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0019] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those skilled in the art that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description. It should be noted that the embodiments in the present disclosure and the features in the embodiments can be combined with each other unless there is a conflict.

[0020] In addition, in the technical solutions involved in the present disclosure, the acquisition, storage, use, processing, transportation, provision and disclosure of the user personal information involved (for example, in some possible scenarios, the "big language model" in the present disclosure may be used to process the user's personal data. Accordingly, in order to train a "big language model" that meets such capabilities, the sample questions and sample answers corresponding to the sample questions involved in the present disclosure may also include at least part of the personal data) shall comply with the provisions of relevant laws and regulations and shall not violate public order and good morals.

[0021] Figure 1 An exemplary system architecture 100 is shown to which embodiments of the method, apparatus, electronic device, and computer-readable storage medium for evaluating the capability of a large language model and aligning a large language model of the present disclosure can be applied.

[0022] like Figure 1 As shown, system architecture 100 may include terminal devices 101, 102, 103, a network 104, and a server 105. Network 104 is a medium for providing communication links between terminal devices 101, 102, 103 and server 105. Network 104 may include various connection types, such as wired or wireless communication links or fiber optic cables.

[0023] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Various applications for enabling information communication between the terminal devices 101, 102, and 103 and server 105 can be installed, such as large language model capability assessment applications, large language model training applications, and instant messaging applications.

[0024] Terminal devices 101, 102, 103 and server 105 can be either hardware or software. When terminal devices 101, 102, 103 are hardware, they can be various electronic devices with display screens, including but not limited to smartphones, tablet computers, laptop computers, and desktop computers. When terminal devices 101, 102, 103 are software, they can be installed in the electronic devices listed above. They can be implemented as multiple software or software modules, or as a single software or software module, and are not specifically limited here. When server 105 is hardware, it can be implemented as a distributed server cluster consisting of multiple servers, or as a single server. When the server is software, it can be implemented as multiple software or software modules, or as a single software or software module, and are not specifically limited here.

[0025] The server 105 can provide various services through various built-in applications. Taking a large language model capability evaluation application that can provide an evaluation of the current model capability of a large language model as an example, the server 105 can achieve the following effects when running the large language model capability evaluation application: first, the server 105 uses the large language model to be evaluated to process the sample question and obtain at least two answers to be evaluated; then, the server 105 uses the sample answers corresponding to the sample question to determine a correct answer set from the at least two answers to be evaluated; next, if the correct answer set includes at least two correct answers, the server 105 responds to this by generating a first capability evaluation value based on the similarity comparison result between the correct answers, and generates a second capability evaluation value based on the quantitative relationship between the correct answers and the answers to be evaluated in the correct answer set; finally, the server 105 generates a target capability evaluation value for evaluating the model capability of the large language model to be evaluated based on the first capability evaluation value and the second capability evaluation value.

[0026] Similarly, when the above-mentioned application is, for example, a large language model training application that can provide training and alignment of a large language model, the server 105 can achieve the following effects when running the large language model training application: the server 105 uses the first alignment strategy to adjust the initial large language model to obtain at least two large language models to be evaluated; then, the server 105 generates a first target capability evaluation value corresponding to the large language model to be evaluated, wherein the first target capability evaluation value can be generated based on the above-mentioned process of evaluating the capability of the large language model; next, the server 105 selects a first intermediate large language model from the at least two large language models to be evaluated based on the ranking result of the first target capability evaluation value; finally, the server 105 uses the second alignment strategy to adjust the first intermediate large language model to obtain the target large language model.

[0027] It should be noted that the large language model, sample questions, sample answers, and the like can all be obtained by the server 105 from other devices, such as the terminal devices 101, 102, and 103, via the network 104. However, in addition to being obtained from the terminal devices 101, 102, and 103 via the network 104, they can also be pre-stored locally on the server 105 in various ways. Therefore, when the server 105 detects that this data is already stored locally (for example, when it begins processing a previously saved evaluation task for a large language model to be evaluated), it can choose to obtain this data directly from the local storage. In this case, the exemplary system architecture 100 may also not include the terminal devices 101, 102, and 103 and the network 104.

[0028] Since evaluating and aligning a large language model may require a large number of computing resources and strong computing power, the methods for evaluating and aligning a large language model provided in the subsequent embodiments of the present disclosure are generally performed by a server 105 having strong computing power and more computing resources. Accordingly, the apparatus for evaluating and aligning a large language model is generally also provided in the server 105. However, it should also be pointed out that when the terminal devices 101, 102, and 103 also have computing power and computing resources that meet the requirements, the terminal devices 101, 102, and 103 can also complete the various calculations originally assigned to the server 105 through the large language model capability evaluation application and the large language model training application installed thereon, and then output the same results as the server 105. Especially when there are multiple terminal devices with different computing capabilities, but the large language model capability assessment application and the large language model training application determine that the terminal device has stronger computing capabilities and more remaining computing resources, the terminal device can be allowed to perform the above-mentioned operations, thereby appropriately reducing the computing pressure on the server 105. Accordingly, the device for assessing the capability of the large language model and aligning the large language model can also be set in the terminal devices 101, 102, and 103. In this case, the exemplary system architecture 100 may also not include the server 105 and the network 104.

[0029] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is merely illustrative. Any number of terminal devices, networks and servers may be provided as required.

[0030] Next, we will first discuss the process of evaluating the capabilities of large language models.

[0031] Please refer to Figure 2 , Figure 2 A flowchart of a process for evaluating the capability of a large language model provided by an embodiment of the present disclosure includes process 200 .

[0032] The process 200 specifically includes the following steps: Step 201: Process the sample question using the large language model to be evaluated to obtain at least two answers to be evaluated; In the embodiment of the present disclosure, this step can be performed by an execution subject of the method for evaluating the capability of a large language model (e.g. Figure 1 The server 105 shown uses the large language model to be evaluated to process the sample question (for example, calls the large language model to be evaluated to process the sample question) to obtain at least two answers to be evaluated.

[0033] Typically, the large language model to be evaluated can be a partially trained large language model that has a certain processing capability (e.g., the ability to process sample questions). For example, the large language model to be evaluated can be a pre-trained large language model that is expected to be post-trained.

[0034] The sample questions may correspond to the desired capabilities of the large language model to be evaluated, such as those developed after the aforementioned pre-training. For example, if the pre-training is intended to train the large language model to process text (e.g., to extract key information from a text, such as a summary, or semantics, or to identify editing or writing errors within the text), the sample questions may include the corresponding text and a "prompt" for processing the text accordingly.

[0035] For example, in the case of expecting to read key information from a text, sample questions may be “What is said in the text?”, “What key information is included in the text?”, and so on.

[0036] For example, when the purpose of pre-training is to train a large language model to have corresponding answering and execution capabilities for input text, images, audio, etc., the sample question can actually be text, images, audio information that specifically includes the question to be answered, or instructions to perform corresponding actions.

[0037] For example, a sample question may be “What is XX?”. Accordingly, by using this sample question, the large language model to be evaluated may be instructed to find explanation information about “XX”.

[0038] Accordingly, in the embodiments of the present disclosure, when processing sample questions using the large language model to be evaluated, different prompt words, such as "generate at least two different language forms" or "generate at least two different versions of answers", can be used to indicate that after the large language model to be evaluated processes the sample question, two or more answers to be evaluated are generated or obtained. For example, for a sample question such as "What is XX?", the "explanatory information" generated by the large language model to be evaluated as the answer to be evaluated can be "XX is A1", "XX is A2", "XX is A3", etc.

[0039] It should be noted that, as discussed above, in some scenarios, even if the large language model to be evaluated is deployed locally on the execution subject, the sample problem can be obtained directly from the local storage device by the execution subject, or from a non-local storage device (such as Figure 1 The sample questions are obtained from the terminal devices 101, 102, and 103 shown in the figure (thereby enriching the source of "sample questions" so that the capabilities of the large language model to be evaluated can be evaluated more appropriately in the scenario). The local storage device can be a data storage module provided in the above-mentioned execution entity, such as a server hard disk. In this case, the sample questions can be quickly read locally; the non-local storage device can also be any other electronic device configured to store data, such as some user terminals. In this case, the above-mentioned execution entity can obtain the required sample questions by sending a obtain command to the electronic device.

[0040] Step 202: using the sample answers corresponding to the sample questions, determine a correct answer set from at least two answers to be evaluated; In an embodiment of the present disclosure, after obtaining at least two answers to be evaluated based on step 201 above, the execution entity may compare them with sample answers corresponding to the sample questions to determine a correct answer set from the at least two answers to be evaluated.

[0041] The sample answer can be a pre-determined, standard answer that corresponds to the sample question, is used to answer the sample question, and satisfies the requirements. For example, referring to the example discussed in step 201 above, if the sample question is "What is XX?", the sample answer can be the corresponding explanation information of "XX is YY." For example, such explanation information can be a widely recognized and accepted definition of "XX" (e.g., "YY").

[0042] Accordingly, in this step, the execution entity can use the "sample answer" to determine a "correct answer" that is correct and meets the requirements from at least two answers to be evaluated, thereby determining a correct answer set. For example, the execution entity can use textual similarity or semantic similarity to select answers to be evaluated that meet predetermined similarity requirements with the sample answer in terms of textual similarity or semantic similarity, and then determine these qualified answers to be evaluated as the correct answers to form the correct answer set.

[0043] For example, in the above example, where the "explanatory information" of the answer to be evaluated is "XX is A1", "XX is A2", and "XX is A3", the exemplary execution entity can determine the "correct answer" from these answers to be evaluated by comparing the semantic similarity of "A1", "A2", and "A3" with "YY" respectively.

[0044] Next, the execution entity may read and determine the number of elements included in the correct answer set (or more specifically, the number of “correct answers” ​​included).

[0045] If the number of elements included in the correct answer set is at least two, the execution entity may continue to execute step 203 in response thereto.

[0046] Step 203: In response to the correct answer set including at least two correct answers, generating a first ability evaluation value based on a similarity comparison result between the correct answers, and generating a second ability evaluation value based on a quantitative relationship between the correct answers in the correct answer set and the answers to be evaluated; In an embodiment of the present disclosure, in this step, if the execution entity includes at least two (i.e., two or more) "correct answers" in the correct answer set determined in the above step 202, the execution entity can respond thereto by generating a first ability evaluation value based on the similarity comparison result between the correct answers, and generating a second ability evaluation value based on the quantitative relationship between the correct answers in the correct answer set and the answers to be evaluated.

[0047] For the first ability evaluation value, for example, the execution entity can determine the text similarity between each "correct answer pair" based on a pairwise comparison between the correct answers. Then, the first ability evaluation value is generated by aggregating the text similarities, for example, taking an "average" approach.

[0048] For the second ability evaluation value, the execution entity may first determine the difference between the number of correct answers in the correct answer set and the number of answers to be evaluated (for example, the total number of answers to be evaluated minus the total number of correct answers). The execution entity then maps this "difference" to the corresponding second ability evaluation value based on a predetermined relationship between the difference and the ability evaluation value.

[0049] In some optional implementations of this embodiment, the execution entity may also choose to use the edit distance between correct answers as the "similarity comparison result" for the first ability evaluation value actually associated with "similarity." That is, the execution entity may use a correct answer as a benchmark and determine the edit distance between it and another correct answer in the same "correct answer pair" (for example, the number of operations required to adjust it to the other correct answer from the perspective of characters or words).

[0050] The execution entity can then use the edit distance as the similarity comparison result to determine the first ability evaluation value (for example, similarly, based on a predetermined correspondence between the edit distance and the first ability evaluation value, the edit distance is associated with the first ability evaluation value). In this way, the "edit distance" of semantic and glyph differences is combined to more comprehensively reflect the similarity between correct answers.

[0051] In practice, the correlation between the level of "similarity" and "similarity" and the level of the first ability evaluation value can be set accordingly based on different needs. For example, because the answers to be evaluated for the first ability evaluation value are all "correct answers", in order to enable the large language model to have stronger divergent ability and answer-finding ability, the level of the first ability evaluation value can be inversely proportional to the "similarity" and "similarity" in the above-mentioned similarity comparison results. That is, the lower the similarity between the correct answers output, the better the diversity of the "correct answers" generated by the large language model to be evaluated (or, in other words, it has a higher "answer diversity"), and it can correspondingly have a higher first ability evaluation value.

[0052] In some optional implementations of this embodiment, for the aforementioned "second ability evaluation value," the executing entity may choose to determine the quantitative relationship between the correct answers and the answers to be evaluated in the correct answer set based on the ratio of the number of correct answers to the number of answers to be evaluated. Furthermore, based on this quantitative relationship, the second ability evaluation value is generated.

[0053] Therefore, the execution entity can determine its performance in providing “correct answers” ​​by the proportion of the answers to be evaluated output by the large language model to be evaluated that are correct answers.

[0054] Accordingly, the second ability evaluation value is generally positively correlated with the ratio. Thus, the second ability evaluation value can be used to intuitively and directly reflect the correct answer output capability of the large language model (to be evaluated).

[0055] Step 204: Based on the first capability evaluation value and the second capability evaluation value, a target capability evaluation value for evaluating the model capability of the large language model to be evaluated is generated.

[0056] In an embodiment of the present disclosure, after the first capability evaluation value and the second capability evaluation value are determined and generated based on the above step 203, the execution entity generates a target capability evaluation value for evaluating the model capability of the large language model to be evaluated based on the first capability evaluation value and the second capability evaluation value.

[0057] For example, the execution entity may directly add the first capability evaluation value and the second capability evaluation value to obtain the target capability evaluation value. Alternatively, the execution entity may combine the first capability evaluation value and the second capability evaluation value based on a pre-configured reference coefficient, such as by weighted addition, to generate a target capability evaluation value for evaluating the model capability of the large language model to be evaluated.

[0058] Correspondingly, through this target capability evaluation value, the "overall capability" of the large language model to be evaluated in the dimension that needs to be evaluated can be intuitively fed back. For example, for an evaluated large language model with a higher target capability evaluation value, it can be understood as having better and higher answer diversity, as well as a higher "correct answer ratio".

[0059] The method for evaluating the capability of a large language model provided by an embodiment of the present disclosure first processes a sample question using the large language model to be evaluated to obtain at least two answers to be evaluated; then, using the sample answers corresponding to the sample questions, a correct answer set is determined from the at least two answers to be evaluated; next, if the correct answer set includes at least two correct answers, a first capability evaluation value is generated based on a similarity comparison result between the correct answers, and a second capability evaluation value is generated based on a quantitative relationship between the correct answers and the answers to be evaluated in the correct answer set; finally, based on the first capability evaluation value and the second capability evaluation value, a target capability evaluation value for evaluating the model capability of the large language model to be evaluated is generated. Thus, the present disclosure can evaluate the model capability of the large language model in a more comprehensive, high-quality, and efficient manner.

[0060] In some optional implementations of this embodiment, if in the above step 202 the execution subject only "believes" that there is only one correct answer among the at least two answers to be evaluated, that is, the correct answer set only includes a single correct answer.

[0061] In such cases, the execution entity can respond by selecting and utilizing only the "second capability evaluation value" to generate the aforementioned "target capability evaluation value," without further evaluating the similarity from, for example, a "diversity" perspective. Specifically, if the correct answer set includes a single correct answer, the execution entity can respond by generating the second capability evaluation value based solely on the quantitative relationship between the correct answer in the correct answer set and the answer to be evaluated.

[0062] Then, the execution entity generates a target capability evaluation value for evaluating the model capability of the large language model to be evaluated based only on the second capability evaluation value, so as to avoid operation errors caused by the lack of a "correct answer pair" for generating the first capability evaluation value.

[0063] In some embodiments, the executing entity may not determine at least one "correct answer" from at least two answers to be evaluated in the above-mentioned step 202 (that is, the executing entity "believes" that all the answers to be evaluated are wrong). In such a case, the executing entity may generate a prompt message for this situation, and then communicate the prompt message to the target device (for example, the terminal device used by the training party that provides training for the large language model to be evaluated) through a pre-configured communication path, so that the situation that the large language model to be evaluated may not currently generate the "correct answer" can be fed back to the "training party" for reference, so as to timely adjust the large language model to be evaluated (for example, adjust the "pre-training" environment).

[0064] In some embodiments, in order to improve the evaluation quality, the execution entity may also determine the target capability evaluation sub-value corresponding to each "processing round" in a manner similar to repeating and looping the above process, at least in a plurality of "processing rounds".

[0065] Then, the final target capability evaluation value is obtained by integrating these target capability evaluation sub-values, thereby avoiding inaccurate evaluation caused by, for example, random conditions in individual processing rounds.

[0066] For this, please refer to Figure 3 , Figure 3 A flowchart of another process for evaluating the capability of a large language model provided by an embodiment of the present disclosure includes process 300 .

[0067] The process 300 specifically includes the following steps: Step 301: Repeatedly process the sample question using the large language model to be evaluated to obtain at least two first answers to be evaluated corresponding to each processing round; Specifically, in this step, as discussed above, the execution entity can first determine at least two first answers to be evaluated corresponding to each processing round in a "processing round" manner (for ease of understanding, the "answers to be evaluated" for the processing round are described as "first answers to be evaluated").

[0068] For example, the execution entity may repeatedly provide sample questions to the large language model to be evaluated in processing rounds, so that the large language model to be evaluated completes the processing corresponding to the "processing round" and obtains at least two first answers to be evaluated corresponding to each processing round.

[0069] Step 302: In each processing round, using the sample answers corresponding to the sample questions, determine a first correct answer set corresponding to the processing round from at least two first answers to be evaluated; Specifically, in this step, the execution entity, in each processing round, is similar to the content discussed in step 202 above. In each processing round, the execution entity determines the (first) correct answer among the at least two first answers to be evaluated corresponding to the processing round to determine the first correct answer set corresponding to the processing round.

[0070] Step 303: In each processing round, in response to the corresponding first correct answer set including at least two first correct answers, a first ability evaluation sub-value corresponding to the processing round is generated based on the similarity comparison result between the first correct answers, and a second ability evaluation sub-value corresponding to the processing round is generated based on the quantitative relationship between the first correct answers and the first answer to be evaluated corresponding to the processing round; Specifically, in this step, the execution entity may, in each processing round, similarly to the content discussed in step 203 above, if the corresponding first correct answer set in the processing round includes at least two first correct answers, then the execution entity shall correspondingly generate the first ability evaluation sub-value corresponding to the processing round (i.e., determined based on the similarity comparison result between the first correct answers) and the second ability evaluation sub-value (i.e., determined based on the quantitative relationship between the first correct answer and the first answer to be evaluated corresponding to the processing round) using the process discussed in step 203 above.

[0071] Accordingly, if the first correct answer set includes only one first correct answer, the execution entity can similarly determine only the corresponding second capability evaluation sub-value, and in subsequent steps (for example, step 304), only use the second capability evaluation sub-value as the target capability evaluation sub-value of the corresponding processing round, which will not be repeated here.

[0072] Step 304: generating target capability evaluation sub-values ​​corresponding to each processing round based on the first capability evaluation sub-value and the second capability evaluation sub-value corresponding to the processing round; Specifically, after determining the first capability evaluation sub-value and the second capability evaluation sub-value corresponding to each processing round based on the above step 303, the execution entity can be similar to the discussion in the above step 204. In each processing round, the execution entity obtains the target capability evaluation sub-value corresponding to the processing round by combining the first capability evaluation sub-value and the second capability evaluation sub-value.

[0073] Step 305: Based on each target capability evaluation sub-value, generate a target capability evaluation value for the large language model to be evaluated.

[0074] Specifically, after obtaining the target capability evaluation sub-values ​​corresponding to each processing round based on the above step 304, the execution entity can similarly generate the target capability evaluation value for the large language model to be evaluated by, for example, adding all the target capability evaluation sub-values.

[0075] Therefore, the execution entity can reduce the impact of the randomness of the model itself on the effect evaluation through repeated evaluation based on multiple processing rounds.

[0076] In some embodiments, not only can the quality be improved through multiple "processing rounds", but the execution entity can also further choose to use different sample questions in different processing rounds to utilize more sample questions to more accurately test the large language model to be evaluated, so as to more accurately evaluate the ability.

[0077] For this, please refer to Figure 4 , Figure 4 A flowchart of another process for evaluating the capability of a large language model provided in an embodiment of the present disclosure includes process 400 .

[0078] The process 400 specifically includes the following steps: Step 401: Process at least two sample questions respectively using the large language model to be evaluated to obtain at least two second answers to be evaluated corresponding to each sample question; Specifically, in this step, the execution entity can use the large language model to be evaluated to process at least two (different) sample questions respectively, and obtain at least two second answers to be evaluated corresponding to each sample question (for ease of understanding, the "answers to be evaluated" for the sample questions are described as "second answers to be evaluated").

[0079] Step 402: For each sample question, using the sample answer corresponding to the sample question, determine a second correct answer set corresponding to the sample question from at least two second answers to be evaluated; Specifically, in this step, the execution entity can determine the "(second) correct answer" among the at least two second answers to be evaluated corresponding to each sample answer, similar to the content discussed in step 202 above, to obtain a set of second correct answers corresponding to each sample question.

[0080] Step 403: For each sample question, in response to the corresponding second correct answer set including at least two second correct answers, based on the similarity comparison result between the second correct answers, generate a third ability evaluation sub-value corresponding to the sample question, and based on the quantitative relationship between the second correct answers and the second answer to be evaluated, generate a fourth ability evaluation sub-value corresponding to the sample question; Specifically, in this step, the execution entity can, similar to the content discussed in step 203 above, for each sample question, if the corresponding second correct answer set includes at least two second correct answers, then the execution entity will use these second correct answers to determine the third ability evaluation sub-value corresponding to the sample question (that is, determined based on the similarity comparison result between the second correct answers, and the difference from the above is provided in form only for the convenience of understanding) and the fourth ability evaluation sub-value (that is, determined based on the quantitative relationship between the second correct answer and the second answer to be evaluated corresponding to the processing round, and the difference from the above is provided in form only for the convenience of understanding).

[0081] Step 404: Generate a target ability evaluation sub-value corresponding to each sample question based on the third ability evaluation sub-value and the fourth ability evaluation sub-value corresponding to the sample question; Specifically, after determining the third ability evaluation sub-value and the fourth ability evaluation sub-value corresponding to each sample problem based on the above step 403, the execution entity can, similar to the discussion in the above step 204, obtain the target ability evaluation sub-value corresponding to each sample problem by combining the third ability evaluation sub-value and the fourth ability evaluation sub-value for each sample problem.

[0082] Step 405: Based on each target capability evaluation sub-value, generate a target capability evaluation value for the large language model to be evaluated.

[0083] Specifically, after obtaining the target capability evaluation sub-value corresponding to each sample question based on the above step 404, the execution entity can similarly generate the target capability evaluation value for the large language model to be evaluated by, for example, adding all the target capability evaluation sub-values.

[0084] In some embodiments, when the execution entity generates a target capability evaluation value for the large language model to be evaluated based on each target capability evaluation sub-value, for example, during step 305 or step 405, the execution entity may select at least one of the mean, variance, and range of each target capability evaluation sub-value to generate the target capability evaluation value for the large language model to be evaluated. Thus, through further computational processing, a target capability evaluation value for evaluating the model capability of the large language model to be evaluated can be more accurately and reasonably provided.

[0085] In some embodiments, after the execution entity determines the corresponding second ability evaluation sub-value (or, in different scenarios, the fourth ability evaluation sub-value, here only the text form of "second ability evaluation sub-value" is used as an example) for each sample question in each processing round, the execution entity may also choose to determine a total result based on the distribution results of each second ability evaluation sub-value to replace the part of the target ability evaluation value corresponding to each ability evaluation sub-value. In this way, the "correctness" of the answer to be evaluated provided by the large language model and its ability to provide correct answers can be reflected through the distribution of the "correctness rate" by the execution entity.

[0086] Next, embodiments of the present disclosure also provide a method for aligning a large language model. This process can be used to ensure that the output of a large language model complies with ethical, legal, and social norms during training and use. For example, the large language model can be trained (or post-trained) using manually labeled sample alignment inputs and sample alignment results to achieve alignment.

[0087] For this, please refer to Figure 5 . Figure 5 A flowchart of a process of aligning a large language model provided in an embodiment of the present disclosure includes process 500 .

[0088] The process 500 specifically includes the following steps: Step 501: Using a first alignment strategy to adjust the initial large language model to obtain at least two large language models to be evaluated; In the embodiment of the present disclosure, this step can be performed by the execution subject of the method for aligning the large language model (for example Figure 1 The server 105 shown in FIG. 1 adjusts the initial large language model based on the first alignment strategy. The initial large language model may be, for example, a pre-trained large language model that needs to be aligned.

[0089] The first alignment strategy can be one of the supervised fine-tuning alignment strategy (SFT), the direct preference optimization alignment strategy (DPO), and the proximal policy optimization alignment strategy (PPO).

[0090] SFT is the process of training a pre-trained large language model for a specific application or task by introducing labeled data. This allows the model to better adapt to specific applications or tasks. This process is typically used to further adjust the model parameters of a pre-trained large language model using supervised learning methods to achieve better performance in specific domains or tasks.

[0091] When aligning the initial large language model, DPO aims to directly optimize the output of the initial large language model to make it more consistent with human preferences.

[0092] PPO is also a behavioral strategy that can be used to train large language models, especially when dealing with large-scale, high-complexity tasks. It can reduce instability during training through more effective strategy optimization.

[0093] In this step, the execution entity can actually use the first alignment strategy to adjust the initial large language model to obtain at least two large language models to be evaluated. For example, even for the same first alignment adjustment strategy (e.g., DPO), the execution entity can use different sample alignment inputs and sample alignment results to train and adjust at least two different large language models to be evaluated based on the same initial large language model. Step 502: Generate a first target capability evaluation value corresponding to the large language model to be evaluated; In an embodiment of the present disclosure, based on step 501, the executing entity in this step may generate a first target capability evaluation value corresponding to each of the at least two large language models to be evaluated generated in the above step 501 (for the convenience of description, the target capability evaluation value corresponding to each of the large language models to be evaluated may be referred to as a first target capability evaluation value).

[0094] As for the "first target capability evaluation value", it can actually be generated by the executing entity based at least on the method and process for evaluating the capability of the large language model discussed above for process 200 (that is, the executing entity generates the "first target capability evaluation value" by the process of generating the "target capability evaluation value" discussed above), and will not be repeated here.

[0095] Step 503: Selecting a first intermediate large language model from at least two large language models to be evaluated based on the ranking result of the first target capability evaluation value; In an embodiment of the present disclosure, after generating each first target capability evaluation value based on step 502, the execution entity may sort each first target capability evaluation value in descending order, for example. The execution entity then selects the first intermediate large language model for subsequent "alignment" from the sorted results.

[0096] For example, after sorting in descending order, the execution entity may select the first-ranked large language model (i.e., the large language model to be evaluated with the highest first target capability evaluation value as the first intermediate large language model), or the large language model to be evaluated with a preset rank before sorting as the first intermediate large language model.

[0097] Step 504: Use the second alignment strategy to adjust the first intermediate large language model to obtain a target large language model.

[0098] In an embodiment of the present disclosure, after the execution entity selects the first intermediate large language model based on the above step 503, it can continue to adjust the first intermediate large language model in the next round or stage based on the second alignment strategy to obtain the target large language model.

[0099] In some optional implementations of this embodiment, the "second alignment strategy" may be another alignment strategy different from the first alignment strategy, for example, another alignment strategy among the above-mentioned SFT, DPO, and PPO that is different from the first alignment strategy.

[0100] Therefore, by combining the first alignment strategy and the second alignment strategy (for example, using two different alignment strategies in series), the execution entity can combine the advantages of each alignment strategy to better complete the alignment and post-training of the large language model.

[0101] In some optional implementations of this embodiment, when the execution entity uses the second alignment strategy for alignment, that is, in the process of executing this step, as an alternative or substitute, the execution entity may also similarly choose to simultaneously use the second alignment strategy to adjust the first intermediate large language model to obtain at least two second intermediate large language models.

[0102] Then, after determining and generating the second target ability evaluation values ​​corresponding to each second intermediate large language model (for example, the second target ability evaluation values ​​corresponding to the second intermediate large language model are generated in the same manner as discussed in the above process 200), a target large language model is selected from at least two second intermediate large language models based at least on the ranking results of the second target ability evaluation values.

[0103] This avoids the situation where the alignment quality is affected by, for example, occasional errors, deviations, etc., in a method of directly generating a target large language model as a result.

[0104] In some embodiments, in the process of selecting a target large language model from at least two second intermediate large language models based on at least the sorting result of the second target capability evaluation value, the execution entity may also, as an alternative or alternative, choose to select a target large language model from the first intermediate large language model and at least two second intermediate large language models based on the joint sorting result of the first target capability evaluation value and the second target capability evaluation value.

[0105] Specifically, the execution entity can use the "first intermediate large language model" before adjustment using the second alignment strategy, along with each "second intermediate large language model," as candidates for the target large language model. These models can then be compared with at least two second intermediate large language models obtained after adjustment using the second alignment strategy to determine the optimal large language model. This prevents undesirable "alignment" of large language models due to inappropriate second alignment strategies, which could negatively impact model quality.

[0106] The method for aligning large language models provided in the embodiments of the present disclosure first adjusts an initial large language model using a first alignment strategy to obtain at least two large language models to be evaluated. Then, the method for evaluating the capabilities of large language models described above is used to generate a first target capability evaluation value corresponding to each of the large language models to be evaluated. Next, based on the ranking results of the first target capability evaluation values, a first intermediate large language model is selected from the at least two large language models to be evaluated. Finally, the second alignment strategy is used to adjust the first intermediate large language model to obtain a target large language model.

[0107] Therefore, the present disclosure can achieve higher-quality alignment of large language models by combining at least two alignment strategies, thereby improving alignment quality. Furthermore, in this process, to further improve the alignment effect of each alignment strategy, the above-mentioned more comprehensive and efficient evaluation method can be used to generate multiple candidates and select the best one to determine the large language model that will enter the next alignment stage and alignment strategy. This effectively improves the alignment quality of large language models.

[0108] In some embodiments, after the first and second alignment strategies, third and fourth alignment strategies can be utilized, sequentially and in stages, to achieve alignment that meets different objectives and criteria. For example, after obtaining at least two second intermediate large language models, the execution entity can select the second intermediate large language model with the highest corresponding second target capability evaluation value and continue to align it using the third alignment strategy. This description will not be repeated here.

[0109] On the basis of any of the above embodiments, the "target large language model" obtained after alignment can be used accordingly to process and complete corresponding tasks. For example, as discussed above, if the target large language model is trained to have text processing capabilities (for example, the ability to read key information from the text) after pre-training (and post-training, alignment), then in such a case, the "target large language model" can, after obtaining the text to be processed, perform text processing actions accordingly based on prompts such as "Please extract the key information in the text", or based on pre-configured processing purposes to obtain corresponding results. For example, in this scenario, the target large language model can extract the "key information" such as synopsis and summary included in the "text to be processed" after obtaining the text to be processed.

[0110] To deepen understanding, this disclosure also combines a specific application scenario and provides a specific implementation solution including the two processes of evaluating the large language model capabilities and aligning the large language model. Figure 6 . Figure 6 A flowchart illustrating two processes, namely, evaluating the capabilities of a large language model and aligning the large language model, is provided for an embodiment of the present disclosure in an application scenario. Figure 6 Process 600 is included.

[0111] In process 600 , for the sake of convenience, the server 105 (not shown in the figure) may be used as the complete “executing entity” of the two processes.

[0112] First, in the process 600, the large language model 610 can be an initial "initial large language model" that needs to be subsequently "aligned".

[0113] Accordingly, for the large language model 610 , the execution entity may adjust the large language model 610 using the first alignment strategy by executing S601 to obtain large language models 621 , 622 to 62N to be evaluated (where N is a positive integer).

[0114] Then, the execution subject can use the sample question 630 and the sample answer 635 to respectively “evaluate” the large language models 621, 622 to 62N to be evaluated. For example, the execution subject can perform the evaluation based on at least the above Figure 2 As discussed in the illustrated process 200, for each large language model to be evaluated, its processing sample question 630 is called separately; then, the execution body determines the correct answer in the processing result of each large language model to be evaluated for the sample question 630 based on the sample answer 635; and then, based on the "correct answer (set)", the target capability evaluation value of the large language model to be evaluated is correspondingly determined.

[0115] For example, the result of the “evaluation” of the large language model to be evaluated 621 may be the first target capability evaluation value 641. Similarly, the result of the “evaluation” of the large language model to be evaluated 622 may be the first target capability evaluation value 642. The result of the “evaluation” of the large language model to be evaluated 62N may be the first target capability evaluation value 64N.

[0116] Then, the execution body executes S603 to select a first intermediate large language model from the large language models 621 , 622 to 62N to be evaluated.

[0117] For example, the execution entity may select the large language model to be evaluated with the highest first target capability evaluation value as the first intermediate large language model based on the descending sorting result of the first target capability evaluation values ​​641 , 642 to 64N.

[0118] For example, in this scenario, the first target capability evaluation value 642 corresponding to the large language model 622 to be evaluated is the highest. Then, after executing S603 , the execution entity may use the large language model 622 to be evaluated as the “first intermediate large language model”.

[0119] Next, the execution entity may continue to execute S604 to further “align” the large language model to be evaluated 622 using a second alignment strategy (different from the first alignment strategy) to obtain an aligned target large language model 650 .

[0120] Further references Figure 7 As an implementation of the above-mentioned method for evaluating the capability of a large language model, the present disclosure provides an embodiment of a device for evaluating the capability of a large language model. Figure 2 Corresponding to the method embodiment shown, the device can be specifically applied to various electronic devices.

[0121] like Figure 7As shown, the apparatus 700 for evaluating the capability of a large language model in this embodiment may include: an answer-to-be-evaluated generation unit 701, a correct answer determination unit 702, a sub-ability evaluation value generation unit 703, and a total capability evaluation value generation unit 704. The answer-to-be-evaluated generation unit 701 is configured to process a sample question using the large language model to be evaluated to obtain at least two answers to be evaluated; the correct answer determination unit 702 is configured to determine a correct answer set from the at least two answers to be evaluated using the sample answers corresponding to the sample question; the sub-ability evaluation value generation unit 703 is configured to generate a first capability evaluation value based on a similarity comparison result between the correct answers in response to the correct answer set including at least two correct answers, and to generate a second capability evaluation value based on a quantitative relationship between the correct answers in the correct answer set and the answers to be evaluated; and the total capability evaluation value generation unit 704 generates a target capability evaluation value for evaluating the model capability of the large language model to be evaluated based on the first capability evaluation value and the second capability evaluation value.

[0122] In this embodiment, the specific processing and technical effects of the following in the apparatus 700 for aligning a large language model: the to-be-evaluated answer generating unit 701, the correct answer determining unit 702, the sub-ability evaluation value generating unit 703 and the total ability evaluation value generating unit 704 can be referred to respectively. Figure 2 The relevant descriptions of steps 201-204 in the corresponding embodiment are not repeated here.

[0123] In some optional implementations of this embodiment, the apparatus 700 further includes: a similarity comparison unit configured to determine a similarity comparison result between the correct answers based on the edit distances between the correct answers.

[0124] In some optional implementations of this embodiment, the apparatus 700 further includes: a quantitative relationship determination unit configured to determine the quantitative relationship between the correct answers and the answers to be evaluated in the correct answer set based on the ratio of the number of correct answers to the number of answers to be evaluated.

[0125] In some optional implementations of this embodiment, the answer generation unit 701 is further configured to repeatedly process the sample question using the large language model to be evaluated to obtain at least two first answers to be evaluated corresponding to each processing round; and the correct answer determination unit 702 is further configured to, in each processing round, respectively use the sample answers corresponding to the sample question to determine the first correct answer set corresponding to the processing round from the at least two first answers to be evaluated; and the ability evaluation value generation unit 703 is further configured to, in each processing round, respond to the corresponding first correct answer. The set includes at least two first correct answers, and based on the similarity comparison result between the first correct answers, a first ability evaluation sub-value corresponding to the processing round is generated, and based on the quantitative relationship between the first correct answer and the first answer to be evaluated corresponding to the processing round, a second ability evaluation sub-value corresponding to the processing round is generated; the total ability evaluation value generation unit 704 is further configured to generate a target ability evaluation sub-value corresponding to each processing round based on the first ability evaluation sub-value and the second ability evaluation sub-value corresponding to the processing round; based on each target ability evaluation sub-value, a target ability evaluation value for the large language model to be evaluated is generated.

[0126] In some optional implementations of this embodiment, the answer-to-be-evaluated generation unit 701 is further configured to process at least two sample questions respectively using the large language model to be evaluated to obtain at least two second answers to be evaluated corresponding to each sample question; and the correct answer determination unit 702 is further configured to, for each sample question, use the sample answer corresponding to the sample question to determine a second correct answer set corresponding to the sample question from the at least two second answers to be evaluated; and the sub-ability evaluation value generation unit 703 is further configured to, for each sample question, in response to the corresponding second correct answer set including at least two second correct answers, generate a third ability evaluation sub-value corresponding to the sample question based on a similarity comparison result between the second correct answers, and generate a fourth ability evaluation sub-value corresponding to the sample question based on a quantitative relationship between the second correct answer and the second answer to be evaluated; and the total ability evaluation value generation unit 704 is further configured to generate a target ability evaluation sub-value corresponding to each sample question based on the third ability evaluation sub-value and the fourth ability evaluation sub-value corresponding to the sample question; and generate a target ability evaluation value for the large language model to be evaluated based on each target ability evaluation sub-value.

[0127] In some optional implementations of this embodiment, a target capability evaluation value for the large language model to be evaluated is generated based on each target capability evaluation sub-value, including: generating a target capability evaluation value for the large language model to be evaluated based on at least one of the mean, variance, and range of each target capability evaluation sub-value.

[0128] In some optional implementations of this embodiment, the sub-ability evaluation value generation unit 703 is further configured to, in response to the correct answer set including a unique correct answer, generate a second ability evaluation value based on the quantitative relationship between the correct answers in the correct answer set and the answers to be evaluated; the total ability evaluation value generation unit 704 is further configured to generate a target ability evaluation value for evaluating the model ability of the large language model to be evaluated based on the second ability evaluation value.

[0129] This embodiment exists as an apparatus embodiment corresponding to the above-mentioned method embodiment. The apparatus for evaluating the capability of a large language model provided by this embodiment first processes a sample question using the large language model to be evaluated to obtain at least two answers to be evaluated; then, a set of correct answers is determined from the at least two answers to be evaluated using the sample answers corresponding to the sample questions; next, if the correct answer set includes at least two correct answers, a first capability evaluation value is generated based on the similarity comparison result between the correct answers, and a second capability evaluation value is generated based on the quantitative relationship between the correct answers and the answers to be evaluated in the correct answer set; finally, a target capability evaluation value for evaluating the model capability of the large language model to be evaluated is generated based on the first capability evaluation value and the second capability evaluation value. Thus, the present disclosure can evaluate the model capability of the large language model in a more comprehensive, high-quality and efficient manner.

[0130] Further references Figure 8 As an implementation of the above-mentioned method for aligning a large language model, the present disclosure provides an embodiment of a device for aligning a large language model. Figure 5 Corresponding to the method embodiment shown, the device can be specifically applied to various electronic devices.

[0131] like Figure 8As shown, the device 800 for aligning large language models in this embodiment may include: a first model adjustment unit 801, a first model evaluation unit 802, an intermediate model selection unit 803, and a second model adjustment unit 804. The first model adjustment unit 801 is configured to adjust the initial large language model using a first alignment strategy to obtain at least two large language models to be evaluated; the first model evaluation unit 802 is configured to generate a first target capability evaluation value corresponding to the large language model to be evaluated, wherein the first target capability evaluation value is generated based on the device for evaluating the capability of a large language model of claim 10; the intermediate model selection unit 803 is configured to select a first intermediate large language model from at least two large language models to be evaluated based on the ranking result of the first target capability evaluation value; the second model adjustment unit 804 is configured to adjust the first intermediate large language model using a second alignment strategy to obtain a target large language model. In this embodiment, in the device 800 for aligning large language models: the specific processing of the first model adjustment unit 801, the first model evaluation unit 802, the intermediate model selection unit 803, and the second model adjustment unit 804 and the technical effects brought about by them can be referred to respectively. Figure 5 The relevant descriptions of steps 501-504 in the corresponding embodiment are not repeated here.

[0132] In some optional implementations of this embodiment, the second model adjustment unit 804 is further configured to adjust the first intermediate large language model using a second alignment strategy to obtain at least two second intermediate large language models; generate a second target capability evaluation value corresponding to each second intermediate large language model; and select a target large language model from at least the at least two second intermediate large language models based at least on the ranking result of the second target capability evaluation value.

[0133] In some optional implementations of this embodiment, at least based on the sorting result of the second target ability evaluation value, the target large language model is selected from at least two second intermediate large language models, including: based on the common sorting result of the first target ability evaluation value and the second target ability evaluation value, the target large language model is selected from the first intermediate large language model and at least two second intermediate large language models.

[0134] In some optional implementations of this embodiment, the first alignment strategy is different from the second alignment strategy, and the alignment strategies include: a supervised fine-tuning alignment strategy, a direct preference optimization alignment strategy, and a proximal strategy optimization alignment strategy.

[0135] This embodiment exists as an apparatus embodiment corresponding to the above-mentioned method embodiment. The apparatus for aligning large language models provided in this embodiment first adjusts an initial large language model using a first alignment strategy to obtain at least two large language models to be evaluated. Then, the apparatus uses the above-mentioned method for evaluating the capabilities of large language models to generate a first target capability evaluation value corresponding to each of the (each) large language model to be evaluated. Next, based on the ranking results of the first target capability evaluation values, a first intermediate large language model is selected from the at least two large language models to be evaluated. Finally, the apparatus adjusts the first intermediate large language model using a second alignment strategy to obtain a target large language model.

[0136] Therefore, the present disclosure can achieve higher-quality alignment of large language models by combining at least two alignment strategies, thereby improving alignment quality. Furthermore, in this process, to further improve the alignment effect of each alignment strategy, the above-mentioned more comprehensive and efficient evaluation method can be used to generate multiple candidates and select the best one to determine the large language model that will enter the next alignment stage and alignment strategy. This effectively improves the alignment quality of large language models.

[0137] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0138] Figure 9 A schematic block diagram of an example electronic device 900 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are provided as examples only and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0139] like Figure 9 As shown, device 900 includes a computing unit 901, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 902 or a computer program loaded from a storage unit 908 into a random access memory (RAM) 903. RAM 903 may also store various programs and data required for the operation of device 900. Computing unit 901, ROM 902, and RAM 903 are connected to each other via a bus 904. An input / output (I / O) interface 905 is also connected to bus 904.

[0140] Various components in the device 900 are connected to the I / O interface 905, including an input unit 906, such as a keyboard, a mouse, etc.; an output unit 907, such as various types of displays, speakers, etc.; a storage unit 908, such as a magnetic disk, an optical disk, etc.; and a communication unit 909, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 909 allows the device 900 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0141] The computing unit 901 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 901 performs the various methods and processes described above, such as the methods for evaluating the capabilities of a large language model and aligning a large language model. For example, in some embodiments, the methods for evaluating the capabilities of a large language model and aligning a large language model can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 908. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 900 via the ROM 902 and / or the communication unit 909. When the computer program is loaded into the RAM 903 and executed by the computing unit 901, one or more steps of the methods for evaluating the capabilities of a large language model and aligning a large language model described above can be performed. Alternatively, in other embodiments, the computing unit 901 may be configured in any other appropriate manner (for example, by means of firmware) to execute the method for evaluating the capability of a large language model and aligning a large language model.

[0142] Various embodiments of the systems and techniques described above can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0143] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0144] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fibers, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0145] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0146] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0147] A computer system may include clients and servers. The clients and servers are typically remote from each other and typically interact via a communication network. This client-server relationship arises through computer programs running on the respective computers, creating a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host. This server is a host product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosts and virtual private server (VPS) services. Servers can also be classified as distributed system servers or servers integrated with blockchain.

[0148] According to the technical solutions of the embodiments of the present disclosure, not only can the model capabilities of the large language model be evaluated more comprehensively, with higher quality and efficiency, but the large language model can also be aligned more efficiently by combining at least two alignment strategies to improve the alignment quality. Moreover, in such a process, in order to further improve the alignment effect of each alignment strategy, it is also possible to determine the large language model that enters the next alignment stage and alignment strategy by generating multiple candidates and selecting the best one based on the more high-quality, comprehensive and efficient evaluation method provided above. In this way, the alignment quality of the large language model can be effectively improved.

[0149] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions provided by this disclosure can be achieved. This is not limited herein.

[0150] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.

Claims

1. A method for evaluating the capabilities of a large language model, comprising: Use the large language model to process the sample question and obtain at least two answers to be evaluated; Determining a correct answer set from at least two answers to be evaluated using sample answers corresponding to the sample question; In response to the correct answer set including at least two correct answers, generating a first ability evaluation value based on a similarity comparison result between the correct answers, and generating a second ability evaluation value based on a quantitative relationship between the correct answers in the correct answer set and the answer to be evaluated; A target capability evaluation value for evaluating the model capability of the large language model to be evaluated is generated based on the first capability evaluation value and the second capability evaluation value.

2. The method according to claim 1, further comprising: Based on the edit distances between the correct answers, a similarity comparison result between the correct answers is determined.

3. The method according to claim 1, further comprising: Based on the ratio of the number of the correct answers to the number of the answers to be evaluated, a quantitative relationship between the correct answers and the answers to be evaluated in the correct answer set is determined.

4. The method according to claim 1, wherein The method of processing the sample question using the large language model to be evaluated to obtain at least two answers to be evaluated includes: Repeatedly processing the sample question using the large language model to be evaluated to obtain at least two first answers to be evaluated corresponding to each processing round; and The method of determining a correct answer set from at least two answers to be evaluated using the sample answers corresponding to the sample questions includes: In each of the processing rounds, using the sample answers corresponding to the sample questions, respectively, determining a first correct answer set corresponding to the processing round from at least two of the first answers to be evaluated; In response to the correct answer set including at least two correct answers, generating a first ability evaluation value based on a similarity comparison result between the correct answers, and generating a second ability evaluation value based on a quantitative relationship between the correct answers in the correct answer set and the answer to be evaluated, including: In each of the processing rounds, in response to the corresponding first correct answer set including at least two first correct answers, generating a first ability evaluation sub-value corresponding to the processing round based on a similarity comparison result between the first correct answers, and generating a second ability evaluation sub-value corresponding to the processing round based on a quantitative relationship between the first correct answers and the first answer to be evaluated corresponding to the processing round; And generating a target capability evaluation value for evaluating the model capability of the large language model to be evaluated based on the first capability evaluation value and the second capability evaluation value includes: generating target capability evaluation sub-values ​​corresponding to the respective processing rounds based on the first capability evaluation sub-values ​​and the second capability evaluation sub-values ​​corresponding to the processing rounds; Based on each of the target capability evaluation sub-values, a target capability evaluation value for the large language model to be evaluated is generated.

5. The method according to claim 1, wherein The method of processing the sample question using the large language model to be evaluated to obtain at least two answers to be evaluated includes: Processing at least two sample questions respectively using the large language model to be evaluated to obtain at least two second answers to be evaluated corresponding to each of the sample questions; and The method of determining a correct answer set from at least two answers to be evaluated using the sample answers corresponding to the sample questions includes: For each of the sample questions, using the sample answers corresponding to the sample questions, correspondingly determining a second correct answer set corresponding to the sample question from at least two second answers to be evaluated; and generating a capability evaluation value for evaluating the model capability of the large language model to be evaluated using the correct answer set, including: For each of the sample questions, in response to the corresponding second correct answer set including at least two second correct answers, generating a third ability evaluation sub-value corresponding to the sample question based on a similarity comparison result between the second correct answers, and generating a fourth ability evaluation sub-value corresponding to the sample question based on a quantitative relationship between the second correct answers and the second answer to be evaluated; And generating a target capability evaluation value for evaluating the model capability of the large language model to be evaluated based on the first capability evaluation value and the second capability evaluation value includes: generating a target ability evaluation sub-value corresponding to each of the sample questions based on the third ability evaluation sub-value and the fourth ability evaluation sub-value corresponding to the sample questions; Based on each of the target capability evaluation sub-values, a target capability evaluation value for the large language model to be evaluated is generated.

6. The method according to claim 4 or 5, wherein: Generating a target capability evaluation value for the large language model to be evaluated based on each of the target capability evaluation sub-values ​​includes: Based on at least one of the mean difference, variance, and range of each of the target capability evaluation sub-values, a target capability evaluation value for the large language model to be evaluated is generated.

7. The method according to claim 1, further comprising: In response to the correct answer set including a unique correct answer, generating a second ability evaluation value based on a quantitative relationship between the correct answer in the correct answer set and the answer to be evaluated; Based on the second capability evaluation value, a target capability evaluation value for evaluating the model capability of the large language model to be evaluated is generated.

8. A method for aligning a large language model, comprising: Using the first alignment strategy to adjust the initial large language model to obtain at least two large language models to be evaluated; generating a first target capability evaluation value corresponding to the large language model to be evaluated, wherein the first target capability evaluation value is generated based on the method for evaluating the capability of a large language model according to any one of claims 1 to 7; Selecting a first intermediate large language model from at least two of the large language models to be evaluated based on the ranking result of the first target capability evaluation values; The first intermediate large language model is adjusted using a second alignment strategy to obtain a target large language model.

9. The method according to claim 8, wherein The step of adjusting the intermediate large language model using the second alignment strategy to obtain a target large language model includes: adjusting the first intermediate large language model using a second alignment strategy to obtain at least two second intermediate large language models; generating a second target capability evaluation value corresponding to each of the second large intermediate language models; A target large language model is selected from at least two of the second intermediate large language models based at least on the ranking result of the second target capability evaluation values.

10. The method according to claim 9, wherein: The selecting a target large language model from at least two of the second intermediate large language models based on at least the ranking result of the second target capability evaluation value includes: Based on a common ranking result of the first target capability evaluation value and the second target capability evaluation value, a target large language model is selected from the first intermediate large language model and at least two of the second intermediate large language models.

11. The method according to any one of claims 8 to 10, wherein: The first alignment strategy is different from the second alignment strategy, and the alignment strategy includes: a supervised fine-tuning alignment strategy, a direct preference optimization alignment strategy, and a proximal strategy optimization alignment strategy.

12. A device for evaluating the capabilities of a large language model, comprising: an answer generation unit to be evaluated, configured to process the sample question using the large language model to be evaluated to obtain at least two answers to be evaluated; a correct answer determination unit configured to determine a correct answer set from at least two answers to be evaluated using the sample answers corresponding to the sample questions; a capability evaluation value generating unit configured to generate a first capability evaluation value based on a similarity comparison result between the correct answers in response to the correct answer set including at least two correct answers, and to generate a second capability evaluation value based on a quantitative relationship between the correct answers in the correct answer set and the answer to be evaluated; The total capability evaluation value generating unit generates a target capability evaluation value for evaluating the model capability of the large language model to be evaluated based on the first capability evaluation value and the second capability evaluation value.

13. A device for aligning a large language model, comprising: A first model adjustment unit is configured to adjust the initial large language model using a first alignment strategy to obtain at least two large language models to be evaluated; a first model evaluation unit configured to generate a first target capability evaluation value corresponding to the large language model to be evaluated, wherein the first target capability evaluation value is generated based on the apparatus for evaluating the capability of a large language model according to claim 10; an intermediate model selection unit, configured to select a first intermediate large language model from at least two of the large language models to be evaluated based on a ranking result of the first target capability evaluation value; The second model adjustment unit is configured to adjust the first intermediate large language model using a second alignment strategy to obtain a target large language model.

14. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method for evaluating the capability of a large language model according to any one of claims 1-7, and / or execute the method for aligning a large language model according to any one of claims 8-11.

15. A non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the method for evaluating the capability of a large language model according to any one of claims 1 to 7, and / or the method for aligning a large language model according to any one of claims 8 to 11.

16. A computer program product, comprising a computer program, which, when executed by a processor, implements the method for evaluating the capability of a large language model according to any one of claims 1 to 7, and / or performs the method for aligning a large language model according to any one of claims 8 to 11.

Citation Information

Patent Citations

  • Human-machine conversation method and device, terminal and server

    CN109543014A

  • Generated text length control method and device based on large language model

    CN117787241A

  • Large-model-oriented multi-dimensional question and answer pair generation task evaluation method and system

    CN118093344A

  • Text evaluation benchmark construction method and device

    CN118093831A

  • Large model fine tuning training system

    CN119443275A