Language model comparison method and device based on preference
Through semantic truncation processing and preference prediction model, the problems of position influence and data sparseness in the preference judgment of large language models are solved, and more accurate model comparison is achieved.
Patent Information
- Application Number
- CN202510612830.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-13
- Publication Date
- 2025-08-15
- Estimated Expiration
- 2045-05-13
AI Technical Summary
Existing large language models are easily affected by content position when making preference judgments, making it difficult to capture effective information in long texts, and the high training cost and sparse data lead to insufficient generalization, resulting in inaccurate preference judgments.
The truncated data is obtained through semantic truncation processing, the first and second data are generated, and the preference prediction model is used to calculate the preference probability of the first and second language models, reducing the positional impact and improving accuracy.
It effectively reduces the impact of content position on the prediction results, retains key information, and improves the accuracy and efficiency of model comparison.
Smart Images

Figure CN120493943A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a preference-based language model comparison method and device. Background Art
[0002] In the field of artificial intelligence, preference refers to a user's preference for model-generated content and is a crucial criterion for evaluating and comparing models. Manual evaluation is the gold standard for directly capturing human preferences, but it also presents challenges such as high cost and low efficiency. As the capabilities of large language models continue to grow, direct preference prediction based on large language models has emerged. This technology leverages a pre-trained large language model to simulate human judges and automatically determine preference for responses generated by different models.
[0003] However, with this approach, the large language model may not base its preference judgments on the quality of the answers themselves, but may instead be more inclined to select answers at fixed positions. Furthermore, due to the limitations of the large language model's capabilities, it struggles to capture effective information when processing long texts, leading to inaccurate preference judgments and difficulty in obtaining accurate model comparison results. Furthermore, training large language models requires a large amount of manually annotated preference data, which is costly and data-sparse. Furthermore, most of this data is in a single language, resulting in insufficient generalization. Summary of the Invention
[0004] In response to the above problems, the purpose of the present invention is to provide a preference-based language model comparison method and device, which can effectively reduce the impact of content location on prediction, retain the core information in the data to improve the accuracy of model comparison.
[0005] In order to solve the above technical problems, the present invention provides the following technical solutions:
[0006] In one aspect, the present invention provides a preference-based language model comparison method, comprising:
[0007] Acquire data to be processed, the data to be processed including a plurality of sub-data, the sub-data including a prompt to be processed, a first answer generated by a first language model based on the prompt to be processed, and a second answer generated by a second language model based on the prompt to be processed;
[0008] For each of the sub-data, based on the length of the sub-data, the length of the data to be processed, and the length limit of the data to be processed set by the preference prediction model, semantically truncate the sub-data to obtain truncated data, wherein the truncated data includes a truncation prompt, a first truncation answer, and a second truncation answer;
[0009] generating first data and second data based on answer difference data between the first answer and the second answer and the truncated data, wherein the first truncated answer and the second truncated answer have different positions in the first data than in the second data;
[0010] Using the preference prediction model to perform prediction processing on the first data and the second data, obtaining a first preference probability corresponding to the first language model and a second preference probability corresponding to the second language model;
[0011] A comparison result between the first language model and the second language model is determined based on the first preference probability and the second preference probability.
[0012] On the other hand, the present invention also provides a preference-based language model comparison device, comprising:
[0013] an acquisition module, configured to acquire data to be processed, the data to be processed comprising a plurality of sub-data, the sub-data comprising a prompt to be processed, a first answer generated by a first language model based on the prompt to be processed, and a second answer generated by a second language model based on the prompt to be processed;
[0014] a truncation module, configured to perform semantic truncation processing on each sub-data based on the length of the sub-data, the length of the data to be processed, and the length limit of the data to be processed set by the preference prediction model, to obtain truncated data, wherein the truncated data includes a truncation prompt, a first truncation answer, and a second truncation answer;
[0015] a generating module for generating first data and second data based on answer difference data between the first answer and the second answer and the truncated data, wherein the first truncated answer and the second truncated answer have different positions in the first data than in the second data;
[0016] A prediction module, configured to perform prediction processing on the first data and the second data using the preference prediction model to obtain a first preference probability corresponding to the first language model and a second preference probability corresponding to the second language model;
[0017] A comparison module is configured to determine a comparison result between the first language model and the second language model based on the first preference probability and the second preference probability.
[0018] On the other hand, the present invention also provides an electronic device comprising a processor and a memory, wherein the memory stores a plurality of instructions; the processor loads instructions from the memory to execute the steps of any one of the preference-based language model comparison methods provided by the present invention.
[0019] On the other hand, the present invention also provides a computer-readable storage medium, which stores a plurality of instructions, and the instructions are suitable for loading by a processor to execute the steps in any one of the preference-based language model comparison methods provided by the present invention.
[0020] On the other hand, the present invention also provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the steps in any one of the preference-based language model comparison methods provided by the present invention.
[0021] The beneficial effects brought about by the technical solution provided by the present invention include at least:
[0022] An embodiment of the present invention can obtain data to be processed containing multiple sub-data, and perform semantic truncation processing on the sub-data according to the length of the sub-data, the length of the data to be processed, and the limited length of the preference prediction model to be processed, to ensure that the truncated data carries a higher level of effective information, and then generate first data and second data based on the truncated data and the answer difference data. The introduction of the answer difference data can assist the model in accurately understanding the difference between the two answers, and the positions of the two answers in the first data are different from those in the second data, which can reduce the impact of the position of the answer on the result. After processing by the preference prediction model, the first preference probability of the first language model and the second preference probability of the second language model can be obtained; finally, the comparison result is directly obtained based on the first preference probability and the second preference probability, and the model comparison can be accurately realized. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.
[0024] Figure 1 Schematic diagram of an application scenario of the preference-based language model comparison method provided by an embodiment of the present invention;
[0025] Figure 2 is a flow chart of a preference-based language model comparison method provided by an embodiment of the present invention;
[0026] Figure 3 is a schematic diagram of semantic truncation processing provided by an embodiment of the present invention;
[0027] Figure 4 is a schematic diagram of generating first data and second data provided by an embodiment of the present invention;
[0028] Figure 5is a structural diagram of a preference-based language model comparison device provided by an embodiment of the present invention;
[0029] Figure 6 It is a structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0030] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making any creative efforts shall fall within the scope of protection of the present invention.
[0031] It is understandable that in the specific implementation of the present invention, data related to user information, etc., requires user permission or consent, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards of relevant countries and regions.
[0032] The embodiment of the present invention provides a preference-based language model comparison method, which can be found in Figure 1 , shows a schematic diagram of an application scenario for the preference-based language model comparison method. The application scenario may include a terminal 101 and a server 102, which can exchange data over a network. Terminal 101 may have a question-and-answer related application installed on it. Terminal 101 may be a mobile phone, tablet computer, smart Bluetooth device, computer, large screen, robot, or other device; server 102 may be a single server or a server cluster consisting of multiple servers.
[0033] The user can send the data to be processed to the server 102 through the terminal 101, wherein the data to be processed includes multiple sub-data, including a prompt to be processed, a first answer generated by the first language model based on the prompt to be processed, and a second answer generated by the second language model based on the prompt to be processed.
[0034] For each sub-data, server 102 performs semantic truncation processing on the sub-data based on the length of the sub-data, the length of the data to be processed, and the length limit of the data to be processed by the preference prediction model to obtain truncated data, where the truncated data includes a truncation prompt, a first truncated answer, and a second truncated answer; generates first data and second data based on the answer difference data between the first answer and the second answer and the truncated data, where the positions of the first truncated answer and the second truncated answer in the first data are different from those in the second data; uses the preference prediction model to perform prediction processing on the first data and the second data to obtain a first preference probability corresponding to the first language model and a second preference probability corresponding to the second language model; and determines a comparison result of the first language model and the second language model based on the first preference probability and the second preference probability.
[0035] Finally, the server 102 may send the comparison result to the terminal 101 so as to be displayed to the user via the terminal 101 .
[0036] In this embodiment, a preference-based language model comparison method is provided, such as Figure 2 As shown, the specific process of the preference-based language model comparison method can be as follows:
[0037] S110: Obtain data to be processed.
[0038] The data to be processed refers to the data that currently needs to be processed. The data to be processed may include multiple sub-data, and the sub-data may include a prompt to be processed, a first answer generated by the first language model based on the prompt to be processed, and a second answer generated by the second language model based on the prompt to be processed.
[0039] The pending prompt can be understood as a description of the task that the model needs to perform. The first language model and the second language model are the two models that need to be compared. These two models can be different versions of the same model or completely different models.
[0040] The first answer is the content output by the first language model after processing the pending prompt, and the second answer is the content output by the second language model after processing the pending prompt. The pending prompt, the first answer, and the second answer can all be referred to as sub-data. The pending data typically includes the pending prompt, the first answer, and the second answer.
[0041] The data to be processed may be a prompt to be processed inputted by the user into the first language model and the second language model respectively in advance to obtain the first answer and the second answer, and the user combines the prompt to be processed, the first answer and the second answer into the data to be processed.
[0042] Of course, the user can also directly provide the data to be processed and input the data to be processed into the first language model and the second language model respectively. The server directly receives the first answer output by the first language model and the second answer output by the second language model, and automatically combines the first answer, the second answer and the prompt to be processed into the data to be processed.
[0043] S120 , for each sub-data, based on the length of the sub-data, the length of the data to be processed, and the length limit of the data to be processed by the preference prediction model, perform semantic truncation processing on the sub-data to obtain truncated data.
[0044] The data to be processed contains multiple sub-data. For each sub-data, the length of each sub-data can be obtained. The length of the data to be processed is the sum of the lengths of all sub-data. In this embodiment of the present invention, the length of the sub-data and the length of the data to be processed can be measured by the number of characters. For example, if a sub-data contains 10 characters, the length of the sub-data is 10.
[0045] The prediction model is a model that compares the output effects of two models based on preference comparison. The prediction model can be pre-trained using a large amount of preference data. It is understood that the length of the model's input data is not unlimited. Typically, the length of the input data cannot exceed the maximum length limit of the model. In embodiments of the present invention, the input data of the prediction model includes data to be processed. A corresponding length limit can be set for the data to be processed. This length limit should be less than the maximum length limit of the prediction model. The specific setting can be based on actual needs or experience and is not specifically limited here.
[0046] If the length of the data to be processed exceeds the length limit, the data to be processed needs to be truncated. For each sub-data, semantic truncation is performed based on the length of the sub-data, the length of the data to be processed, and the length limit of the data to be processed set by the prediction model. This produces truncated data, which may include a truncation prompt, a first truncation answer, and a second truncation answer.
[0047] The truncated prompt is the data obtained by semantically truncating the prompt to be processed, the first truncated answer is the data obtained by semantically truncating the first answer, and the second truncated answer is the data obtained by semantically truncating the second answer. The truncated prompt, the first truncated answer, and the second truncated answer can constitute the truncated data.
[0048] As an implementation method, in order to avoid the loss of important content in each sub-data after truncation, the available length of each sub-data can be calculated, and the sub-data can be truncated in different ways according to the type of sub-data, with priority given to retaining the important content in the sub-data. Specifically, when performing semantic truncation processing, for each sub-data, the ratio of the length of the sub-data to the length of the data to be processed can be multiplied by the limit length to calculate the available length of the sub-data; if the length of the sub-data is greater than the available length of the sub-data and the sub-data is a prompt to be processed, the prompt to be processed is truncated based on the available length of the prompt to be processed and the preset rules to obtain a truncated prompt; if the length of the sub-data is greater than the available length of the sub-data and the sub-data is answer data, the available length of the answer data is used to determine a candidate clause from the answer data, and based on the candidate clause, a truncated answer corresponding to the answer data is generated, the answer data includes a first answer and a second answer, and the truncated answer includes a first truncated answer and a second truncated answer.
[0049] See Figure 3 , shows a schematic diagram of semantic truncation processing. Among them, for each sub-data, the length of the sub-data, the length of the data to be processed and the limited length can be used to calculate the available length of the sub-data. For example, the length ratio of the sub-data in the data to be processed can be calculated, that is, the length ratio of the sub-data is directly divided by the length of the data to be processed to obtain the length ratio, and then the length ratio is multiplied by the limited length to obtain the available length of the sub-data. For example, Figure 3 In the example, if the length of the pending prompt is s1, the available length of the pending prompt is s2; the length of the first answer is s3, then the available length of the first answer is s4; the length of the second answer is s5, then the available length of the second answer is s6.
[0050] The sub-data may include a pending prompt, a first answer, and a second answer. In one embodiment, if the sub-data is a pending prompt and the length of the pending prompt is greater than its corresponding available length, i.e., s1 > s2, data of the available length may be directly extracted from the pending prompt as the truncated prompt word.
[0051] As another implementation, since the pending prompt usually includes the task goal, in order to avoid ambiguity of the task goal, the pending prompt can be truncated based on the available length of the pending prompt and preset rules to obtain a truncated prompt.
[0052] Optionally, when truncating the pending prompt to obtain a truncated prompt, the available length of the pending prompt can be divided into a first length and a second length, and the sum of the first length and the second length is the available length; the content of the first length before the pending prompt is used as the first truncated content; the content of the second length after the pending prompt is used as the second truncated content; the first truncated content and the second truncated content are spliced to obtain a truncated prompt.
[0053] The beginning of the prompt to be processed can usually include the background of the problem, and the end of the prompt to be processed can usually include the concluding requirements. These two parts of content can be retained first. Specifically, the available length s2 of the prompt to be processed can be divided into a first length s21 and a second length s22. The first length s1 can be considered as the length of the content that needs to be retained in the beginning part to be processed, and the second length s2 can be considered as the length of the content that needs to be retained in the end part to be processed. The sum of the first length and the second length is the available length of the prompt to be processed. The first length and the second length can be set according to actual needs. In an embodiment of the present invention, the first length and the second length can be equal.
[0054] The first length of content before the prompt to be processed is used as the first truncated content, the second length of content after the prompt to be processed is used as the second truncated content, and the first truncated content and the second truncated content are then spliced to obtain the truncated prompt.
[0055] If the sub-data is answer data and the answer data exceeds its corresponding available length, the answer data needs to be truncated. When truncating the answer data, the available length of the answer data can be used to determine candidate clauses from the answer data, and the candidate clauses are used to generate a truncated answer corresponding to the answer data. If the answer data includes a first answer and a second answer, the truncated answer includes the corresponding first truncated answer and second truncated answer.
[0056] Optionally, when semantic truncation processing is performed on the answer data to obtain a truncated answer, the answer data may be split into multiple clauses, wherein each clause has a clause number representing the order of the clause in the answer data; the semantic similarity between each clause and the prompt to be processed is calculated; based on the semantic similarity and the available length of the answer data, candidate clauses are determined from the multiple clauses; and the candidate clauses are sorted according to the clause number to obtain a truncated answer.
[0057] The answer data can be split into a series of clauses according to punctuation marks, and a corresponding clause number can be assigned to each clause. The clause number can be used to represent the order of the clauses in the answer data. For example, the clause number corresponding to the first clause in the answer data can be 1, the clause number corresponding to the second clause can be 2, and so on. For example, the answer data is: To reduce carbon emissions, the following measures are recommended: 1. Promote renewable energy (such as solar and wind energy) to reduce dependence on fossil fuels; 2. Implement a carbon tax policy to curb high-carbon emission behaviors through economic leverage; 3. Develop carbon capture technology to offset industrial emissions. In addition, it is necessary to strengthen international cooperation and formulate unified emission reduction standards. The split clauses can be Clause 1: "Promote renewable energy (such as solar energy and wind energy) and reduce dependence on fossil fuels" (ID=1); Clause 2: "Implement carbon tax policies and curb high-carbon emission behaviors through economic leverage" (ID=2); Clause 3: "Develop carbon capture technology to offset industrial emissions" (ID=3); Clause 4: "Strengthen international cooperation and formulate unified emission reduction standards" (ID=4), where ID is the clause number.
[0058] In order to prioritize clauses with a high semantic relevance to the prompt to be processed, the semantic similarity between each clause and the prompt to be processed may be calculated. Then, based on the semantic similarity and the available length of the answer data, candidate clauses may be determined from the multiple clauses. Alternatively, all clauses may be sorted in descending order of semantic similarity to obtain a clause sequence. An estimated length may be obtained by adding the current length of all candidate clauses in the candidate list to the length of a designated clause, where the designated clause is the first clause currently in the clause sequence. If the estimated length is not greater than the available length of the answer data, the designated clause is added as a candidate clause to the candidate list and deleted from the clause sequence until the estimated length exceeds the available length of the answer data, thereby obtaining a candidate clause in the candidate list.
[0059] It should be noted that the higher the semantic similarity, the higher the correlation between the clause and the prompt to be processed. Therefore, the clauses can be directly arranged in order from high to low according to the semantic similarity to obtain a clause sequence. A candidate list is pre-set to store candidate clauses. When determining the candidate clauses, the clause ranked first in the current clause sequence can be used as the designated clause, and then the sum of the length of all candidate clauses in the current candidate list and the length of the designated clause is calculated as the estimated length; if the estimated length is not greater than the available length of the answer data, it can be considered that the answer data will not exceed its available length after adding the designated clause, and the designated clause can be directly moved from the clause sequence to the candidate list, that is, it can be used as a candidate clause; the above process is repeated until the estimated length is greater than the available length of the answer data, indicating that adding the designated clause will exceed the available length. At this time, all candidate clauses in the candidate list can be directly obtained.
[0060] After determining the candidate clauses, they can be sorted based on their clause numbers to restore the natural word order and logical coherence of the text. For example, the candidate clauses can be arranged in ascending order of clause numbers and then concatenated to form a first truncated answer and a second truncated answer that retain the core information while meeting the length limit.
[0061] The above-mentioned truncation method for pending prompts and answer data can ensure that the data retains contextual information and semantically relevant content that are crucial for preference judgment while meeting the length limit, providing higher-quality input for subsequent preference prediction models, improving the accuracy of predictions, and thus improving the accuracy of model comparison.
[0062] S130 : Generate first data and second data based on the answer difference data between the first answer and the second answer and the truncated data.
[0063] The first and second answers are responses to the same prompt from different models. These two answers often differ in some way. To help the model more discerningly capture these differences, we introduce the difference data between the first and second answers as auxiliary input. This difference data can include a similarity difference and a difference summary.
[0064] The answer difference data and the truncated data are concatenated, and the positions of the first and second truncated answers in the concatenated data are adjusted to obtain the first and second data. In other words, the first and second data are essentially the same; only the positions of the first and second truncated answers differ, minimizing the impact of answer position on prediction results.
[0065] Optionally, when generating the first data and the second data based on the answer difference data and the truncated data, the first similarity can be obtained by calculating the average similarity between all the candidate clauses of the first answer and the prompt to be processed; the second similarity can be obtained by calculating the average similarity between all the candidate clauses of the second answer and the prompt to be processed; the difference points between the first answer and the second answer can be identified to obtain a difference summary; the first similarity, the second similarity and the difference summary can be used as the answer difference data; and the truncated prompt, the first truncated answer, the second truncated answer and the answer difference data can be combined in a specified order to generate the first data and the second data.
[0066] For example, see Figure 4 , which shows a schematic diagram of generating the first data and the second data. When the first and second answers are truncated, candidate clauses corresponding to the first answer and candidate clauses corresponding to the second answer are obtained. The similarity difference can be directly calculated using the candidate clauses corresponding to the two answers.
[0067] For each candidate clause corresponding to the first answer, the semantic similarity between each candidate clause and the pending prompt can be calculated to obtain multiple first candidate similarities; the multiple first candidate similarities are averaged to obtain the first similarity. Similarly, for each candidate clause corresponding to the second answer, the semantic similarity between each candidate clause and the pending prompt can be calculated to obtain multiple second candidate similarities; the multiple second candidate similarities are averaged to obtain the second similarity. The first similarity and the second similarity are the similarity difference.
[0068] By identifying the differences between the first and second answers, a summary of the differences can be generated. For example, a large language model can be used to identify key content in the first and second answers. This content can include key arguments, core entities, key values, etc., and then a short text summary can be generated to describe the differences between the two. For example, "Price information is mentioned in the first answer, but not in the second answer," or "Argument x in the first answer contradicts argument y in the second answer." To avoid overly lengthy summaries, the maximum number of tokens in the output length can be controlled to ensure the brevity of the difference summary.
[0069] The answer difference data may include the first similarity, the second similarity, and the difference summary. At this point, the answer difference data, the truncation prompt, the first truncation answer, and the second truncation answer have been obtained. Combining these four parts in a specified order can yield the first data and the second data.
[0070] For example, the specified order can include a first order and a second order, where the first order is the truncated prompt, the first truncated answer, the second truncated answer, and the answer difference data; the second order is the truncated prompt, the second truncated answer, the first truncated answer, and the answer difference data. The first data can be obtained by concatenating the data in the first order, and the second data can be obtained by concatenating the data in the second order. The first and second orders can be set as needed, ensuring that the positions of the first and second answers in the first data are different from those in the second data to reduce the interference of position on the prediction results.
[0071] S140 : Use the preference prediction model to perform prediction processing on the first data and the second data to obtain a first preference probability corresponding to the first language model and a second preference probability corresponding to the second language model.
[0072] The preference prediction model is pre-trained based on a large amount of preference data. This preference prediction model can be used to predict and compare two responses to determine the preference probability for each response. A higher preference probability indicates that the response more closely matches the user's preferences and is more likely to be accepted by the user, resulting in a more effective model for outputting that response. Therefore, by comparing the preference probabilities of two responses, we can compare two language models.
[0073] When the first data is input into the preference prediction model, the preference prediction model can output a first prediction probability corresponding to the first language model under the first data, and a second prediction probability corresponding to the second language model. When the second data is input into the preference prediction model, the preference prediction model can input a third prediction probability corresponding to the first language model under the second data, and a fourth prediction probability corresponding to the second language model.
[0074] The first prediction probability and the third prediction probability corresponding to the first language model are weighted averaged to obtain the first preference probability of the first language model. The second prediction probability and the fourth prediction probability corresponding to the second language model are weighted averaged to obtain the second preference probability of the second language model.
[0075] Before inputting the first data and the second data into the preference prediction model, a large amount of preference data is required to train the basic model to obtain a preference prediction model. Optionally, the preference prediction model can be obtained in the following manner: fine-tuning the basic model using real preference data to obtain a basic prediction model, the target loss function of the fine-tuning process includes cross-entropy loss, entropy regularization and comparison loss, the basic model includes a basic attention layer and a linear classification head, and the real preference data includes real preference labels; constructing multiple predicted preference data based on the sample prompts to be predicted and multiple candidate models, the predicted preference data includes predicted preference labels; using the predicted preference data to perform a first reinforcement process on the discrimination ability of the basic prediction model to obtain an intermediate prediction model; using multilingual real preference data to perform a second reinforcement process on the language ability of the intermediate prediction model to obtain a preference prediction model, the learning rate of the second reinforcement process on the basic attention layer is smaller than that of the first reinforcement process, and the learning rate on the linear classification head is larger than that of the first reinforcement process.
[0076] Preference data typically includes two responses from two models to the same prompt, along with a preference label. The preference label can be used to identify the response that better reflects human preferences. Based on the source of the preference label, preference data can be divided into real preference data and predicted preference data. Real preference data can include real preference labels, which are annotated by humans; predicted preference data can include predicted preference labels, which are predicted by the underlying prediction model.
[0077] The basic model is a model obtained by modifying the pre-trained large language model (LLM). The pre-trained LLM can select a general large model with excellent performance and suitable for downstream task fine-tuning as the backbone network. In order to enable it to perform preference judgment, one or more fully connected layers can be added to the top layer of the LLM, that is, the output representation of the last layer of transformer block, and a nonlinear activation function is used to form a nonlinear classification head. The output layer dimension of the nonlinear classification head can be set to n, corresponding to n output nodes, representing the preference probability of n models respectively. The preference probability can be understood as the winning probability corresponding to each model in the n model comparison. Among them, n can be set according to actual needs. In an embodiment of the present invention, n can be set to 2. The values of these two preference probabilities should theoretically be between 0 and 1, and their sum is close to 1. By modifying the general large language model in this way, the basic model can be obtained. The basic model can include a transformer block, i.e., a basic attention layer, and a linear classification head.
[0078] The real preference data can be obtained from various public websites or obtained through manual annotation. The basic model can be fine-tuned using the real preference data to obtain a basic prediction model. Specifically, when using the real preference data to fine-tune the basic model, multiple real preference data can be obtained, and the real preference data include sample prompts, first sample answers generated by the first sample model based on the sample prompts, second sample answers generated by the second sample model based on the sample prompts, sample answer difference data between the first sample answer and the second sample answer, and real preference labels; semantic truncation is performed on each of the real preference data, and the positions of the first sample answer and the second sample answer are exchanged according to the specified probability to obtain multiple preference data to be used; the preference data to be used are predicted using the basic model to obtain the first sample preference probability of the first sample model; the target loss function is calculated based on the first sample preference probability and the real preference label; when the target loss function converges, the basic prediction model is obtained.
[0079] The actual preference data may include a sample prompt, a first sample answer, a second sample answer, sample answer difference data, and the actual preference label. The first sample answer is generated by the first sample model based on the sample prompt, and the second sample answer is generated by the second sample model based on the sample prompt. The sample answer difference data is generated in the same manner as the answer difference data described above. Please refer to the corresponding description in the previous embodiment and will not be repeated here.
[0080] The actual preference data needs to be semantically truncated in the manner described above. The specific processing process can also be referred to the corresponding description in the previous embodiment. To ensure that the model can learn the quality of the content itself, rather than relying on the order in which the answers appear, the positions of the first and second sample answers can be swapped according to a certain probability, thereby significantly reducing their positional bias. The specified probability can be set according to actual needs and is not specifically limited here.
[0081] The real preference data after semantic truncation and exchange processing can be recorded as the preference data to be used. The preference data to be used is input into the basic model, and the basic model performs prediction processing to output the first sample preference probability corresponding to the first sample model.
[0082] The basic model is strictly organized according to the format of the system prompts and the preference data to be used. The system prompts are a key part and can be set according to actual needs. They are intended to clearly issue instructions to the basic model, requiring it to play the role of an impartial referee, and can provide specific guidance on how to handle inspections, guiding the basic model to focus on the content quality itself. In an embodiment of the present invention, the system prompt words can be as follows:
[0083] Act as an impartial evaluator to assess two AI responses to the user's query.Compare their quality based on:
[0084] Instruction adherence
[0085] Question relevance
[0086] Accuracy and depth
[0087] Creative problem-solving
[0088] Level of detail
[0089] Overall helpfulness
[0090] Analyze both responses objectively,disregarding presentation order,response length,and assistant identities.Provide a concise comparisonhighlighting key strengths / weaknesses,then conclude with your verdict using[[A]]or[[B]]for the superior response.
[0091] The basic model can output the first sample preference probability corresponding to the first sample answer and the second sample preference probability corresponding to the second sample answer. Using the first sample preference probability and the true preference label, a target loss function can be constructed. The target loss function can include cross-entropy loss, entropy regularization, and comparison loss.
[0092] The cross entropy loss is the basic part of the target loss function, which ensures that the basic prediction model can perform basic classification. Assuming that there are N preference data to be used in a training batch, the cross entropy loss can be expressed as:
[0093]
[0094] Among them, L CErepresents the cross entropy loss function; y represents the true preference label, y∈{0,1}, 1 represents the first sample model wins, 0 represents the second sample model wins; p represents the winning probability of the first sample model predicted by the basic model, that is, the first sample preference probability; N represents the number of preference data to be used in the current training batch.
[0095] Entropy measures the uncertainty of a probability distribution. By penalizing high entropy, meaning the model is uncertain about its predictions and the probability distribution is relatively flat, the model can be encouraged to generate sharper, more peaked probability outputs. Therefore, the objective loss function can also include entropy regularization, a penalty term related to the entropy of the model's output probability distribution. Entropy regularization can be expressed as the following formula:
[0096]
[0097] Among them, L entropy represents entropy regularization; p represents the winning probability of the first sample model predicted by the basic model, that is, the first sample preference probability.
[0098] The comparative loss originates from the matching and comparison model in statistics and is suitable for handling preference selection problems. It does not directly optimize the score or probability of a single answer, but instead focuses on optimizing the winner's score to be significantly higher than the loser's score. Specifically, it tends to maximize the sigmoid function of the difference between the two scores. Compared with cross entropy, the comparative loss focuses more directly on the relative ranking relationship between sample pairs, which helps the model form a clearer boundary near the decision boundary. It can especially effectively distinguish samples of similar quality that are difficult to judge, thereby improving the model's robustness and ranking consistency. The comparative loss can be expressed by the following formula:
[0099] L BT =-logσ(s pos -s neg );
[0100] Among them, s pos Represents the sum of the first sample preference probabilities of the data whose true preference label is 1 (the first sample model wins) among the N preference data to be used in the current training batch; neg Represents the sum of the first sample preference probabilities of the data whose true preference label is 0 (the second sample model wins) among the N preference data to be used; σ is the Sigmoid function.
[0101] Therefore, the objective loss function can be expressed as:
[0102] L=L CE +λ1*L entropy +λ2*L BT ;
[0103] Among them, L represents the target loss function; L CE Characterizes cross entropy loss; L entropy Characterization entropy regularization; L BT Characterizes the comparison loss; the initial value of λ1 can be set to a small value to avoid over-suppressing model exploration. If there is a large amount of ambiguous data in the preference data to be used, such as the quality of two answers is close, this parameter can be gradually increased. In the embodiment of the present invention, the initial value of λ1 can be set to 0.1, and then it can be increased to 0.3. The initial value of λ2 can be set to a small value to emphasize the ability to distinguish data. If the validation set shows that the model has poor effect on difficult data, such as data with low confidence predictions, it can be gradually increased. In the embodiment of the present invention, the initial value of λ2 can be set to 0.5, and then it can be gradually increased to 1.0.
[0104] When the target loss function converges, the basic prediction model can be obtained. The basic prediction model has the ability to make preliminary preference judgments and has a certain resistance to certain deviations.
[0105] Because actual preference data is limited and expensive to generate, in order to expand the dataset and improve the model's generalization and coverage of various scenarios, predicted preference data can be constructed based on the sample prompts to be predicted and multiple candidate models. The sample prompts to be predicted can include prompts in different language types to enhance the model's multilingual capabilities. The predicted preference data can include predicted preference labels, meaning that the composition of the predicted preference data and the actual preference data is consistent, differing in that the actual preference labels are actually annotated by the user, while the predicted preference labels are predicted by the model.
[0106] Samples to be predicted refer to existing, large-scale, real, unlabeled user conversations or questions, which can be obtained from existing public datasets. To facilitate subsequent processing, the question data can be deduplicated and divided into English datasets and other multilingual datasets based on language type.
[0107] Multiple candidate models refer to multiple existing models, which cover different development manufacturers, model sizes, etc. as much as possible to ensure that the multiple candidate models have sufficient ability differences and style differences. Prediction preference data can be constructed using the sample prompt to be predicted and the multiple candidate models. Optionally, for each candidate sample prompt, a first candidate model and a second candidate model can be determined from the multiple candidate models; the candidate sample prompt can be processed using the first candidate model and the second candidate model to obtain a first candidate sample answer and a second candidate sample answer; the candidate sample prompt, the first candidate sample answer and the second candidate sample answer are input into the basic prediction model to obtain a first candidate preference probability and a second candidate preference probability; if the absolute value of the difference between the first candidate preference probability and the second candidate preference probability is greater than a specified value, a prediction preference label is determined based on the first candidate preference probability and the second candidate preference probability; the candidate sample prompt, the first candidate sample answer, the second candidate sample answer and the prediction preference label are used as prediction preference data.
[0108] To increase data diversity, each sample prompt to be predicted can be rewritten to generate multiple new sample prompts with similar semantics but different wording, thereby obtaining multiple candidate sample prompts. This rewriting process can utilize an existing large language model. Specifically, the sample prompt to be predicted is input into the large language model, which then outputs multiple rewritten new sample prompts. These new sample prompts, along with the sample prompt to be predicted, are then used as candidate sample prompts.
[0109] For each candidate sample prompt, a first candidate model and a second candidate model can be determined from the candidate models. The first candidate model and the second candidate model can be two models with relatively large differences among the candidate models, and the first candidate model and the second candidate model can set corresponding selection rules according to actual needs. For example, in an embodiment of the present invention, when determining the first candidate model and the second candidate model from multiple candidate models, it can be that for each candidate model, the model size and model source corresponding to the candidate model are obtained; based on the model size and the preset size range, the capability level of each candidate model is determined; one is randomly selected from the candidate models as the first candidate model, and the first capability level and the first model source corresponding to the first candidate model are obtained; and any candidate model among the remaining candidate models except the first candidate model, whose capability level is not the first capability level and whose model source is not the first model source, is determined as the second candidate model.
[0110] For each candidate model, the model size and model source corresponding to the candidate model can be obtained. The model size mainly refers to the parameter scale of the model, such as 10B, 1.5B, etc. The model source may include the developer of the model, the architecture of the model, etc. The capabilities of each candidate model can be roughly graded based on the model size. It is generally believed that the larger the parameter scale of the model, the stronger its capability. In an embodiment of the present invention, preset size ranges corresponding to different capability levels can be pre-set, and the capability level of each candidate model is determined based on the preset size range in which the candidate model is located. The preset size range can be set according to actual needs. For example, a parameter scale of no more than 10B is defined as a weak model; a parameter model greater than 10B and no more than 100B is defined as a medium model; a parameter scale greater than 100B is defined as a strong model.
[0111] Then, the manufacturer and architecture of each candidate model can be used as the source label. For each candidate sample prompt, a random selection of multiple candidate models can be made as the first candidate model, and the capability level corresponding to the first candidate model can be obtained as the first capability level. The source label of the first candidate model can also be obtained as the source of the first model. Then, from the remaining candidate models other than the first candidate model, the candidate models whose capability level is not the first capability level and whose model source is not the source of the first model can be selected as the second candidate model.
[0112] Optionally, the optional ability levels can be determined based on the first ability level. For example, if the first ability level is a medium model, the optional ability levels are a weak model and a strong model. For example, if the first ability level is a strong model, the optional ability levels are a weak model and a medium model. The candidate models in the optional ability levels are filtered based on the first model source, and the candidate models whose model sources are different from the first model source are eliminated. For ease of description, the remaining candidate models can be recorded as intermediate models, and then one can be randomly selected from the intermediate models as the second candidate model. The first candidate model and the second candidate model selected in this way have different model sizes and model sources. There may be more significant quality or style differences between the answers they generate for the same prompt, resulting in more information, which can help the basic prediction model learn training samples with different styles and obtain a more accurate preference prediction model.
[0113] The candidate sample prompts are input into the first candidate model and the second candidate model respectively so that the first candidate model and the second candidate model process them, and a first candidate sample answer output by the first candidate model and a second candidate sample answer output by the second candidate model can be obtained.
[0114] By combining the candidate sample prompt, the first candidate sample answer, and the second candidate sample answer in the same manner as described above to generate the first and second data, corresponding input data can be obtained. This input data is then fed into the basic prediction model to obtain the corresponding probability output by the basic prediction model. This probability indicates which model's answer is more consistent with the user's preference, thus obtaining the predicted preference label.
[0115] Since the basic prediction model is not a completely accurate model, its judgment on certain data may be inaccurate. In order to improve the quality of the predicted preference data, the comparison results with low confidence can be filtered out. According to the relevant description of the basic prediction model obtained by the above fine-tuning, the basic prediction model can predict the first candidate preference probability corresponding to the first candidate sample answer, and the second candidate preference probability corresponding to the second candidate sample answer. Calculate the absolute value of the difference between the two candidate preference probabilities. If the absolute value is not greater than the specified value, it indicates that the basic prediction model's judgment on the quality of the two answers is very close, and the discrimination is not high. Such samples can be considered to be fuzzy and may contain noise and need to be discarded. Among them, the specified value can be set according to actual needs, and is not specifically limited here. In the embodiment of the present invention, the specified value can be set to 0.03.
[0116] If the absolute value is greater than a specified value, the discrimination is considered high, and a predicted preference label can be determined based on the first candidate preference probability and the second candidate preference probability. Similarly, the constructed input data can be divided into first input data and second input data. Similar to the first and second data described above, a predicted preference label can be determined based on the corresponding preference probabilities by fine-tuning the basic prediction model.
[0117] Optionally, in order to further improve the reliability of the predicted preference labels, multiple basic prediction models with different architectures or training batches can be pre-trained, and the two basic prediction models can be used to predict the same input data separately to obtain the predicted preference labels. Only data with consistent two basic preference labels are retained, thereby achieving cross-validation of the results and providing guarantees for the reliability and accuracy of the predicted preference labels.
[0118] The candidate sample prompt, the first candidate sample answer, the second candidate sample answer and the predicted preference label are used as the predicted preference data.
[0119] It should be noted that the process of building the predicted preference data described above can be performed using an optimized LLM inference framework and server engine for offline batch processing. These frameworks support optimization technologies such as request batching, concurrent inference, and KV caching, significantly improving inference throughput and the efficiency of building the predicted preference data.
[0120] The predicted preference data is used to perform a first enhancement process on the discriminative ability of the basic prediction model to obtain an intermediate prediction model. In the aforementioned predicted preference data, each predicted preference data has a corresponding absolute value of the difference in preference probability. Since the larger the absolute value, the easier it is to distinguish between the two candidate sample answers in the predicted preference data, and the higher the discrimination degree, and the smaller the absolute value, the more difficult it is to distinguish between the two candidate sample answers in the predicted preference data, and the lower the discrimination degree, based on the absolute value, multiple predicted preference data can be graded in difficulty, thereby obtaining predicted preference data with multiple discrimination degrees.
[0121] Optionally, if the absolute value is greater than the first specified value, it can be determined as low difficulty data; if the absolute value is greater than the second specified value and not greater than the first specified value, it can be determined as medium difficulty data, wherein the second specified value is less than the first specified value; if the absolute value is not greater than the second specified value, it can be determined as high difficulty data. The first specified value and the second specified value can be set according to actual needs. In an embodiment of the present invention, the first specified value can be set to 0.8 and the second specified value can be set to 0.5. The predicted preference data of each difficulty level is then proportionally balanced to ensure that the predicted preference data for training can include data of various difficulties, so as to ensure that the model can learn to process data of various difficulties and improve the discrimination and judgment ability of the preference prediction model.
[0122] When further fine-tuning the base prediction model, to avoid the inherent noise in the predicted preference data interfering with the model learning process and the problem of overfitting to a single language leading to insufficient generalization ability to other languages, the base prediction model can be fine-tuned in multiple stages using the predicted preference data and the real preference data. First, the discriminative ability of the base prediction model is enhanced using the predicted preference data to obtain an intermediate prediction model. Then, the language ability of the intermediate prediction model is enhanced using the multilingual real preference data to obtain the final preference prediction model. It is understood that the multilingual real preference data used here and the real preference data used to obtain the base prediction model after fine-tuning the base model can be different.
[0123] When obtaining an intermediate prediction model, the generated prediction preference data can be used to further fine-tune the base prediction model. This allows the base prediction model to learn common preference judgment patterns, strengthen the model's discrimination capabilities, and expose the base prediction model to a variety of different answer styles. Specifically, when fine-tuning the base prediction model, it can be trained step by step according to the aforementioned difficulty level, in the order of low difficulty, medium difficulty, and high difficulty, to obtain an intermediate prediction model.
[0124] The intermediate prediction model already has basic preference discrimination capabilities, and some real preference data can be used to perform a second enhancement on its language capabilities. For example, the intermediate prediction model can be fine-tuned using data from the multilingual real preference data for a specific language to ensure that, under accurate supervision signals, the accuracy of the model's preference judgment in the specified language can be aligned with human preferences. On this basis, further fine-tuning can be performed using real preference data other than the specified language, as well as a small amount of unused real preference data for the specified language, to improve the model's generalization capabilities in non-specified languages, allowing it to accurately judge the pros and cons of answers in different language environments, thereby obtaining the final preference prediction model.
[0125] From the basic prediction model to the preference prediction model, a three-stage fine-tuning process was performed. The first stage was to fine-tune the intermediate prediction model. The second stage was to fine-tune the intermediate prediction model using real preference data in a specified language. The third stage was to fine-tune the final preference prediction model based on the second stage using other languages and a small amount of real preference data. The specified language can be set according to actual needs, for example, English or Chinese.
[0126] Because the model may forget the knowledge learned in the early stages during multi-stage fine-tuning, which is called catastrophic forgetting, different learning rates can be used in different training stages to alleviate this problem. For example, in the second and third stages using multilingual real preference data, a smaller learning rate can be used for the base attention layer than in the first stage, and an even smaller learning rate can be used for the linear classification head than in the first stage.
[0127] For example, the learning rate configuration of each stage in the present invention can be seen in Table 1:
[0128] Table 1
[0129] Basic Attention Layer Linear classification head Phase 1 1e-4 1e-4 Phase 2 2e-5 2e-4 Phase 3 2e-5 2e-4
[0130] During the second and third stages of fine-tuning, most of the underlying model parameters (all layers except the last eight Transformer blocks and the classification head) can be completely frozen. Only the top few layers and the final classification head are fine-tuned. This allows the model's core representations to be more strongly protected, allowing the model to focus on learning high-level knowledge related to specific preference judgment patterns and specific languages. By fine-tuning the base prediction model in multiple stages in this manner, a preference prediction model with good performance and accurate prediction of preferences can be obtained.
[0131] Through guided fine-tuning, an optimized loss function, and a multi-stage fine-tuning strategy, the model's sensitivity to non-content factors such as answer position and length is effectively suppressed. This ensures that the preference prediction model's predictions better reflect the quality of the answer content, resulting in more accurate model comparisons that better align with human preferences. By constructing and quality-controlling predicted preference data, large-scale training data can be generated efficiently and cost-effectively. Furthermore, staged fine-tuning ensures that the preference prediction model maintains its accuracy across multiple language environments. Using this preference prediction model for model comparisons improves accuracy.
[0132] S150: Determine a comparison result between the first language model and the second language model based on the first preference probability and the second preference probability.
[0133] The first preference probability can be understood as the probability that the first answer output by the first language model conforms to human preferences, and the second preference probability can be understood as the probability that the second answer output by the second language model conforms to human preferences. By comparing the first preference probability and the second preference probability, the comparison result of the first language model and the second language model can be obtained. For example, if the first preference probability is greater than the second preference probability, it is considered that the first language model is better than the second language model. The preference-based language model comparison method provided by the embodiment of the present invention can be used in a variety of scenarios that require comparison of model output effects and quality. For example, for a global intelligent customer service assistant agent, after the developer updates and iterates it, the agent before the update can be used as the first language model, and the updated agent can be used as the second language model. By automatically calling the preference prediction model, the answer effects of the two versions can be compared. In order to ensure the accuracy of the comparison, multiple groups of different data to be processed can be provided. For each group of data to be processed, the preference prediction model can output a first preference probability and a second preference probability. The average value of the first preference probability corresponding to the multiple groups of data to be processed is compared with the average value of the second preference probability to achieve the effect of comparing the answers of the new and old versions, and then obtain the comparison results, which can greatly shorten the traditional test time. Moreover, since the preference prediction model has been trained in the above manner, it can accurately compare the answers and output accurate comparison results.
[0134] Optionally, a preset threshold may be set in advance. When the average value of the preference probability of the new version is less than the average value of the preference probability of the old version, and the difference between the two is greater than the preset threshold, an alarm may be directly issued.
[0135] The preference-based language model comparison solution provided by the embodiments of the present invention can be applied in various scenarios where the quality and effectiveness of model output are being compared. For example, in the case of the iteration between the old and new versions of an intelligent assistant, the solution provided by the embodiments of the present invention can rely on the preference prediction model to efficiently and accurately compare the response effects of two language models based on preferences, thus obtaining a comparison result. This can replace the traditional method of testing two models more efficiently and accurately.
[0136] The preference-based model comparison method provided in the embodiment of the present invention can obtain a preference prediction model through training in a specific training method. The preference prediction model is applicable to a multi-language environment and can accurately predict preferences. The preference prediction model is used to predict the preference probabilities of the first language model and the second language model, and the final comparison result is determined based on the corresponding preference probabilities, which is more reliable and accurate. In order to better implement the above method, the embodiment of the present invention also provides a preference-based language model comparison device. The preference-based language model comparison device can be specifically integrated into an electronic device, which can be a terminal, server, or other device. Among them, the terminal can be a mobile phone, a tablet computer, a smart Bluetooth device, a laptop computer, a personal computer, or other device; the server can be a single server or a server cluster composed of multiple servers.
[0137] For example, in this embodiment, the method of the embodiment of the present invention will be described in detail by taking the preference-based language model comparison device specifically integrated into the server as an example.
[0138] For example, Figure 5 As shown, the preference-based language model comparison apparatus 200 may include an acquisition module 210, a truncation module 220, a generation module 230, a prediction module 240, and a comparison module 250, as follows:
[0139] An acquisition module 210 is configured to acquire data to be processed, wherein the data to be processed includes a plurality of sub-data, wherein the sub-data includes a prompt to be processed, a first answer generated by a first language model based on the prompt to be processed, and a second answer generated by a second language model based on the prompt to be processed;
[0140] A truncation module 220 is configured to perform semantic truncation processing on each sub-data based on the length of the sub-data, the length of the data to be processed, and the length limit of the data to be processed set by the preference prediction model, to obtain truncated data, wherein the truncated data includes a truncation prompt, a first truncation answer, and a second truncation answer;
[0141] A generating module 230 for generating first data and second data based on answer difference data between the first answer and the second answer and the truncated data, wherein the first truncated answer and the second truncated answer have different positions in the first data than in the second data;
[0142] Prediction module 240, configured to perform prediction processing on the first data and the second data using the preference prediction model to obtain a first preference probability corresponding to the first language model and a second preference probability corresponding to the second language model;
[0143] The comparison module 250 is configured to determine a comparison result between the first language model and the second language model based on the first preference probability and the second preference probability.
[0144] In some embodiments, the truncation module 220 is specifically configured to:
[0145] For each sub-data, multiply the ratio of the length of the sub-data to the length of the data to be processed by the limited length to calculate the available length of the sub-data;
[0146] If the length of the sub-data is greater than the available length of the sub-data and the sub-data is a prompt to be processed, truncating the prompt to be processed based on the available length of the prompt to be processed and a preset rule to obtain a truncated prompt;
[0147] If the length of the sub-data is greater than the available length of the sub-data and the sub-data is answer data, a candidate clause is determined from the answer data using the available length of the answer data, and a truncated answer corresponding to the answer data is generated based on the candidate clause, the answer data includes a first answer and a second answer, and the truncated answer includes a first truncated answer and a second truncated answer.
[0148] In some embodiments, the truncation module 220 is specifically configured to:
[0149] Dividing the available length of the prompt to be processed into a first length and a second length, wherein the sum of the first length and the second length is the available length;
[0150] Using the first length of content before the prompt to be processed as the first truncated content;
[0151] Using the second length of content after the prompt to be processed as the second truncated content;
[0152] The first truncated content and the second truncated content are spliced together to obtain a truncation prompt.
[0153] In some embodiments, the truncation module 220 is specifically configured to:
[0154] splitting the answer data into a plurality of clauses, wherein each clause has a clause number indicating an order of the clause in the answer data;
[0155] Calculating the semantic similarity between each clause and the prompt to be processed;
[0156] determining a candidate clause from the plurality of clauses based on the semantic similarity and an available length of the answer data;
[0157] The candidate clauses are sorted according to the clause numbers to obtain truncated answers.
[0158] In some embodiments, the generation module 230 is specifically configured to:
[0159] Calculating the average similarity between all candidate clauses of the first answer and the prompt to be processed to obtain a first similarity;
[0160] Calculating the average similarity between all candidate clauses of the second answer and the prompt to be processed to obtain a second similarity;
[0161] identifying differences between the first answer and the second answer and obtaining a summary of the differences;
[0162] using the first similarity, the second similarity, and the difference summary as answer difference data;
[0163] The truncated prompt, the first truncated answer, the second truncated answer, and the answer difference data are combined in a specified order to generate first data and second data.
[0164] In some embodiments, the preference-based language model comparison apparatus 200 further includes a training module. Before using the preference prediction model to perform prediction processing on the first data and the second data to obtain a first preference probability corresponding to the first language model and a second preference probability corresponding to the second language model, the training module is specifically configured to:
[0165] Fine-tuning the base model using the real preference data to obtain a base prediction model, wherein the target loss function of the fine-tuning includes a cross-entropy loss, entropy regularization, and comparison loss, the base model includes a basic attention layer and a linear classification head, and the real preference data includes a real preference label;
[0166] Constructing prediction preference data based on the sample prompt to be predicted and multiple candidate models, wherein the prediction preference data includes a prediction preference label;
[0167] Performing a first enhancement process on the discrimination capability of the basic prediction model using the prediction preference data to obtain an intermediate prediction model;
[0168] The language ability of the intermediate prediction model is subjected to a second reinforcement processing using multilingual real preference data to obtain a preference prediction model. The learning rate of the second reinforcement processing on the basic attention layer is smaller than that of the first reinforcement processing, and the learning rate on the linear classification head is larger than that of the first reinforcement processing.
[0169] In some embodiments, the training module is specifically configured to:
[0170] Acquire a plurality of real preference data, the real preference data including a sample prompt, a first sample answer generated by a first sample model based on the sample prompt, a second answer generated by a second sample model based on the sample prompt, sample answer difference data between the first sample answer and the second sample answer, and a real preference label;
[0171] Performing semantic truncation processing on each of the true preference data, and swapping the positions of the first sample answer and the second sample answer according to a specified probability, to obtain a plurality of preference data to be used;
[0172] Using the basic model to perform prediction processing on the preference data to be used, to obtain a first sample preference probability of a first sample model;
[0173] Calculate the target loss function based on the first sample preference probability and the true preference label;
[0174] When the objective loss function converges, a basic prediction model is obtained.
[0175] In some embodiments, the training module is specifically configured to:
[0176] For each of the sample prompts to be predicted, rewriting the sample prompt to be predicted to obtain multiple candidate sample prompts;
[0177] For each candidate sample prompt, determining a first candidate model and a second candidate model from the multiple candidate models;
[0178] Using the first candidate model and the second candidate model, respectively process the candidate sample prompt to obtain a first candidate sample answer and a second candidate sample answer;
[0179] Inputting the candidate sample prompt, the first candidate sample answer, and the second candidate sample answer into the basic prediction model to obtain a first candidate preference probability and a second candidate preference probability;
[0180] If the absolute value of the difference between the first candidate preference probability and the second candidate preference probability is greater than a specified value, determining a predicted preference label based on the first candidate preference probability and the second candidate preference probability;
[0181] The candidate sample prompt, the first candidate sample answer, the second candidate sample answer, and the predicted preference label are used as predicted preference data.
[0182] In some embodiments, the training module is specifically configured to:
[0183] For each candidate model, obtain the model size and model source corresponding to the candidate model;
[0184] Determining a capability level for each candidate model based on the model size and a preset size range;
[0185] Randomly selecting one of the candidate models as a first candidate model, and obtaining a first capability level and a first model source corresponding to the first candidate model;
[0186] Among the remaining candidate models except the first candidate model, any candidate model whose capability level is not the first capability level and whose model source is not the first model source is determined as a second candidate model.
[0187] In specific implementation, the above units can be implemented as independent entities, or can be arbitrarily combined to be implemented as the same or several entities. The specific implementation of the above units can be found in the previous method embodiments and will not be repeated here.
[0188] From the above, it can be seen that the preference-based language model comparison device of this embodiment can obtain data to be processed containing multiple sub-data, and perform semantic truncation processing on the sub-data according to the length of the sub-data, the length of the data to be processed, and the limit length of the preference prediction model for processing, to ensure that the truncated data carries higher effective information, and then generate first data and second data based on the truncated data and the answer difference data. The introduction of the answer difference data can assist the model in accurately understanding the difference between the two answers, and the positions of the two answers in the first data are different from those in the second data, which can reduce the impact of the position of the answer on the result. After processing by the preference prediction model, the first preference probability of the first language model and the second preference probability of the second language model can be obtained; finally, the comparison result is obtained directly based on the first preference probability and the second preference probability, which can accurately realize model comparison.
[0189] An embodiment of the present invention further provides an electronic device, which may be a terminal, a server, or the like. The terminal may be a mobile phone, a tablet computer, a smart Bluetooth device, a laptop computer, a personal computer, or the like; the server may be a single server or a server cluster consisting of multiple servers, or the like.
[0190] In some embodiments, the preference-based language model comparison device can also be integrated into multiple electronic devices. For example, the preference-based language model comparison device can be integrated into multiple servers, and the preference-based language model comparison method of the present invention can be implemented by multiple servers.
[0191] In this embodiment, the electronic device of this embodiment is a server as an example for detailed description, for example, Figure 6, which shows a schematic structural diagram of an electronic device involved in an embodiment of the present invention, specifically:
[0192] The electronic device may include one or more processing core processors 310, one or more computer-readable storage media memories 320, a power supply 330, an input module 340, and a communication module 350. Those skilled in the art will appreciate that Figure 6 The electronic device structure shown in the figure does not constitute a limitation of the electronic device, and may include more or fewer components than shown in the figure, or combine certain components, or arrange components differently.
[0193] The processor 310 is the control center of the electronic device. It connects all parts of the electronic device using various interfaces and circuits. It executes software programs and / or modules stored in the memory 320 and accesses data stored in the memory 320 to perform various functions of the electronic device and process data. In some embodiments, the processor 310 may include one or more processing cores. In some embodiments, the processor 310 may integrate an application processor and a modem processor. The application processor primarily handles the operating system, user interface, and application programs, while the modem processor primarily handles wireless communications. It is understood that the modem processor may not be integrated into the processor 310.
[0194] The memory 320 can be used to store software programs and modules. The processor 310 executes various functional applications and data processing by running the software programs and modules stored in the memory 320. The memory 320 may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area may store data created according to the use of the electronic device, etc. In addition, the memory 320 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage device. Accordingly, the memory 320 may also include a memory controller to provide the processor 310 with access to the memory 320.
[0195] The electronic device also includes a power supply 330 for supplying power to various components. In some embodiments, the power supply 330 can be logically connected to the processor 310 via a power management system, thereby enabling the power management system to manage charging, discharging, and power consumption. The power supply 330 can also include any components such as one or more DC or AC power supplies, a recharging system, a power failure detection circuit, a power converter or inverter, and a power status indicator.
[0196] The electronic device may further include an input module 340 , which may be configured to receive input digital or character information and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function controls.
[0197] The electronic device may further include a communication module 350. In some embodiments, the communication module 350 may include a wireless module. The electronic device may perform short-range wireless transmission via the wireless module of the communication module 350, thereby providing the user with wireless broadband Internet access. For example, the communication module 350 may be used to help the user send and receive emails, browse web pages, and access streaming media.
[0198] Although not shown, the electronic device may further include a display unit, etc., which will not be described in detail herein. Specifically, in this embodiment, the processor 310 in the electronic device loads the executable files corresponding to one or more application processes into the memory 320 according to the following instructions, and the processor 310 runs the application stored in the memory 320, thereby implementing the steps of the method in each embodiment of the present invention.
[0199] The specific implementation of the above operations can be found in the previous embodiments and will not be repeated here.
[0200] From the above, it can be seen that the electronic device provided by the embodiment of the present invention can obtain data to be processed containing multiple sub-data, and perform semantic truncation processing on the sub-data according to the length of the sub-data, the length of the data to be processed, and the limited length of the preference prediction model to be processed, to ensure that the truncated data carries a higher level of effective information, and then generate the first data and the second data based on the truncated data and the answer difference data. The introduction of the answer difference data can assist the model in accurately understanding the difference between the two answers, and the positions of the two answers in the first data are different from those in the second data, which can reduce the impact of the position of the answer on the result. After processing by the preference prediction model, the first preference probability of the first language model and the second preference probability of the second language model can be obtained; finally, the comparison result is obtained directly based on the first preference probability and the second preference probability, which can accurately realize model comparison.
[0201] Those skilled in the art will appreciate that all or part of the steps in the various methods of the above embodiments may be accomplished by instructions, or by controlling related hardware through instructions. The instructions may be stored in a computer-readable storage medium and loaded and executed by a processor.
[0202] To this end, an embodiment of the present invention provides a computer-readable storage medium storing a plurality of instructions, which can be loaded by a processor to execute the steps of any preference-based language model comparison method provided in an embodiment of the present invention.
[0203] The storage medium may include a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0204] According to one aspect of the present invention, a computer program product or computer program is provided, comprising a computer program / instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer program / instructions from the computer-readable storage medium and executes the computer program / instructions, causing the electronic device to perform the methods provided in the various optional implementations of the aforementioned embodiments for preference-based language model comparison, preference prediction model training, or predicted preference data construction.
[0205] Since the instructions stored in the storage medium can execute the steps of any preference-based language model comparison method provided in the embodiments of the present invention, the beneficial effects that can be achieved by any preference-based language model comparison method provided in the embodiments of the present invention can be achieved. Please refer to the previous embodiments for details and will not be repeated here.
[0206] The above is a detailed introduction to a preference-based language model comparison method and device provided in an embodiment of the present invention. Specific examples are used herein to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea. At the same time, for those skilled in the art, based on the ideas of the present invention, there may be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting the present invention.
Claims
1. A preference-based language model comparison method, characterized in that: The method comprises: Acquire data to be processed, the data to be processed including a plurality of sub-data, the sub-data including a prompt to be processed, a first answer generated by a first language model based on the prompt to be processed, and a second answer generated by a second language model based on the prompt to be processed; For each of the sub-data, based on the length of the sub-data, the length of the data to be processed, and the length limit of the data to be processed set by the preference prediction model, semantically truncate the sub-data to obtain truncated data, wherein the truncated data includes a truncation prompt, a first truncation answer, and a second truncation answer; generating first data and second data based on answer difference data between the first answer and the second answer and the truncated data, wherein the first truncated answer and the second truncated answer have different positions in the first data than in the second data; Using the preference prediction model to perform prediction processing on the first data and the second data, obtaining a first preference probability corresponding to the first language model and a second preference probability corresponding to the second language model; A comparison result between the first language model and the second language model is determined based on the first preference probability and the second preference probability.
2. The method according to claim 1, characterized in that For each of the sub-data, based on the length of the sub-data, the length of the data to be processed, and the length limit of the data to be processed by the preference prediction model, semantic truncation processing is performed on the sub-data to obtain truncated data, including: For each sub-data, multiply the ratio of the length of the sub-data to the length of the data to be processed by the limited length to calculate the available length of the sub-data; If the length of the sub-data is greater than the available length of the sub-data and the sub-data is a prompt to be processed, truncating the prompt to be processed based on the available length of the prompt to be processed and a preset rule to obtain a truncated prompt; If the length of the sub-data is greater than the available length of the sub-data and the sub-data is answer data, a candidate clause is determined from the answer data using the available length of the answer data, and a truncated answer corresponding to the answer data is generated based on the candidate clause, the answer data includes a first answer and a second answer, and the truncated answer includes a first truncated answer and a second truncated answer.
3. The method according to claim 2, characterized in that Truncating the prompt to be processed based on the available length of the prompt to be processed and a preset rule to obtain a truncated prompt includes: Dividing the available length of the prompt to be processed into a first length and a second length, wherein the sum of the first length and the second length is the available length; Using the first length of content before the prompt to be processed as the first truncated content; Using the second length of content after the prompt to be processed as the second truncated content; The first truncated content and the second truncated content are spliced together to obtain a truncation prompt.
4. The method according to claim 2, characterized in that The determining of candidate clauses from the answer data using the available length of the answer data, and generating a truncated answer corresponding to the answer data based on the candidate clauses, includes: splitting the answer data into a plurality of clauses, wherein each clause has a clause number indicating an order of the clause in the answer data; Calculating the semantic similarity between each clause and the prompt to be processed; determining a candidate clause from the plurality of clauses based on the semantic similarity and an available length of the answer data; The candidate clauses are sorted according to the clause numbers to obtain truncated answers.
5. The method according to claim 2, characterized in that The generating of first data and second data based on the answer difference data between the first answer and the second answer and the truncated data includes: Calculating the average similarity between all candidate clauses of the first answer and the prompt to be processed to obtain a first similarity; Calculating the average similarity between all candidate clauses of the second answer and the prompt to be processed to obtain a second similarity; identifying differences between the first answer and the second answer and obtaining a summary of the differences; using the first similarity, the second similarity, and the difference summary as answer difference data; The truncated prompt, the first truncated answer, the second truncated answer, and the answer difference data are combined in a specified order to generate first data and second data.
6. The method according to claim 1, characterized in that Before performing prediction processing on the first data and the second data using the preference prediction model to obtain a first preference probability corresponding to the first language model and a second preference probability corresponding to the second language model, the method further includes: Fine-tuning the base model using the real preference data to obtain a base prediction model, wherein the target loss function of the fine-tuning includes a cross-entropy loss, entropy regularization, and comparison loss, the base model includes a basic attention layer and a linear classification head, and the real preference data includes a real preference label; Constructing prediction preference data based on the sample prompt to be predicted and multiple candidate models, wherein the prediction preference data includes a prediction preference label; Performing a first enhancement process on the discrimination capability of the basic prediction model using the prediction preference data to obtain an intermediate prediction model; The language ability of the intermediate prediction model is subjected to a second reinforcement processing using multilingual real preference data to obtain a preference prediction model. The learning rate of the second reinforcement processing on the basic attention layer is smaller than that of the first reinforcement processing, and the learning rate on the linear classification head is larger than that of the first reinforcement processing.
7. The method according to claim 6, characterized in that The method of fine-tuning the basic model using the real preference data to obtain the basic prediction model includes: Acquire a plurality of real preference data, the real preference data including a sample prompt, a first sample answer generated by a first sample model based on the sample prompt, a second answer generated by a second sample model based on the sample prompt, sample answer difference data between the first sample answer and the second sample answer, and a real preference label; Performing semantic truncation processing on each of the true preference data, and swapping the positions of the first sample answer and the second sample answer according to a specified probability, to obtain a plurality of preference data to be used; Using the basic model to perform prediction processing on the preference data to be used, to obtain a first sample preference probability of a first sample model; Calculate the target loss function based on the first sample preference probability and the true preference label; When the objective loss function converges, a basic prediction model is obtained.
8. The method according to claim 6, characterized in that The method of constructing prediction preference data based on the sample prompt to be predicted and multiple candidate models includes: For each of the sample prompts to be predicted, rewriting the sample prompt to be predicted to obtain multiple candidate sample prompts; For each candidate sample prompt, determining a first candidate model and a second candidate model from the multiple candidate models; Using the first candidate model and the second candidate model, respectively process the candidate sample prompt to obtain a first candidate sample answer and a second candidate sample answer; Inputting the candidate sample prompt, the first candidate sample answer, and the second candidate sample answer into the basic prediction model to obtain a first candidate preference probability and a second candidate preference probability; If the absolute value of the difference between the first candidate preference probability and the second candidate preference probability is greater than a specified value, determining a predicted preference label based on the first candidate preference probability and the second candidate preference probability; The candidate sample prompt, the first candidate sample answer, the second candidate sample answer, and the predicted preference label are used as predicted preference data.
9. The method according to claim 8, characterized in that The determining of the first candidate model and the second candidate model from the multiple candidate models includes: For each candidate model, obtain the model size and model source corresponding to the candidate model; Determining a capability level for each candidate model based on the model size and a preset size range; Randomly selecting one of the candidate models as a first candidate model, and obtaining a first capability level and a first model source corresponding to the first candidate model; Among the remaining candidate models except the first candidate model, any candidate model whose capability level is not the first capability level and whose model source is not the first model source is determined as a second candidate model.
10. A preference-based language model comparison device, the device being used to implement the method according to any one of claims 1 to 9, characterized in that: The device comprises: an acquisition module, configured to acquire data to be processed, the data to be processed comprising a plurality of sub-data, the sub-data comprising a prompt to be processed, a first answer generated by a first language model based on the prompt to be processed, and a second answer generated by a second language model based on the prompt to be processed; a truncation module, configured to perform semantic truncation processing on each sub-data based on the length of the sub-data, the length of the data to be processed, and the length limit of the data to be processed set by the preference prediction model, to obtain truncated data, wherein the truncated data includes a truncation prompt, a first truncation answer, and a second truncation answer; a generating module for generating first data and second data based on answer difference data between the first answer and the second answer and the truncated data, wherein the first truncated answer and the second truncated answer have different positions in the first data than in the second data; A prediction module, configured to perform prediction processing on the first data and the second data using the preference prediction model to obtain a first preference probability corresponding to the first language model and a second preference probability corresponding to the second language model; A comparison module is configured to determine a comparison result between the first language model and the second language model based on the first preference probability and the second preference probability.
Citation Information
Patent Citations
Large model preference alignment method improved by using online synchronization strategy
CN119539082A
A language model training method, device, storage medium and electronic device
CN119740024A
Question and answer generation method and system based on big language model preference alignment
CN119938863A
Universal Language Segment Representations Learning with Conditional Masked Language Model
US20220198144A1
Privacy-preserving text insight mining in a closed domain
US20230161972A1