Large language model determination method and device and program product

By constructing the preferred data set and constructing loss functions based on probability, the problem of insufficient balance in multiple dimensions of large language model replies is solved, and a more balanced reply output is achieved.

CN120216643AActive Publication Date: 2025-06-27SHENZHEN MEITUAN TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510289383.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-06-27
Estimated Expiration
2045-03-11

AI Technical Summary

Technical Problem

The current large language model's responses are not balanced enough in multiple dimensions, resulting in an uneven dimension of the reply.

Method used

By constructing a preferred data set in the data collection stage and building a loss function based on the probability of generating data in each dimension in the training stage of the large language model, adaptive multi-objective training is performed until the loss function converges, and the target large language model is obtained.

Benefits of technology

The balance of large language model replies in multiple dimensions is improved, ensuring that the difference between data from model outputs is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216643A_ABST
    Figure CN120216643A_ABST
Patent Text Reader

Abstract

The invention provides a large language model determination method and device and a program product, and relates to the technical field of computers.The method comprises the steps that in the data collection stage, a preference data set is constructed for multiple dimensions, and a loss function corresponding to a large language model is constructed at a large language model training node based on the probability of generating data of each dimension, according to the method, adaptive multi-target training is carried out on the preference data set until the loss function is converged, the target large language model is obtained, and the obtained target large language model can improve the balance of reply in multiple dimensions.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0002] A large language model refers to a model constructed using artificial intelligence technology that can automatically learn, adapt, and optimize to achieve more accurate and efficient prediction and decision-making. It is usually constructed based on technologies such as machine learning and deep learning, and can learn patterns and rules from a large amount of data and make predictions and decisions. With the development of the times, the application of large language models is becoming more and more extensive, but the current responses of large language models do not focus on the same dimensions. As a result, the dimensional emphasis of the current responses is unbalanced.

[0003] It should be noted that the information disclosed in the above background art section is only used to enhance the understanding of the background of the present disclosure, and therefore may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention

[0004] The present disclosure provides a method, apparatus, and program product for determining a large language model, which at least improves the balance of the responses of the large language model in multiple dimensions.

[0005] Other features and advantages of the present disclosure will become apparent through the following detailed description, or will be partially learned through the practice of the present disclosure.

[0006] According to one aspect of the present disclosure, there is provided a method for determining a large language model, including:

[0007] In the data collection stage, a preference data set is constructed for multiple dimensions;

[0008] In the large language model training stage, a loss function corresponding to the large language model is constructed based on the probability of generating data for each dimension, and adaptive multi-objective training is performed on the preference data set until the loss function converges to obtain the target large language model.

[0009] In an embodiment of the present disclosure, the method further includes:

[0010] In the case where the loss function value corresponding to the loss function does not converge, weights corresponding to each dimension are determined based on the probability of generating data for each dimension and a Gaussian sampling model;

[0011] Adjust the weights corresponding to each dimension to reduce the difference between the multi-dimensional data corresponding to the response output by the large language model;

[0012] Update the loss function based on the adjusted weights corresponding to each dimension.

[0013] In an embodiment of the present disclosure, determining the weights corresponding to each dimension based on the probability of generating data for each dimension and a Gaussian sampling model includes:

[0014] Determine the probability mean and probability variance of generating data for each dimension based on the probability of generating data for each dimension;

[0015] Construct a Gaussian distribution model based on the probability mean and probability variance of generating data for each dimension;

[0016] Perform Gaussian sampling on the Gaussian distribution model to obtain the weights corresponding to each dimension.

[0017] In one embodiment of the present disclosure, constructing a loss function corresponding to the large language model based on the probability of generating data for each dimension includes:

[0018] Construct a loss function based on the weight corresponding to each dimension and the probability mean of generating data for each dimension.

[0019] In one embodiment of the present disclosure, the method further includes:

[0020] In the case where the loss function value corresponding to the loss function does not converge, adjust the model parameters to reduce the difference between the multi-dimensional data corresponding to the responses output by the large language model.

[0021] In one embodiment of the present disclosure,

[0022] The preference dataset includes prompts, multiple responses corresponding to the prompts, and multi-dimensional data corresponding to each response. Constructing a preference dataset for multiple dimensions includes:

[0023] Divide the multiple responses into winning responses and losing responses based on the multi-dimensional data corresponding to each response;

[0024] Concatenate the multi-dimensional prompts with the winning responses and losing responses respectively to obtain multiple reconstructed losing prompts and winning prompts.

[0025] In one embodiment of the present disclosure, the method further includes:

[0026] Input the multiple reconstructed losing prompts and winning prompts into the large language model respectively, and determine the probability of generating data for each dimension based on the probability of the winning prompts in the responses for each dimension output by the large language model.

[0027] In one embodiment of the present disclosure, the method further includes:

[0028] In response to a user inputting a question at the input end of the large language model, display a multi-dimensionally balanced response corresponding to the question.

[0029] According to another aspect of the present disclosure, the present disclosure provides a large language model determination device, including:

[0030] A construction module for constructing a preference dataset for multiple dimensions during the data collection phase;

[0031] A training module for constructing a loss function corresponding to the large language model based on the probability of generating data for each dimension during the large language model training phase, and performing adaptive multi-objective training on the preference dataset until the loss function converges to obtain the target large language model.

[0032] In an embodiment of the present disclosure, the device further includes:

[0033] A first determination module for determining the weight corresponding to each dimension based on the probability of generating data for each dimension and the Gaussian sampling model when the loss function value corresponding to the loss function has not converged;

[0034] An adjustment module for adjusting the weight corresponding to each dimension to reduce the difference between the multiple dimension data corresponding to the reply output by the large language model;

[0035] An update module for updating the loss function based on the adjusted weight corresponding to each dimension.

[0036] In an embodiment of the present disclosure, the first determination module includes:

[0037] A determination unit for determining the probability mean and probability variance of generating data for each dimension based on the probability of generating data for each dimension;

[0038] A first construction unit for constructing a Gaussian distribution model generation space based on the probability mean and probability variance of generating data for each dimension;

[0039] A sampling unit for performing Gaussian weight sampling on the Gaussian distribution model generation space to obtain the weight corresponding to each dimension.

[0040] In an embodiment of the present disclosure, the training module includes:

[0041] A second construction unit for constructing a loss function based on the weight corresponding to each dimension and the probability mean of generating data for each dimension, so that when the loss function value corresponding to the loss function gradually converges, the difference between the multiple dimension data is reduced.

[0042] In an embodiment of the present disclosure, the device further includes:

[0043] An adjustment module for adjusting the model parameters when the loss function value corresponding to the loss function has not converged, so as to reduce the difference between the multiple dimension data corresponding to the reply output by the large language model.

[0044] In an embodiment of the present disclosure, the construction module further includes:

[0045] A division unit, configured to divide multiple responses into winning responses and losing responses based on multi-dimensional data corresponding to each response;

[0046] A monocular splicing unit, configured to splice multi-dimensional prompts with winning responses and losing responses respectively to obtain multiple reconstructed losing prompt words and winning prompt words.

[0047] In an embodiment of the present disclosure, the apparatus further includes:

[0048] A second determination module, configured to input the multiple reconstructed losing prompt words and winning prompt words into a large language model respectively, and determine the probability of generating each dimension of data based on the probability of the winning prompt words in the responses of each dimension output by the large language model.

[0049] In an embodiment of the present disclosure, the apparatus further includes:

[0050] A display module, configured to display a multi-dimensionally balanced response corresponding to a question in response to a user inputting the question at the input end of the large language model.

[0051] According to still another aspect of the present disclosure, there is provided an electronic device, including: a processor; and a memory for storing executable instructions of the processor; wherein, the processor is configured to execute the above-mentioned large language model determination method of any one via executing the executable instructions.

[0052] According to yet another aspect of the present disclosure, there is provided a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the above-mentioned model determination method of any one is implemented.

[0053] According to yet another aspect of the present disclosure, there is provided a computer program product, which includes a computer program or computer instructions, and the computer program or computer instructions are loaded and executed by a processor to enable a computer to implement the above-mentioned large language model determination method of any one.

[0054] In the embodiment of the present disclosure, in the data collection stage, a preference data set is constructed for multiple dimensions, and at the large language model training node, a loss function corresponding to the large language model is constructed based on the probability of generating each dimension of data, and adaptive multi-objective training is performed on the preference data set until the loss function converges to obtain a target large language model. The obtained target large language model can improve the balance of responses in multiple dimensions.

[0055] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. Description of the Drawings

[0056] The accompanying drawings here are incorporated into and constitute a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure. Obviously, the accompanying drawings in the following description are only some embodiments of the present disclosure, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.

[0057] Figure 1 Schematic diagram of a large language model determination system structure in an embodiment of the present disclosure.

[0058] Figure 2 Flowchart of a method for determining a large language model in an embodiment of the present disclosure.

[0059] Figure 3 Flowchart of another method for determining a large language model in an embodiment of the present disclosure.

[0060] Figure 4 Flowchart of yet another method for determining a large language model in an embodiment of the present disclosure.

[0061] Figure 5 Flowchart of still another method for determining a large language model in an embodiment of the present disclosure.

[0062] Figure 6 Flowchart of still another method for determining a large language model in an embodiment of the present disclosure.

[0063] Figure 7 Schematic diagram of a device for determining a large language model in an embodiment of the present disclosure.

[0064] Figure 8 Block diagram of the structure of an electronic device in an embodiment of the present disclosure.

[0065] Figure 9 Schematic diagram of a computer-readable storage medium in an embodiment of the present disclosure. Detailed implementation

[0066] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be more comprehensive and complete, and will fully convey the concept of the example embodiments to those skilled in the art. The features, structures, or characteristics described can be combined in any suitable manner in one or more embodiments.

[0067] In addition, the accompanying drawings are only schematic illustrations of the present disclosure and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and thus repeated descriptions thereof will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities may be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0068] It should be understood that the steps described in the method embodiments of the present disclosure may be executed in a different order and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this regard.

[0069] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0070] It should be noted that the modifiers "a" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly stated in the context, it should be understood as "one or more".

[0071] In order to improve the balance of the responses of large language models in multiple dimensions, the following multiple solutions are provided in the related art. One of them is the instruction control method, which controls each preference by inserting control tokens in the prompt, but this method has poor adaptability. Another is the multi-model integration method, which optimizes each preference dimension by training multiple reward models, but this method has high complexity.

[0072] To solve the above problems, the present disclosure provides a large language model determination method, apparatus and computer program product.

[0073] It should be noted that, without conflict, the embodiments of the present disclosure and the technical features in the embodiments may be combined with each other.

[0074] The following will describe in detail the specific implementation methods of the embodiments of the present disclosure with reference to the accompanying drawings.

[0075] Figure 1 A schematic structural diagram of a large language model determination system in an embodiment of the present disclosure is shown. This system can be applied to the model determination method and model determination apparatus in various embodiments of the present disclosure.

[0076] As Figure 1As shown, the model determination system 10 may include a front-end display module 101 and a model processing module 102. In some embodiments, the front-end display module 101 and the model processing module 102 may be located on different devices. The front-end display module 101 may be a module on an electronic device with data acquisition capabilities. The model processing module 102 may be a module on an electronic device with data processing capabilities such as a computer or a server. The front-end display module 101 and the model processing module 102 may be located on the same device. For example, the front-end display module 101 and the model processing module 102 may be an input module and a processing module on a computer or a mobile phone.

[0077] A network realizes a communication connection between the front-end display module 101 and the model processing module 102. This network may be a wired network or a wireless network.

[0078] Optionally, the above-mentioned wireless network or wired network uses standard communication technologies and / or protocols. The network is usually the Internet, but it can also be any network, including but not limited to any combination of a local area network (LAN), a metropolitan area network (MAN), a wide area network (WAN), a mobile, wired or wireless network, a private network or a virtual private network. In some embodiments, technologies and / or formats including Hyper Text Mark-up Language (HTML), Extensible Markup Language (XML), etc. are used to represent the data exchanged through the network. In addition, conventional encryption technologies such as Secure Socket Layer (SSL), Transport Layer Security (TLS), Virtual Private Network (VPN), Internet Protocol Security (IPsec), etc. may be used to encrypt all or some of the links. In other embodiments, customized and / or dedicated data communication technologies may be used to replace or supplement the above data communication technologies.

[0079] The case where the front-end display module 101 and the model processing module 102 are located on two different devices will be described below.

[0080] The front-end display module 101 may be located on a terminal device, and the terminal device may be various electronic devices, including but not limited to smart phones, tablet computers, laptop portable computers, desktop computers, wearable devices, augmented reality devices, virtual reality devices, etc.

[0081] Optionally, the clients of the application programs installed in different terminal devices are the same, or the clients of the same type of application programs based on different operating systems. Depending on the different terminal platforms, the specific form of the client of the application program can also be different. For example, the client of the application program can be a mobile client, a PC client, etc.

[0082] The model processing module 102 can be located on a server, which can be a server that provides various services. For example, it can be a background management server that supports the devices for the operations performed by users using terminal devices. The background management server can analyze and process data such as requests received, and feedback the processing results to the terminal devices.

[0083] Optionally, the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), as well as big data and artificial intelligence platforms.

[0084] Figure 2 The flowchart of a method for determining a large language model in an embodiment of the present disclosure is shown. As Figure 2 shown, the method for determining a large language model in an embodiment of the present disclosure may include:

[0085] S210, in the data collection phase, construct a preference data set for multiple dimensions.

[0086] In some embodiments, the preference data set may include prompt words, responses, and multi-dimensional data for each response in multiple dimensions.

[0087] The prompt words may include the information input by the user to the large language model. A prompt can be a keyword, phrase, or sentence used to guide the model to generate a specific type or style of text, perform a specific task, or provide specific information.

[0088] In some embodiments, the response may include the reply of the large language model to the input prompt words. The response can be a specific type or style of text generated by the model, a keyword, phrase, or sentence for performing a specific task or specific information.

[0089] In some embodiments, the multi-dimensional data may include dimensions and dimension scores. Among them, the dimensions may include the perspectives for evaluating responses. Exemplarily, the multiple dimensions may be helpfulness, correctness, and instruction compliance respectively. The dimension scores may be the helpfulness score, the correctness score, and the instruction compliance score respectively.

[0090] In some embodiments, the prompt can be input into the large language model to obtain multiple responses, and then an instruction can be sent to the large language model to enable the large language model to perform multi-dimensional evaluation on the responses and obtain multiple dimension data.

[0091] Exemplarily, the prompt can be: How many r in strawberry How many "r"s are there in the word "strawberry".

[0092] Inputting the above prompt into the large language model, the first response obtained is: Counting: 1.s 2.t 3.r 4.a 5.w 6.b 7.e 8.r 9.r 10.y There are 3 “r”s in the word “strawberry”. The second response obtained is: there are 2 “r”s in the word “strawberry”.

[0093] Then input the above prompt and the two responses corresponding to the prompt into the large language model, and input respectively: Please score according to the scoring criteria in terms of helpfulness, Please score according to the scoring criteria in terms of correctness, and Please score according to the scoring criteria in terms of instruction compliance. to guide the large language model to generate multiple dimension data.

[0094] S220, in the large language model training stage, based on the probability of generating each dimension data, construct the loss function corresponding to the large language model, and perform adaptive multi-objective training on the preference dataset until the loss function converges to obtain the target large language model.

[0095] In some embodiments, the probability of generating each dimension data can be determined based on the large language model.

[0096] In some embodiments, constructing the loss function corresponding to the large language model based on the probability of generating each dimension data may include determining the loss function based on the mean and variance of the probability of generating each dimension data.

[0097] In some embodiments, constructing the loss function corresponding to the large language model based on the probability of each dimension data can use the generation metric as an implicit reward, so that the large language model constructed based on the above loss function can inherently understand and align the meanings of different dimensions.

[0098] In some embodiments, the loss function is a function used to measure the difference between the prediction result of the large language model and the true value, and is used to optimize the large language model.

[0099] In some embodiments, constructing the loss function corresponding to the large language model based on the probability of generating data for each dimension can enable the target large language model optimized based on the above loss function to align multiple dimensions.

[0100] In some embodiments, the adaptive multi-objective training may include inputting the prompt into the model under training to obtain the reply output by the model under training and the multi-dimensional data corresponding to each reply, and then comparing the output of the model with the multiple replies in the training samples and the multi-dimensional data corresponding to each reply based on the loss function. Among them, the output of the model includes the reply output by the model and the multi-dimensional data corresponding to each reply.

[0101] In the embodiments of the present disclosure, in the data collection stage, a preference data set is constructed for multiple dimensions. At the large language model training node, the loss function corresponding to the large language model is constructed based on the probability of generating data for each dimension, and adaptive multi-objective training is performed on the preference data set until the loss function converges to obtain the target large language model. The obtained target large language model can improve the balance of the reply in multiple dimensions.

[0102] Figure 3 Show the flowchart of another method for determining a large language model in the embodiments of the present disclosure. As Figure 3 shown, the method for determining a large language model in the embodiments of the present disclosure may include:

[0103] S310, when the loss function value corresponding to the loss function has not converged, determine the weight corresponding to each dimension based on the probability of generating data for each dimension and the Gaussian sampling model.

[0104] In some embodiments, the convergence of the loss function value may include that the difference between the loss function values obtained continuously for multiple times is less than a preset threshold.

[0105] In some embodiments, S310 may include:

[0106] Determine the probability mean and probability variance of generating data for each dimension based on the probability of generating data for each dimension;

[0107] Construct a Gaussian distribution model generation space based on the probability mean and probability variance of generating data for each dimension;

[0108] Perform Gaussian weight sampling on the Gaussian distribution model generation space to obtain the weight corresponding to each dimension.

[0109] In some embodiments, after determining the probability mean and the probability variance, the Gaussian distribution of the probability can be determined.

[0110] In some embodiments, Gaussian sampling refers to the process of randomly extracting samples from a Gaussian distribution.

[0111] Exemplarily, Gaussian sampling of the Gaussian distribution model can be performed based on the Box-Muller transform, the Ziggurat algorithm, and the central limit theorem.

[0112] In some embodiments, after Gaussian sampling, the sampled data can be used as the weight corresponding to each dimension.

[0113] In some embodiments, after Gaussian sampling, the sampled data can be used as the weight corresponding to each dimension after mathematical transformation.

[0114] S320. Adjust the weight corresponding to each dimension to reduce the difference between the multi-dimensional data corresponding to the reply output by the large language model.

[0115] In some embodiments, the weight corresponding to each dimension can be adjusted according to a preset period, or the weight corresponding to each dimension can be adjusted in real time based on user definition.

[0116] In some embodiments, the dimension data includes dimension scores, and reducing the difference between each dimension data can include reducing the difference between the dimension scores between each dimension data.

[0117] In some embodiments, the weight corresponding to each dimension can be adjusted based on the dimension data corresponding to each dimension. Exemplarily, the weight corresponding to each dimension can be adjusted based on the inverse proportion relationship and the dimension data corresponding to each dimension.

[0118] S330. Update the loss function based on the adjusted weight corresponding to each dimension.

[0119] In some embodiments, the loss function is composed of the weight corresponding to each dimension and the probability mean of generating each dimension data.

[0120] Exemplarily, the loss function can be:

[0121]

[0122] Where x is the prompt, y a is the first reply, y b is the second reply, m, h, c, if are different dimensions, πθ is the large language model, β is the temperature parameter used to scale the difference between the implicit rewards of the preferred reply, is the prompt of different dimensions reconstructed corresponding to the first reply, The prompting words for different dimensions after reconstructing the second response, α {i} is the weight factor for different dimensions.

[0123] In some embodiments, in order to enable the constructed target large language model to obtain corresponding rewards when generating multi-dimensional balanced responses, the present disclosure constructs an objective function to obtain the optimal solution. The present disclosure sets the weight vector α = [α1,..., α k , and the weight vector is obtained through the policy α ∼ ρ, where and α k > 0. Then, the Pareto optimal solution is determined by solving a scalar optimization problem:

[0124]

[0125] where is the preference consistency function for the k-th dimension, y w is the winning response, y l is the losing response. x is the prompting word, and x k represents the modified prompting word by focusing on the k-th target, ρ represents sampling, that is, sampling a value from the distribution, and π represents the output of the target large language model.

[0126] To effectively optimize and balance multiple goals, the present disclosure adopts dynamic weight sampling. Different from the traditional method of fusing multiple dimension scores into a single score, the target large language model of the present disclosure adopts the sum of dynamic weighted scores across different dimensions to achieve balanced scores among multiple dimensions and realize adaptive goal optimization.

[0127]

[0128] In some embodiments, α is derived from the weight policy ρ, is the probability that the output y w of the target large language model is greater than y l , r k is the scoring model, and f represents the input x and dimension d, and outputs the rewrite of x with respect to the d dimension.

[0129] In the embodiments of the present disclosure, the average value and variance of each dimension k can be calculated based on the token generation probability in the target large language model. These data are used to parameterize a Gaussian distribution ρ k from which the weight α k is sampled, and the average value μ k and the variance can be:

[0130]

[0131] where is the probability of the t-th token in dimension k.

[0132] The formula for the log-likelihood of generating output sequence y given output sequence x can be:

[0133]

[0134] Combining formulas (3), (4), and (5) gives the loss function of the present disclosure, i.e., formula (1).

[0135] In some embodiments, since the loss function is constructed from the weights corresponding to each dimension and the probability mean of the data in each dimension, the loss function can be updated according to the adjusted weights corresponding to each dimension.

[0136] In the embodiments of the present disclosure, by determining the probability of the data in each dimension to determine the weights corresponding to each dimension, adjusting the weights corresponding to each dimension to reduce the difference between the multi-dimensional data corresponding to the reply output by the model, and updating the loss function based on the adjusted weights corresponding to each dimension can make the reply output by the large language model trained based on the loss function multi-dimensionally balanced.

[0137] Figure 4 Shows a flowchart of another large language model determination method in the embodiments of the present disclosure. As Figure 4 shown, the model determination method in the embodiments of the present disclosure may include:

[0138] S410, in the data collection stage, construct a preference dataset for multiple dimensions;

[0139] S420, construct a loss function based on the weights corresponding to each dimension and the probability mean of generating the data in each dimension, so that when the loss function value corresponding to the loss function gradually converges, the difference between the multi-dimensional data decreases;

[0140] S430, perform adaptive multi-objective training on the preference dataset until the loss function converges to obtain the target large language model.

[0141] In some embodiments, the adjusted model parameters may include: the weights and bias terms of the model, the learning rate, the regularization parameter, the batch size, the optimizer parameters, the network structure parameters, and the initialization parameters.

[0142] In the embodiments of the present disclosure, when the loss function corresponding to the model does not reach the preset threshold, adjusting the model parameters can reduce the difference between the multi-dimensional data corresponding to the reply output by the model, making the reply output by the model more balanced in multiple dimensions.

[0143] Figure 5The flowchart of yet another large language model determination method in an embodiment of the present disclosure is shown. As Figure 5 shown, the large language model determination method in an embodiment of the present disclosure may include:

[0144] S510, dividing multiple responses into winning responses and losing responses based on multiple-dimensional data corresponding to each response.

[0145] In some embodiments, it may be determined whether a response is a winning response or a losing response based on the magnitude of the dimensional score.

[0146] In some embodiments, since each response can correspond to scores in multiple dimensions, the scores in multiple dimensions can be weighted averaged and sorted according to the magnitude of the weighted average value, and the response corresponding to the weighted average value with a higher ranking in the sorted sequence is determined as the winning response.

[0147] In some embodiments, the prompt words may include prompt words corresponding to multiple dimensions.

[0148] S520, splicing the multi-dimensional prompts with the winning responses and the losing responses respectively to obtain multiple reconstructed losing prompt words and winning prompt words.

[0149] In some embodiments, the method of using the target prompt word, the target response corresponding to the target prompt word, and the target dimensional data corresponding to each target response as training samples may be the same as the method of generating training samples based on the prompt word, multiple responses corresponding to the prompt word, and multiple-dimensional data corresponding to each response.

[0150] In some embodiments, the constructed winning prompt words can be: Please pay attention to helpfulness and score according to the scoring criteria. Counting: 1.s 2.t 3.r 4.a 5.w 6.b 7.e 8.r 9.r 10.y There are 3 “r” in the word “strawberry”. How many points can this reply get? Please pay attention to correctness and score according to the scoring criteria. Counting: 1.s 2.t 3.r 4.a 5.w 6.b 7.e 8.r 9.r 10.y There are 3 “r” in the word “strawberry”. How many points can this reply get? And please pay attention to instruction following and score according to the scoring criteria. Counting: 1.s 2.t 3.r 4.a 5.w 6.b 7.e 8.r 9.r 10.y There are 3 “r” in the word “strawberry”. How many points can it get? The constructed losing prompt words can be: Please pay attention to helpfulness and score according to the scoring criteria. There are 2 “r” in the word “strawberry”. How many points can this reply get? Please pay attention to correctness and score according to the scoring criteria. There are 2 “r” in the word “strawberry”. How many points can this reply get? And please pay attention to instruction following and score according to the scoring criteria. There are 2 “r” in the word “strawberry”. How many points can this reply get? Input the above reconstructed prompt words into the target large language model according to the dimensions respectively, so that the target large language model outputs replies with higher scores. Exemplarily, in the correctness dimension, the sentences “Please pay attention to correctness and score according to the scoring criteria. Counting: 1.s 2.t 3.r 4.a 5.w 6.b 7.e 8.r 9.r 10.y There are 3 “r” in the word “strawberry”. How many points can this reply get” and “Please pay attention to correctness and score according to the scoring criteria. There are 2 “r” in the word “strawberry”. How many points can this reply get”, as well as the replies with higher scores output, are input into the large language model, and then the probability of the winning reply in the output result of the large language model is determined. The probability of the obtained winning reply in the output result of the large language model is the probability in the correctness dimension.

[0151] In the embodiments of the present disclosure, multiple responses are divided into winning responses and losing responses based on multi-dimensional data corresponding to each response. Then, multi-dimensional prompt words are respectively concatenated with the winning responses and the losing responses to obtain multiple reconstructed losing prompt words and winning prompt words. The above-mentioned prompt words are input into the large language model to obtain the responses output by the large language model, so as to determine the probability of each dimension data, making the determined probability of the dimension data more accurate.

[0152] Figure 6 Show a flowchart of another large language model determination method in the embodiments of the present disclosure. As Figure 6 shown, the model determination method in the embodiments of the present disclosure may include:

[0153] S610, in the data collection stage, construct a preference data set for multiple dimensions;

[0154] In the large language model training stage, construct a loss function corresponding to the large language model based on the probability of generating each dimension data, and perform adaptive multi-objective training on the preference data set until the loss function converges to obtain the target large language model.

[0155] S620, in the large language model training stage, construct a loss function corresponding to the large language model based on the probability of generating each dimension data, and perform adaptive multi-objective training on the preference data set until the loss function converges to obtain the target large language model.

[0156] S630, in response to a user inputting a question at the input end of the large language model, display a multi-dimensionally balanced response corresponding to the question.

[0157] In some embodiments, after the large language model training is completed, the preference large language model can be set in an electronic device. After the user inputs a question, the electronic device inputs the input question to the input end of the preference large language model, and then the preference large language model processes the input question to obtain a multi-dimensionally balanced response. The electronic device displays the above response.

[0158] Based on the same inventive concept, an apparatus for determining a large language model is also provided in the embodiments of the present disclosure, as described in the following embodiments. Since the principle of solving problems in this apparatus embodiment is similar to that of the above method embodiment, the implementation of this apparatus embodiment can refer to the implementation of the above method embodiment, and the repeated parts will not be described again.

[0159] Figure 7 Show a schematic diagram of an apparatus for determining a large language model in the embodiments of the present disclosure, as Figure 7 shown, the apparatus 700 for determining the large language model includes:

[0160] A construction module 710, configured to construct a preference data set for multiple dimensions in the data collection stage;

[0161] A training module 720, configured to, in the large language model training stage, construct a loss function corresponding to the large language model based on the probability of generating data for each dimension, perform adaptive multi-objective training on the preference dataset until the loss function converges, and obtain the target large language model.

[0162] In an embodiment of the present disclosure, the apparatus further includes:

[0163] A first determination module, configured to, when the loss function value corresponding to the loss function has not converged, determine the weight corresponding to each dimension based on the probability of generating data for each dimension and the Gaussian sampling model;

[0164] An adjustment module, configured to adjust the weight corresponding to each dimension to reduce the difference between the multi-dimensional data corresponding to the reply output by the large language model;

[0165] An update module, configured to update the loss function based on the adjusted weight corresponding to each dimension.

[0166] In an embodiment of the present disclosure, the first determination module includes:

[0167] A determination unit, configured to determine the probability mean and probability variance of generating data for each dimension based on the probability of generating data for each dimension;

[0168] A first construction unit, configured to construct a Gaussian distribution model generation space based on the probability mean and probability variance of generating data for each dimension;

[0169] A sampling unit, configured to perform Gaussian weight sampling on the Gaussian distribution model generation space to obtain the weight corresponding to each dimension.

[0170] In an embodiment of the present disclosure, the training module includes:

[0171] A second construction unit, configured to construct a loss function based on the weight corresponding to each dimension and the probability mean of generating data for each dimension, so that when the loss function value corresponding to the loss function gradually converges, the difference between the multi-dimensional data is reduced.

[0172] In an embodiment of the present disclosure, the apparatus further includes:

[0173] An adjustment module, configured to, when the loss function value corresponding to the loss function has not converged, adjust the model parameters to reduce the difference between the multi-dimensional data corresponding to the reply output by the large language model.

[0174] In an embodiment of the present disclosure, the construction module further includes:

[0175] A division unit for dividing multiple responses into winning responses and losing responses based on multi-dimensional data corresponding to each response;

[0176] A splicing unit for splicing multi-dimensional prompts with winning responses and losing responses respectively to obtain multiple reconstructed losing prompt words and winning prompt words.

[0177] In an embodiment of the present disclosure, the apparatus further includes:

[0178] A second determination module for inputting the multiple reconstructed losing prompt words and winning prompt words into a large language model respectively, and determining the probability of generating each dimension of data based on the probability of the winning prompt words in the responses of each dimension output by the large language model.

[0179] In an embodiment of the present disclosure, the apparatus further includes:

[0180] A display module for displaying a multi-dimensionally balanced response corresponding to a question in response to the user inputting the question at the input end of the large language model.

[0181] In the implementation manner of the present disclosure, in the data collection stage, a preference data set is constructed for multiple dimensions, and at the large language model training node, a loss function corresponding to the large language model is constructed based on the probability of generating each dimension of data, and adaptive multi-objective training is performed on the preference data set until the loss function converges to obtain a target large language model. The obtained target large language model can improve the balance of responses in multiple dimensions.

[0182] Those skilled in the art can understand that various aspects of the present disclosure can be implemented as a system, a method, or a program product. Therefore, various aspects of the present disclosure can be specifically implemented in the following forms, namely: a complete hardware implementation manner, a complete software implementation manner (including firmware, microcode, etc.), or an implementation manner combining hardware and software aspects, which can be collectively referred to as "circuit", "module", or "system" here.

[0183] An electronic device provided by the present disclosure includes: a processor; and a memory for storing executable instructions of the processor; wherein, the processor is configured to execute the above-mentioned display method by executing the executable instructions.

[0184] Exemplarily, the following refers to Figure 8 to describe the electronic device 800 according to this implementation manner of the present disclosure. Figure 8 The displayed electronic device 800 is only an example and should not bring any limitation to the functions and usage scope of the embodiments of the present disclosure.

[0185] As Figure 8As shown, the electronic device 800 is presented in the form of a general-purpose computing device. The components of the electronic device 800 may include, but are not limited to: at least one of the above-mentioned processing units 810, at least one of the above-mentioned storage units 820, and a bus 830 connecting different system components (including the storage unit 820 and the processing unit 810).

[0186] Among them, the storage unit stores program code, which can be executed by the processing unit 810, so that the processing unit 810 executes the steps according to various exemplary embodiments of the present disclosure described in the "Exemplary Method" section of this specification. For example, the processing unit 810 may execute the following steps of the above method embodiments:

[0187] In the data collection stage, a preference dataset is constructed for multiple dimensions;

[0188] In the large language model training stage, a loss function corresponding to the large language model is constructed based on the probability of generating data for each dimension, and adaptive multi-objective training is performed on the preference dataset until the loss function converges to obtain the target large language model.

[0189] The storage unit 820 may include a readable medium in the form of a volatile storage unit, such as a random access storage unit (RAM) 8201 and / or a cache storage unit 8202, and may further include a read-only storage unit (ROM) 8203.

[0190] The storage unit 820 may further include a program / utility 8204 having a set (at least one) of program modules 8205. Such program modules 8205 include, but are not limited to: an operating system, one or more application programs, other program modules, and program data. The implementation of a network environment may be included in each or some combination of these examples.

[0191] The bus 830 may represent one or more of several types of bus structures, including a storage unit bus or a storage unit controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any bus structure in a variety of bus structures.

[0192] The electronic device 800 can also communicate with one or more external devices 840 (such as a keyboard, a pointing device, a Bluetooth device, etc.), and can also communicate with one or more devices that enable a user to interact with the electronic device 800, and / or communicate with any device that enables the electronic device 800 to communicate with one or more other computing devices (such as a router, a modem, etc.). Such communication can be carried out through the input / output (I / O) interface 850. Moreover, the electronic device 800 can also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through the network adapter 860. As shown in the figure, the network adapter 860 communicates with other modules of the electronic device 800 through the bus 830. It should be understood that although not shown in the figure, other hardware and / or software modules can be used in combination with the electronic device 800, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.

[0193] Through the description of the above embodiments, those skilled in the art can easily understand that the exemplary embodiments described herein can be implemented by software, or can be implemented by the way of software combined with necessary hardware. Therefore, the technical solutions according to the embodiments of the present disclosure can be embodied in the form of a software product, and the software product can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which can be a personal computer, a server, a terminal device, or a network device, etc.) to execute the methods according to the embodiments of the present disclosure.

[0194] In the disclosed exemplary embodiments, a computer-readable storage medium is also provided, and the computer-readable storage medium can be a readable signal medium or a readable storage medium. Figure 9 The schematic diagram of a computer-readable storage medium in the embodiments of the present disclosure is shown, as Figure 9 shown, a program product capable of implementing the above methods of the present disclosure is stored on the computer-readable storage medium 900.

[0195] In some possible implementation manners, various aspects of the present disclosure can also be implemented in the form of a program product, which includes program code. When the program product runs on a terminal device, the program code is used to enable the terminal device to execute the steps according to various exemplary embodiments of the present disclosure described in the above "Specific Embodiments" section of this specification.

[0196] More specific examples of the computer-readable storage medium in the present disclosure may include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above.

[0197] In the present disclosure, the computer-readable storage medium may include a data signal propagated in a baseband or as part of a carrier wave, in which the readable program code is carried. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The readable signal medium may also be any readable medium other than the readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0198] Optionally, the program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to wireless, wired, optical fiber cable, RF, etc., or any suitable combination of the above.

[0199] In specific implementation, the program code for performing the operations of the present disclosure may be written in any combination of one or more programming languages. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, executed as an independent software package, partially on the user computing device and partially on a remote computing device, or entirely on the remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., by connecting through the Internet using an Internet service provider).

[0200] The embodiments of the present disclosure provide a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the display method provided in any of the various alternative manners in the embodiments of the present disclosure.

[0201] It should be noted that although several modules or units of the device for action execution are mentioned in the above detailed description, such a division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more of the above-described modules or units can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0202] Through the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software or by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several instructions to enable a computing device (such as a personal computer, a server, a mobile terminal, or a network device, etc.) to execute the method according to the embodiments of the present disclosure.

[0203] After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily conceive of other embodiments of the present disclosure. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the technical field not disclosed by the present disclosure. The specification and examples are only illustrative, and the true scope of the present disclosure is pointed out by the appended claims.

Claims

1. A method for determining a large language model, characterized in that: include: During the data collection phase, preference datasets were constructed for multiple dimensions; In the large language model training stage, a loss function corresponding to the large language model is constructed based on the probability of generating data in each dimension, and adaptive multi-objective training is performed on the preferred data set until the loss function converges to obtain the target large language model.

2. The method according to claim 1, characterized in that The method further comprises: When the loss function value corresponding to the loss function has not converged, determining the weight corresponding to each dimension based on the probability of generating data of each dimension and the Gaussian sampling model; Adjusting the weight corresponding to each dimension so as to reduce the difference between the multiple dimensional data corresponding to the response output by the large language model; The loss function is updated based on the adjusted weight corresponding to each dimension.

3. The method according to claim 2, characterized in that The determining the weight corresponding to each dimension based on the probability of generating data of each dimension and the Gaussian sampling model includes: Determine the probability mean and probability variance of generating each dimensional data based on the probability of generating each dimensional data; Constructing a Gaussian distribution model generation space based on the probability mean and probability variance of generating each dimensional data; Gaussian weight sampling is performed on the Gaussian distribution model generation space to obtain the weight corresponding to each dimension.

4. The method according to claim 3, characterized in that The loss function corresponding to the large language model is constructed based on the probability of generating each dimension of data, including: The loss function is constructed based on the weight corresponding to each dimension and the probability mean of generating each dimensional data, so that the difference between multiple dimensional data is reduced when the loss function value corresponding to the loss function gradually converges.

5. The method according to claim 1, characterized in that The method further comprises: When the loss function value corresponding to the loss function has not converged, the model parameters are adjusted to reduce the difference between the multiple dimensional data corresponding to the responses output by the large language model.

6. The method according to claim 1, characterized in that The preference data set includes a prompt word, multiple replies corresponding to the prompt word, and multiple dimension data corresponding to each reply. The construction of the preference data set for multiple dimensions includes: Based on multiple dimension data corresponding to each reply, multiple replies are divided into victory replies and failure replies; The multi-dimensional prompt words are respectively concatenated with the victory response and the failure response to obtain a plurality of reconstructed failure prompt words and victory prompt words.

7. The method according to claim 6, characterized in that The method further comprises: The reconstructed multiple failure prompt words and victory prompt words are respectively input into the large language model, and the probability of generating each dimension of data is determined based on the probability of the victory prompt word in the response of each dimension output by the large language model.

8. The method according to claim 6, characterized in that The method further comprises: In response to a user inputting a question at the large language model input terminal, a multi-dimensional balanced reply corresponding to the question is displayed.

9. A large language model determination device, characterized in that: include: The construction module is used to construct preference datasets for multiple dimensions during the data collection phase; The training module is used to construct a loss function corresponding to the large language model based on the probability of generating data of each dimension during the large language model training phase, and to perform adaptive multi-objective training on the preferred data set until the loss function converges to obtain a target large language model.

10. A computer program product, comprising a computer program or computer instructions, characterized in that: The computer program or the computer instruction is loaded and executed by a processor so that the computer implements the large language model determination method as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Natural language processing method and device based on multiple tools and medium

    CN114580387A

  • Sensitive question identification method and device based on large language model, equipment and medium

    CN118821789A

  • Large language model alignment method and device, electronic equipment and readable storage medium

    CN119513306A

  • System and method for testing machine learning

    US20210319338A1