Generative pre-training model GPT configuration method, apparatus and device

By combining the actor model and reference model, or critic model and reward model and sharing model parameters, the problem of high storage resources consumption when training large language models based on PPO algorithm is solved, and more efficient storage utilization is achieved.

CN120146061APending Publication Date: 2025-06-13QINGDAO HAIER TECH +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311715870.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-13
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

When training large language models based on PPO algorithm, multiple models (actor model, reference model, reward model, critic model) need to be loaded, resulting in a large consumption of storage resources.

Method used

By merging the actor model and reference model, or merging the critic model and reward model, the merged model shares model parameters, thereby reducing the consumption of storage resources.

Benefits of technology

When training the target processing model based on the PPO algorithm, the storage space occupation is reduced and the loading rate is improved by merging the relevant models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120146061A_ABST
    Figure CN120146061A_ABST
Patent Text Reader

Abstract

The invention discloses a generative pre-training model GPT configuration method, device and equipment, and the method comprises the steps: obtaining sample data used for training a target processing model; the target processing model comprises an actor model, a reviewer model, a reference model corresponding to the actor model, and a reward model corresponding to the reviewer model; processing the sample data using a first model to obtain a first loss, the first model referring to any one of a reference model and a reward model; processing the output of the first model by using the first sub-network to obtain a second loss, the first model and the first sub-network forming a second model, the second model being one of the actor model and the commentator model corresponding to the first model; determining the total loss of the model according to the first loss, the second loss and the loss of other models in the target processing model; and training at least part of parameters of the target processing model according to the total model loss.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of data processing, and particularly relates to a method, apparatus, and device for configuring a Generative Pretrained Transformer (GPT). Background Art

[0002] With the development and application of large language models, how to make the generative pre-trained model more in line with the human value system is related to the application value and scope of the model. At present, a solution to this problem is to optimize the large language model using the Reinforcement learning with human feedback (RLHF) technique. Among them, the Proximal Policy Optimization (PPO) algorithm in deep reinforcement learning is the key algorithm for implementing RLHF.

[0003] When training a large language model based on the PPO algorithm, it is necessary to combine the actor model, the ref model, the reward model, and the critic model, and train with reference to the outputs of these four models. Therefore, these four models need to be loaded into the storage space of the computer device used for training during the training process, resulting in relatively high storage resource consumption when training a large language model based on the PPO algorithm. Summary of the Invention

[0004] To this end, this application discloses a method, apparatus, and device for configuring a Generative Pretrained Transformer (GPT) to reduce the storage resources consumed when training a large language model based on the PPO algorithm.

[0005] The first aspect of this application provides a method for configuring a Generative Pretrained Transformer (GPT), including:

[0006] Obtaining sample data for training a target processing model; wherein, the target processing model includes an actor model, a critic model, a ref model corresponding to the actor model, and a reward model corresponding to the critic model;

[0007] Processing the sample data using a first model to obtain a first loss, where the first model refers to any one of the ref model and the reward model, and the first model is a pre-trained model;

[0008] Processing the output of the first model using a first sub-network to obtain a second loss, where the first model and the first sub-network form a second model, and the second model is the one corresponding to the first model among the actor model and the critic model;

[0009] Determine the total model loss based on the first loss, the second loss, the third loss obtained by a third model, and the fourth loss obtained by a fourth model; wherein, the third model refers to the one of the reference model and the reward model that is different from the first model, and the fourth model refers to the one of the actor model and the critic model that does not correspond to the second model;

[0010] Train at least part of the parameters of the target processing model according to the total model loss.

[0011] Optionally, the process of obtaining the third loss and the fourth loss includes:

[0012] Process the sample data using the third model to obtain the third loss;

[0013] Process the output of the third model using a second sub-network to obtain the fourth loss, where the third model and the second sub-network form the fourth model.

[0014] Optionally, the first model includes a shared network unit and a first network unit, and the third model includes the shared network unit and a third network unit;

[0015] The process of using the first model to process the sample data to obtain the first loss includes:

[0016] Process the sample data using the shared network unit;

[0017] Process the output of the shared network unit using the first network unit to obtain the first loss;

[0018] The process of using the third model to process the sample data to obtain the third loss includes:

[0019] Process the output of the shared network unit using the third network unit to obtain the third loss.

[0020] Optionally, the process of training at least part of the parameters of the target processing model according to the total model loss includes:

[0021] Train the parameters of the first sub-network and the second sub-network according to the total model loss.

[0022] A second aspect of the present application provides a Generative Pretrained Transformer (GPT) configuration device, including:

[0023] An acquisition unit, configured to acquire sample data for training a target processing model; wherein, the target processing model includes an actor model, a critic model, a reference model corresponding to the actor model, and a reward model corresponding to the critic model;

[0024] A processing unit for processing the sample data using a first model to obtain a first loss, where the first model refers to either the reference model or the reward model, and the first model is a pre-trained model;

[0025] The processing unit is configured to process the output of the first model using a first sub-network to obtain a second loss, where the first model and the first sub-network form a second model, and the second model is the one corresponding to the first model among the actor model and the critic model;

[0026] A determination unit for determining a total model loss based on the first loss, the second loss, a third loss obtained by a third model, and a fourth loss obtained by a fourth model; wherein the third model refers to the one different from the first model among the reference model and the reward model, and the fourth model refers to the one not corresponding to the second model among the actor model and the critic model;

[0027] A training unit for training at least part of the parameters of the target processing model according to the total model loss.

[0028] Optionally, the process by which the processing unit obtains the third loss and the fourth loss includes:

[0029] Processing the sample data using the third model to obtain the third loss, where the third model is a pre-trained model;

[0030] Processing the output of the third model using a second sub-network to obtain the fourth loss, where the third model and the second sub-network form the fourth model.

[0031] Optionally, the first model includes a shared network unit and a first network unit, and the third model includes the shared network unit and a third network unit;

[0032] When the processing unit processes the sample data using the first model to obtain the first loss, it is specifically configured to:

[0033] Process the sample data using the shared network unit;

[0034] Process the output of the shared network unit using the first network unit to obtain the first loss;

[0035] When the processing unit processes the sample data using the third model to obtain the third loss, it is specifically configured to:

[0036] Process the output of the shared network unit using the third network unit to obtain the third loss.

[0037] Optionally, when the training unit trains at least some parameters of the target processing model according to the total model loss, it is specifically configured to:

[0038] Train the parameters of the first sub-network and the second sub-network according to the total model loss.

[0039] A third aspect of the present application provides an electronic device, including a memory and a processor;

[0040] The memory is used to store a computer program;

[0041] The processor is used to execute the computer program, and is specifically configured to implement the generative pre-training model GPT configuration method provided in any item of the first aspect of the present application.

[0042] A fourth aspect of the present application provides a computer-readable storage device for storing a computer program, which when executed, is specifically configured to implement the generative pre-training model GPT configuration method provided in any item of the first aspect of the present application.

[0043] The beneficial effects of this solution are as follows:

[0044] When training the target processing model based on the PPO algorithm, at least the actor model and the corresponding reference model are merged, or the critic model and the corresponding reward model are merged, so that the merged models share model parameters. Description of the Drawings

[0045] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.

[0046] Figure 1 It is a flowchart of a generative pre-training model GPT configuration method provided by an embodiment of the present application;

[0047] Figure 2 It is a schematic diagram of the architecture of a large language model provided by an embodiment of the present application;

[0048] Figure 3 It is a schematic diagram of the architecture of another large language model provided by an embodiment of the present application;

[0049] Figure 4 It is a schematic diagram of the architecture of yet another large language model provided by an embodiment of the present application;

[0050] Figure 5It is a schematic structural diagram of a generative pre-trained model GPT configuration device provided by an embodiment of the present application;

[0051] Figure 6 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Specific embodiments

[0052] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0053] This embodiment provides a method for configuring a generative pre-trained model GPT. Please refer to Figure 1 ., which is a flowchart of this method. This method may include the following steps.

[0054] S101, Obtain sample data for training the target processing model.

[0055] Among them, the target processing model includes an actor model, a critic model, a reference model corresponding to the actor model, and a reward model corresponding to the critic model.

[0056] The target processing model in this embodiment may be a large language model based on RLHF.

[0057] The sample data may include multiple pieces of prompt text obtained through various methods such as manual editing and algorithm generation. Each piece of prompt text corresponds to a pre-annotated label result text.

[0058] When using the sample data to train the target processing model, the inputs of the actor model and the reference model include the prompt text, and the outputs include the predicted result text corresponding to the prompt text.

[0059] The input of the reward model may include the predicted result text output by the actor model and the reference model, as well as the prompt text and label result text of the sample data, and the output may include a reward value reflecting the matching degree between the predicted result text and the label result text.

[0060] The input of the critic model may include the predicted result text output by the actor model and the reference model, as well as the prompt text and label result text of the sample data, and the output may include a vector representing the reward accumulation value.

[0061] S102, Process the sample data using the first model to obtain the first loss. The first model refers to any one of the reference model and the reward model, and the first model is a pre-trained model.

[0062] In this embodiment, the architecture of the target processing model can be any one of the following:

[0063] Architecture One: The actor model consists of a reference model and a first sub-network. The critic model and the reward model are two independent models. In this case, the reference model is equivalent to the first model in step S102.

[0064] Architecture Two: The critic model consists of a reward model and a first sub-network. The actor model and the reference model are two independent models. In this case, the reward model is equivalent to the first model in step S102.

[0065] Exemplarily, when the architecture of the target processing model is Architecture One, the schematic diagram of its architecture can be referred to Figure 2 .

[0066] Optionally, the first model and the first sub-network can have the same network structure.

[0067] The first model can be a pre-trained large language model base (LLMBase).

[0068] In step S102, the sample data can be processed by the first model to obtain the output of the first model, and then the first loss corresponding to the first model can be calculated based on the output of the first model.

[0069] The specific calculation method of the first loss can be referred to the related technology and will not be elaborated here.

[0070] S103: Process the output of the first model using the first sub-network to obtain a second loss.

[0071] The first model and the first sub-network form a second model, and the second model is one of the actor model and the critic model corresponding to the first model.

[0072] In step S103, the output of the first model can be processed by the first sub-network to obtain the output of the first sub-network, and then the second loss can be calculated based on the output of the first sub-network.

[0073] In step S103, if the target processing model belongs to the case of Architecture One, the second model is equivalent to the actor model corresponding to the reference model, the output of the first sub-network is equivalent to the output of the actor model, and the second loss is equivalent to the loss of the actor model.

[0074] If the target processing model belongs to the case of Architecture Two, the second model is equivalent to the critic model corresponding to the reward model, the output of the first sub-network is equivalent to the output of the critic model, and the second loss is equivalent to the loss of the critic model.

[0075] The specific calculation method of the second loss can be referred to the related technology and will not be elaborated here.

[0076] S104. Determine the total model loss according to the first loss, the second loss, the third loss obtained by the third model, and the fourth loss obtained by the fourth model.

[0077] Among them, the third model refers to the one that is different from the first model in the reference model and the reward model, and the fourth model refers to the one that does not correspond to the second model in the actor model and the critic model.

[0078] In step S104, on the one hand, the data input to the third model can be processed by the third model to obtain the output of the third model, and then according to the output of the third model, the third loss corresponding to the third model can be calculated; on the other hand, the data input to the fourth model can be processed by the fourth model to obtain the output of the fourth model, and then according to the output of the fourth model, the fourth loss corresponding to the fourth model can be calculated.

[0079] The specific process of calculating the loss corresponding to the model according to the output of the model can be referred to the related technology and will not be elaborated here.

[0080] After obtaining the first loss to the fourth loss, the total model loss of the target processing model can be calculated based on these four losses. The specific calculation process of the total model loss can be referred to the related technology.

[0081] In S104, if the architecture of the target processing model belongs to architecture one, the third model can be the reward model and the fourth model can be the critic model.

[0082] If the architecture of the target processing model belongs to architecture two, the third model can be the reference model and the fourth model can be the actor model.

[0083] S105. Train at least part of the parameters of the target processing model according to the total model loss.

[0084] In step S105, it can be first determined whether the current total model loss meets the preset convergence condition.

[0085] This convergence condition can have various forms and is not limited in this embodiment. As an example, the convergence condition can be that the total model loss is not greater than the preset loss convergence threshold.

[0086] If the convergence condition is not met, according to the total model loss, calculate the update gradient of the part of the parameters that need to be trained in the target processing model, and then update the parameters that need to be trained one by one according to the update gradient. After the update, return to step S102 and execute S102 to S105 again until the convergence condition is met.

[0087] If the convergence condition is satisfied, stop the training, and the method for configuring the generative pre-training model GPT in this embodiment ends.

[0088] The beneficial effects of this embodiment are as follows:

[0089] When training the target processing model based on the PPO algorithm, at least merge the actor model and the corresponding reference model, or merge the critic model and the corresponding reward model, so that the merged models share model parameters.

[0090] In some alternative embodiments, in order to further reduce the storage space occupied by the target processing model, with reference to the structures of the first model and the second model, the third model and the fourth model can also be merged. That is to say, construct a second sub-network, and use the model composed of the series connection of the third model and the second sub-network as the fourth model, where the output of the third model is used as the input of the second sub-network, and the output of the second sub-network is used as the output of the fourth model.

[0091] In this case, the architecture of the target processing model can be referred to Figure 3 .

[0092] In Figure 3 In the shown architecture, the reference model is equivalent to the first model of this embodiment, the model composed of the reference model and the first sub-network is equivalent to the second model of this embodiment, that is, the actor model, the reward model is equivalent to the third model of this embodiment, and the model composed of the reward model and the second sub-network is equivalent to the fourth model of this embodiment.

[0093] In this embodiment, the reward model can be a pre-trained basic large language model, and the reward model and the second sub-network can have the same network structure.

[0094] Correspondingly, when adopting the Figure 3 shown architecture, the process of using the third model and the fourth model to process sample data to obtain the corresponding third loss and fourth loss can include:

[0095] Use the third model to process the sample data to obtain the third loss;

[0096] Use the second sub-network to process the output of the third model to obtain the fourth loss, and the third model and the second sub-network form the fourth model.

[0097] The above process of obtaining the third loss and the fourth loss can refer to the process of obtaining the first loss and the second loss in steps S102 and S103 of the previous embodiment, and will not be elaborated here.

[0098] The beneficial effect of this embodiment is that by combining the third model and the fourth model, the number of parameters in the target processing model is further reduced, saving the occupied storage space.

[0099] Meanwhile, when training the target processing model of this embodiment, only the Figure 3 shown actor model and critic model need to be loaded into the memory, and the number of models to be loaded is reduced from 4 in the related art to two, improving the loading rate and reducing the storage space occupied by the target processing model.

[0100] Optionally, during the training process, if the target processing model includes pre-trained models, then in step S105, the parameters of these pre-trained models may not be trained, that is, the parameters of these pre-trained models are not updated.

[0101] Combined with the foregoing example, when the target processing model adopts the model architecture as Figure 2 shown, the reference model therein belongs to a pre-trained model. Therefore, in step S105, only the parameters of the part other than the reference model can be trained according to the total model loss, that is, the first sub-network and the critic model are trained according to the total model loss.

[0102] When the target processing model adopts the model architecture as Figure 3 shown, both the reference model and the reward model therein belong to pre-trained models. Therefore, in step S105, only the parameters of the actor model and the critic model that do not belong to the reference model and the reward model are trained. That is, at this time, S105 may include:

[0103] Training the parameters of the first sub-network and the second sub-network according to the total model loss.

[0104] The specific process of training the parameters according to the loss of the neural network model can refer to the related art, which will not be elaborated in this embodiment.

[0105] The advantage of adopting this training method is that by reducing the parameters to be updated during the training process, the amount of calculation during the training process can be reduced, saving computing resources and improving the training speed.

[0106] In some optional embodiments, the first model and the third model of the target processing model as Figure 3 shown can share a part of the parameters. Please refer to Figure 4 , which is a schematic structural diagram of a target processing model provided for this embodiment.

[0107] In the target processing model of this embodiment, the first model (i.e., the reference model) consists of a shared network unit and a first network unit, where the output of the shared network unit is used as part of the input of the first network unit, and the output of the first network unit is used as the output of the first model.

[0108] At the same time, the third model (i.e., the reward model) consists of a shared network unit and a third network unit, and the output of the shared network unit is used as part of the input of the third network unit, and the output of the third network unit is used as the output of the third model.

[0109] That is to say, in the target processing model of this embodiment, the reference model and the reward model share the parameters in the shared network unit, and at the same time, they respectively use their own unique first network unit and third network unit to process the output of the shared network unit to obtain the output of the first model and the output of the third model.

[0110] In this case, using the first model to process the sample data to obtain the first loss, specifically, it may include:

[0111] Using the shared network unit to process the sample data;

[0112] Using the first network unit to process the output of the shared network unit to obtain the first loss.

[0113] Correspondingly, using the third model to process the sample data to obtain the third loss, including:

[0114] Using the third network unit to process the output of the shared network unit to obtain the third loss.

[0115] The beneficial effect of this embodiment is that by sharing a part of the reference model and the reward model, the storage space occupied by the target processing model is further reduced.

[0116] According to the method for configuring the generative pre-training model GPT provided by the embodiment of the present application, the embodiment of the present application also provides a device for configuring the generative pre-training model GPT. Please refer to Figure 5 For the structural schematic diagram of this device, this device may include the following units.

[0117] An obtaining unit 501, configured to obtain sample data for training a target processing model; wherein, the target processing model includes an actor model, a critic model, a reference model corresponding to the actor model, and a reward model corresponding to the critic model;

[0118] A processing unit 502, configured to use the first model to process the sample data to obtain a first loss, where the first model refers to any one of the reference model and the reward model, and the first model is a pre-trained model;

[0119] A processing unit 502, configured to process the output of the first model by using a first sub-network to obtain a second loss, where the first model and the first sub-network form a second model, and the second model is one of the actor model and the critic model corresponding to the first model;

[0120] A determination unit 503, configured to determine a total model loss according to the first loss, the second loss, a third loss obtained by a third model, and a fourth loss obtained by a fourth model; where the third model refers to one of the reference model and the reward model that is different from the first model, and the fourth model refers to one of the actor model and the critic model that does not correspond to the second model;

[0121] A training unit 504, configured to train at least some parameters of the target processing model according to the total model loss.

[0122] Optionally, the process by which the processing unit 502 obtains the third loss and the fourth loss includes:

[0123] Processing the sample data by using the third model to obtain a third loss, where the third model is a pre-trained model;

[0124] Processing the output of the third model by using a second sub-network to obtain a fourth loss, where the third model and the second sub-network form a fourth model.

[0125] Optionally, the first model includes a shared network unit and a first network unit, and the third model includes a shared network unit and a third network unit;

[0126] When the processing unit 502 processes the sample data by using the first model to obtain the first loss, it is specifically configured to:

[0127] Process the sample data by using the shared network unit;

[0128] Process the output of the shared network unit by using the first network unit to obtain the first loss;

[0129] When the processing unit 502 processes the sample data by using the third model to obtain the third loss, it is specifically configured to:

[0130] Process the output of the shared network unit by using the third network unit to obtain the third loss.

[0131] Optionally, when the training unit 504 trains at least some parameters of the target processing model according to the total model loss, it is specifically configured to:

[0132] Train the parameters of the first sub-network and the second sub-network according to the total model loss.

[0133] The GPT configuration device for the generative pre-training model provided by the embodiments of the present application, the specific working principle of which can refer to the relevant steps of the GPT configuration method for the generative pre-training model provided in any embodiment of the present application, will not be elaborated here.

[0134] The embodiments of the present application further provide an electronic device. Please refer to Figure 6 , which is a schematic structural diagram of the electronic device. The electronic device may include a memory 601 and a processor 602.

[0135] The memory 601 is used to store computer programs;

[0136] The processor 602 is used to execute the computer programs, and specifically used to implement the GPT configuration method for the generative pre-training model provided in any embodiment of the present application.

[0137] The embodiments of the present application further provide a computer-readable storage device for storing computer programs, which, when executed, are specifically used to implement the GPT configuration method for the generative pre-training model provided in any embodiment of the present application.

[0138] The embodiments of the present application further provide a computer program product, including a computer program, which, when executed, is specifically used to implement the GPT configuration method for the generative pre-training model provided in any embodiment of the present application.

[0139] It should be noted that the various embodiments in this specification are all described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other.

[0140] For the convenience of description, when describing the above system or device, it is described by function as various modules or units respectively. Of course, when implementing the present application, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0141] From the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a storage device, such as ROM / RAM, magnetic disk, optical disc, etc., including several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of the present application.

[0142] Finally, it should also be noted that in this text, relational terms such as first, second, third, and fourth are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including", or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article, or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the said element.

[0143] The above are only the preferred embodiments of the present application. It should be pointed out that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present application.

Claims

1. A method for configuring a Generative Pretrained Transformer (GPT) model, characterized in that, it includes: Obtaining sample data for training a target processing model; wherein, the target processing model includes an actor model, a critic model, a reference model corresponding to the actor model, and a reward model corresponding to the critic model; Processing the sample data using a first model to obtain a first loss, where the first model refers to either the reference model or the reward model, and the first model is a pre-trained model; Processing the output of the first model using a first sub-network to obtain a second loss, where the first model and the first sub-network form a second model, and the second model is one of the actor model and the critic model corresponding to the first model; Determining a total model loss based on the first loss, the second loss, a third loss obtained by a third model, and a fourth loss obtained by a fourth model; wherein, the third model refers to the one of the reference model and the reward model that is different from the first model, and the fourth model refers to the one of the actor model and the critic model that does not correspond to the second model; Training at least some parameters of the target processing model based on the total model loss.

2. The method according to claim 1, characterized in that, The process of obtaining the third loss and the fourth loss includes: Processing the sample data using the third model to obtain the third loss, where the third model is a pre-trained model; Processing the output of the third model using a second sub-network to obtain the fourth loss, where the third model and the second sub-network form the fourth model.

3. The method according to claim 2, characterized in that, The first model includes a shared network unit and a first network unit, and the third model includes the shared network unit and a third network unit; The step of processing the sample data using the first model to obtain a first loss includes: Processing the sample data using the shared network unit; Processing the output of the shared network unit using the first network unit to obtain a first loss; The step of processing the sample data using the third model to obtain the third loss includes: Processing the output of the shared network unit using the third network unit to obtain the third loss.

4. The method according to claim 2, characterized in that, The step of training at least some parameters of the target processing model based on the total model loss includes: Training the parameters of the first sub-network and the second sub-network based on the total model loss.

5. A device for configuring a Generative Pretrained Transformer (GPT) model, characterized in that, it includes: An obtaining unit for obtaining sample data for training a target processing model; wherein, the target processing model includes an actor model, a critic model, a reference model corresponding to the actor model, and a reward model corresponding to the critic model; A processing unit for processing the sample data using a first model to obtain a first loss, where the first model refers to any one of the reference model and the reward model, and the first model is a pre-trained model; The processing unit is configured to process the output of the first model using a first sub-network to obtain a second loss, where the first model and the first sub-network form a second model, and the second model is the one corresponding to the first model among the actor model and the critic model; A determination unit for determining the total model loss according to the first loss, the second loss, the third loss obtained by the third model, and the fourth loss obtained by the fourth model; wherein the third model refers to the one different from the first model among the reference model and the reward model, and the fourth model refers to the one not corresponding to the second model among the actor model and the critic model; A training unit for training at least some parameters of the target processing model according to the total model loss.

6. The apparatus according to claim 5, wherein, The process by which the processing unit obtains the third loss and the fourth loss includes: Processing the sample data using the third model to obtain the third loss, where the third model is a pre-trained model; Processing the output of the third model using a second sub-network to obtain the fourth loss, where the third model and the second sub-network form the fourth model.

7. The apparatus according to claim 6, wherein, The first model includes a shared network unit and a first network unit, and the third model includes the shared network unit and a third network unit; When the processing unit processes the sample data using the first model to obtain a first loss, it is specifically configured to: Process the sample data using the shared network unit; Process the output of the shared network unit using the first network unit to obtain a first loss; When the processing unit processes the sample data using the third model to obtain the third loss, it is specifically configured to: Process the output of the shared network unit using the third network unit to obtain the third loss.

8. The apparatus according to claim 6, wherein, When the training unit trains at least some parameters of the target processing model according to the total model loss, it is specifically configured to: Train the parameters of the first sub-network and the second sub-network according to the total model loss.

9. An electronic device, wherein, It includes a memory and a processor; The memory is used to store a computer program; The processor is used to execute the computer program, specifically for implementing the generative pre-training model GPT configuration method according to any one of claims 1 to 4.

10. A computer-readable storage device, wherein, It is used to store a computer program, and when the computer program is executed, it is specifically for implementing the generative pre-training model GPT configuration method according to any one of claims 1 to 4.